Content of review 1, reviewed on December 20, 2021
Sadoul et al. present an extremely interesting study that is providing important insights into the feedback loop between behavior, the environment, and physiology. Their experiments are well-designed and, crucially, allow for assessments of behavior-environment interactions over a longer period of life history than prior work. I am altogether enthusiastic about their work but do have some concerns.
Various aspects of methodology are at times confusingly presented or inadequately detailed. I have outlined specific examples in the detailed comments below.
There are also some aspects particularly of the transcriptomic analyses and interpretation that could be improved or better-justified. The bulk of the intepretations rely on GO enrichment analyses that have methodological issues and are also largely underpowered. Enriched GO categories appear to be based on uncorrected p values, which makes the false positive rate a substantial concern. There are a few things the authors could do to mitigate these concerns:
- If there is a true signal of biological underpinnings of transcriptome differences between bold and shy, there should be an enrichment of low P values in the GO enrichment analyses. Visualizing the enrichment P values using quantile-quantile plots would therefore help alleviate concerns. Given the number of enriched terms (0 in brain, 2 in pituitary, 9 in head kidney), I suspect that the overall signal is quite reak.
- The power of the hypergeometric tests is limited by the number of differentially expressed genes in each analysis (6 in brain, 556 in pituitary, 141 in head kidney), which are determined based on thresholding (padj < 0.05). I think in these cases a reasonable compromise could be to relax the threshold for differentially expressed genes, which has the drawback of increasing the false positive rate beyond 5%, but with the tradeoff that potentially many more genes pass the threshold and improve power of the hypergeometric tests (since false positives should be random with respect to GO terms, this could be a reasonable tradeoff—with the limitations clearly discussed). An alternative or complementary approach could be to use threshold-independent tests such as a KS test to improve power.
I do think that the lack of correction for multiple hypothesis testing is important. If this is not possible, the interpretations based on these findings should be curtailed accordingly and the limitations clearly discussed.
The data accessibility is insufficient. Raw sequencing reads and processed data (i.e., gene expression matrices) should be provided through NCBI GEO or other respositories. There are several files in the supplements that have extension CSV (comma seperated values) but are actually semicolon-separated (which is fine, but should be made clear). Numerous column headers should also be defined. A README documentation file would address these issues and be much appreciated.
Detailed comments:
I was not familiar with the "head kidney". A brief description would likely be useful for non-experts, as well as a discussion of its homology (e.g., to mammalian adrenal gland) and implications for cross-taxonomic insights.
The abstract (lines 24-25) refers to "chronic stress" as an environment. Environmental variation can induce chronic stress, but stress is not an environment. Rephrase.
Fish were characterized as bold vs. shy at GRT1 (255dpf) based on thresholding of their latency to exit (greater or less than median). These characterizations are used throughout the manuscript. What does the distribution of this measure look like and does it make sense to analyze it as a binary variable? Could it make more sense to treat boldness as a continuum, i.e., by modeling as a continuous variable?
Lines 205-208. More details on the tissue sampling (are the entire organs sampled?) should be included. The total sample size should also be mentioned. I believe it is 30 (5 individuals x 2 behavioral phenotypes x 3 tissues) but it was not clear. Were all gene-expression individuals sexed? What was the sex breakdown and was it balanced between bold and shy samples?
Lines 306-308. The mean latency differed over time, which the authors took to mean that risk-taking behavior is plastic, potentially varying due to environment or the life course. I do not find this very convincing. The measure of risk-taking behavior (latency to exit) can vary due to many factors that do not necessarily relate to risk-taking. One analysis of interest could be to compare the mean latency for GRT experiments which were conducted on separate tanks (e.g., GRT2). Are there any differences in mean latency between tanks? This could be taken as a measure of random variation, which would provide important context for the variation in latency over time.
Lines 386-387. This sentence states the gene expression analyses (genes related to learning and memory have lower expression in shy individuals) support other studies finding that bold individuals are faster learners. This seems like a stretch that is not justified by the evidence presented here. Rephrase.
Lines 407-411. Similar to the previous comment, this interpretation is also dubious. First off, immune function is a complex phenotype comprising tradeoffs among various function. "Higher immunity" is a weak variable. Second, genes related to immune function can be positive or negative regulators of particular immune functions. Higher expression in the immune system is therefore not necessarily an indication of increased immunity.
Figure 1: This is a useful figure but confusingly presented. A few issues:
- Could the events be written out? There seems to be ample room to do so, making the abbrevations and the key unnecessary.
- Apart from the chronic stress test, the colors are undefined.
- N appears to refer to the number of groups/tanks, but this is unclear. Could this be defined?
Figure 3: The background gradient for the hypoxia challenge period is unnecessary (isn't the information redundant with section D?) and makes the data harder to read. I would suggest removing it and indicating the start of the hypoxia challenge using a vertical line.
Figures 4-5: I would suggest using an improved color gradient for heatmaps. The viridis scale is a great percetually uniform option (available through the viridis R package.
Supplementary Material and Methods: The distinction between methods in the main manuscript and in the supplements seems a bit arbitrary. The four models described in SMM1, for example, are numbered and referenced in Section 3.1 (lines 231-246) but are not defined in the main manuscript, making this confusing to follow. A better approach might be to briefly and succicntly describe the overall approach (including listing the models) in the main manuscript, and then relegating technical details to the supplements. Along the same lines, several modeling decisions (e.g., behavioral analysis and gene expression analysis) are distinct decisions that are highly relevant to the findings of the paper but are only found in the supplements.
SMM1: For Bayesian models, what was the rationale for the burn-in vs sampling interval (2500 iterations each)? Was there any assessment of convergence of these models?
SMM4:
"genes with less than 25 reads (cumulating all the analyzed samples) were filtered out": Why this number? This threshold is extremely hard to interpret because it's dependent on the degree of sampling. What does it correspond to in terms of CPM or TPM?
The use of the mouse genome (mm9) for functional annotation of genes is justified, but the annotations appear to be entirely based on the shared use of gene names between species. This is a problem, as the nature of orthology has a strong bearing on the extent to which functions are conserved. A better approach would be to use existing orthology databases (e.g., in ENSEMBL), to assign genes to curated orthology databases (e.g., TreeFam), or to match based on reciprocal best hits between genomes.
Source
© 2021 the Reviewer.
Content of review 2, reviewed on March 07, 2022
I thank the authors for responding to my concerns. The revised manuscript is much better-organized and improved. I am satisfied by the authors' responses to my comments and only have a handful of minor comments. I am happy to recommend this work for publication.
Line 51: "observed to adapt" makes it unclear if this sentence is referring to plasticity or natural selection. Suggest rephrasing.
Line 71 has an unpaired parenthesis.
Lines 400-401: I think still think "higher immunity" should be rephrased, especially given the following sentence. Perhaps "altered immune function"
Figure 2: "Black points" in the caption does not match the plots, which are in color.
Source
© 2022 the Reviewer.
References
Bastien, S., Sebastien, A., Conor, G., Marine, P., Stephanie, R., Benjamin, G., Marie-Laure, B. 2022. Transcriptomic profiles of consistent risk-taking behaviour across time and contexts in European sea bass. Proceedings of the Royal Society B: Biological Sciences.
