Content of review 1, reviewed on February 17, 2025
Increasing attention is being focused on understanding how urbanization affects phenotypic variation. The authors of this manuscript seek to add to this body of work by conducting a meta-analysis on how the traits of two bird species vary within and among urban and forested regions. This manuscript takes previous research a step further by analyzing phenotypic traits across a large spatial scale and using data collected in similar ways. The authors demonstrate support for more variation within and among urban subpopulations, but not more heterogeneous variation in urban areas.
In general, I find the topic interesting, the paper is well-written, and the results are well-supported by the analyses. Combining this vast dataset into one analysis is a large undertaking and provides our best understanding of this system. The heterogeneity in heterogeneity hypothesis is a novel addition and interesting to think about. The figures were useful, although the boxes seemed unnecessary to me. I have a few suggestions to improve the paper, but I did not find anything major.
Main
The study evaluates a large dataset on bird phenotypes collected in similar ways. Though I agree that this approach is desirable, I don’t believe that applying the term, mega-analysis, is warranted, even if others have applied it. Mega-analysis would imply a very large number of studies rather than a large analysis of similar data, so the terminology is off. Even aggregating over many studies would then still be a meta-analysis, although I wouldn’t call it that unless explicitly using results from multiple studies. Many studies aggregate and analyze large amounts of data (all of the ‘Big Data’) and don’t call it a mega-analysis. I don’t think the name adds anything, so I would just remove this minor part.
Subpopulations are defined here as a group of individuals located in the same area. A real subpopulation would be moderately genetically distinct. I understand that this is a way to operationalize a clustering structure for which no genetic information exists, but more should be done to evaluate how these decisions affect results and some caveats to the conclusions. What if we knew that urban subpopulations were bigger or smaller than nonurban ones – how would that change results and interpretations?
H3: heterogeneity in heterogeneity is an interesting hypothesis, but its reasoning is not well-developed in the introduction. Why should cities have more variation in their within-population variation?
I’m a bit confused about how a site is labeled as urban or not given the strong overlap between urban and nonurban sites in impervious surface area (p. 11). I assume that you are using the data owners’ description, which might be imperfect? What happens if you use impervious surfaces and other characteristics to assign subpopulations to urban vs. nonurban? Do results change? Get stronger?
Minor
L 86-7, add: ‘or performing common garden or transplant experiments.’
L 278, but only 1 individual was assessed (L 274), then how is within-individual variation captured?
L 315, 352, why were the Harjavalta and Barcelona datasets excluded?
L 419, delete ‘population’ – happens at various levels above and below population
L 438, add ‘and evolutionary’
L 441-3, but the first reason is environmental variation rather than higher plasticity? Assuming similar plasticity and more environmental variation, then you get the same patterns. I would think this is the most likely reason.
L 577, change negative to non-significant – negative would imply a directionality to me.
Source
© 2025 the Reviewer.
Content of review 2, reviewed on April 30, 2025
I read this manuscript previously and thought the questions and approach were strong, but some aspects needed clarification, modification, or evaluation for sensitivity. Although the manuscript has improved in clarity, some issues remain.
I take issue with the response to reviewer 1, who called out the authors on claiming support for hypotheses when credible intervals overlapped with the expected values. The authors’ response was: “We see CI’s excluding zero as strong statistical evidence in support of our hypotheses, while cases where CI’s overlap zero could be interpreted as weaker evidence rather than no evidence.”
I also use Bayesian statistics, but I do not think this method frees one from adopting a priori thresholds for determining support for a hypothesis. One should distinguish between using credible intervals to indicate certainty in parameter estimation from using statistics to determine support for a hypothesis. I think we need to continue to support thresholds in statistics for hypothesis testing. If CI’s overlap zero, then there is no support for the hypothesis. One-sided tests are only appropriate if the opposite direction cannot be true, which I don’t think is the case here. If it’s close, then I might discuss it as not being supported but close, and I would not suggest support for the hypothesis as moderate evidence, no matter how much we want our hypotheses to be true.
I raised concerns about using mega-analysis for this study since the term does not seem to fit. The authors suggested that they will use it anyway because a few other papers have used the term and to differentiate it from standard meta-analyses, which are often (but not always) conducted on studies with different types of data and summary statistics. In the end, it’s a semantics debate, so I won’t hold up the paper on that account, but I continue to feel that promoting this term will only lead to confusion.
I had hoped they could test the sensitivity of results to different spatial definitions of subpopulations, but they did not. Based on Table S4, it does look like it would not matter very much given a changing definition, but it would have been nice if they had tested it.
I questioned what would explain H3: heterogeneity in heterogeneity. The authors added some text, but no citations. If there is some discussion of this hypothesis and its mechanisms in the literature, it would be great to add that. I suggest also providing an example of a mechanism in the city that would increase the variation within subpopulations differently across subpopulations.
Source
© 2025 the Reviewer.
References
J., T. M., A., M. J. G., C., B., J., B., J., B. C., P., C., J., D. N., M., D. D., M., E., T., E., L., E. K., C., I., A., L., S., M., E., M., A., M., S., P., C., S. J., G., S., M., S., E., V., H., W., D., R., A., C. 2025. Continental Patterns of Phenotypic Variation Along Replicated Urban Gradients: A Mega-Analysis. Ecology Letters.
