Content of review 1, reviewed on April 21, 2022

This is an impressive study that examined the effect urbanisation on avian life-history trait means and variances via a newly derived effect size. This meta-analysis used variance difference effects (also the traditional mean difference effect sizes) and found interesting results – urban birds have higher laying date variation than non-urban birds. Generally, I like the idea of this research synthesis. I also like the writing of this manuscript. However, I have concerns on the reporting and methodologies in this manuscript. Ecology Letter has a high publication requirement. Before formal acceptance, the authors need to address the following issues carefully.

Abstract:
Generally, the current instruction is not well written. The core idea of current introduction relies too much on one review – Thompson et al. 2022. It also lacks a full justification of their predictions (hypotheses). I would suggest the authors changing their framing and address their questions fully. See bellow for some specific comments.
Lines 43 – 45, the authors said the hypothesis ‘urban populations might have higher levels of phenotypic variation than non-urban counterparts’ has never been rested across taxa. However, in following what the authors said and did are contradictory – the authors also only used avian to test this hypothesis. I did not see any ‘across taxa’ stuff. Please be careful with your wording. Also see other places, e.g., lines 77 – 78.
Lines 55 – 56: pls clarify or extend the latest sentence in the abstract – What are the important implications for the understanding of how species adapt to urban environments? The current sentence is too broad, which did not convey any useful information.

The second paragraph (lines 74 – 80) is too small and unbalanced comparing with other paragraphs (especially the following paragraph). Pls merge this to other paragraphs or extend it. I suggest the authors to introduction the importance of phenotypic variation in the field of ecology generally and then switch to your specific field – urbanisation. At least you need to acknowledge previous works which used the similar meta-analytic technique (i.e., meta-analysis of variation) to investigate the effects of selection pressure on the phenotypic variation.

The authors declared that they used within species urban – non-urban comparisons to test their predictions. So it is not true that they said they test phenotypic variation hypotheses across species.

Method – literature search and data extractions
The authors did not do a good job in searching literature and data extraction. Their literature is incomplete and not comprehensive. The data extractions are not transparent and rigorous. There are four major flaws in their literature search, which might lead to literature bias.
First, I think it is a basic rule for an ecological meta-analyst to use multiple databases to search aimed literature. At least the authors should use two major databases: Web of Science (WOS) and Scopus. Because different databases catalogue different literature sources and have different breadths. There are a lot of papers comparing the coverage of different databases, for example, for some specific fields, overlap between Web of Science and Scopus can be as low as 40–50%. See
P. Mongeon, A. Paul-Hus. The journal coverage of Web of Science and Scopus: a comparative analysis Scientometrics, 106 (2016), pp. 213-228

Second, the authors should have performed benchmarking to ensure that your search string is comprehensive. Otherwise, I do think the authors’ literature is comprehensive. I give you a simple and small example to show the process of benchmarking. You can manually searched some papers regarding your topics (based on your experience) and then perform a pilot search in the Scopus database (or WOS) to test your search strategy (e.g., string). For example, if you have 10 benchmark papers, but your search string can only determine 5 benchmark papers in Scopus. It seems that your search string is not comprehensive, then you need to adjust your string. For the similar search string, if you transfer them into WOS, you may capture 7 benchmark papers. The 5 from Scopus and 7 from WOS together may capture all of your benchmark papers. This is also the reason why you should use multiple databases. Please do search your literature comprehensively. The reference bellow is a good one to read.
Foo Y Z, O'Dea R E, Koricheva J, et al. A practical guide to question formation, systematic searching and study screening for literature reviews in ecology and evolution[J]. Methods in Ecology and Evolution, 2021, 12(9): 1705-1720.

Third, the authors should have collected gray literature such as theses and governmental reports. This is an important process of literature search when conducting a systematic review. Please add gray literature in your revised paper.

Fourth, please clarify which authors extracted data from your eligible papers. In current method section, no relevant descriptions. The authors only said “we extracted …..”. I have no idea of “we” means whom? From the author contribution, I can see only one author did all data extractions. So I wonder how the validity of the data extractions can be guaranteed? The most robust way is to let two authors do it independently. I do not think this is not realistic for this paper (65 papers and around 400 effect sizes do not have too much workload). Alternatively, to reduce extraction bias, but they should ask another author (who should have more experience) to check the accuracy of the extractions. They should at least double-check 10% of the extractions of the eligible papers. If the extraction consistency is very high, let’s say, the ratio of error between two reviewers is less than 10%, the extract ions are valid. Otherwise, the second reviewer should extract data independently.

BTW, I also did not see any descriptions of literature screening. How many reviewers screened the initially collected papers? What are the procedure of literature screening? Please clarify in your paper.

Method – statistical analysis
First, the authors used two sets of effect size, log response ratio (lnRR) and the log coefficient of variation ratio (lnCVR), to investigate the phenotypic mean and phenotypic variation. It is good that the authors can take advantage of the newly derived variance difference effect sizes (e.g., lnCVR, lnVR, lnSD). They author know the basic rules of how to choose different variance difference effect size measures – when there is a mean-variance relationship, a mean-adjusted version should be used (i.e., lnCVR). However, the authors neglected some basic rules for calculations of effect size. Theoretically, for the calculations of lnRR and lnCVR, only normally distributed and ratio scale data is applicable. The authors examined three major outcomes – breeding phenology (i.e., lay date) and reproductive effort (i.e., clutch size, number of fledglings). Mathematically, all of them are not generated or sampled from a normal distribution. If your sample size is very large, it is kindly fine to use them as normal data to calculate effect sizes. But it really depends on your data itself. So, I would suggest the authors conducting a sensitivity analysis to examine the robustness of their results.

(i) Transform their data prior to effect size calculations. For count data, log-transformation is a good option. The transformation is a bit difficult for a researcher with less statistical knowledge. Because both the point estimate and sampling variance estimate need to be transformed. The authors can use delta method to derive the corresponding transformation formulas. See
Nakagawa S, Johnson P C D, Schielzeth H. The coefficient of determination R 2 and intra-class correlation coefficient from generalized linear mixed-effects models revisited and expanded[J]. Journal of the Royal Society Interface, 2017, 14(134): 20170213.
Alternatively, the authors can use formulas in following reference to transform their data:
Lagisz M, Zidar J, Nakagawa S, et al. Optimism, pessimism and judgement bias in animals: A systematic review and meta-analysis[J]. Neuroscience & Biobehavioral Reviews, 2020, 118: 3-17.

(ii) The authors chose lnRR as their mean difference effect size over others. I agree with their choice because lnRR is more powerful than SMD (e.g., Hedge’s d) and has many good statistical property:
Yang Y, Hillebrand H, Lagisz M, et al. Low statistical power and overestimated anthropogenic impacts, exacerbated by publication bias, dominate field studies in global change biology[J]. Global change biology, 2022, 28(3): 969-989.

But I still suggest the authors do use SMD as a complementary effect size to check the robustness of mean differences. They can report the corresponding results as supplementary figures or tables. If the authors have concerns on the effect on heteroscedasticity on SMD, they can use account for the heteroscedastic population variances in the two groups – SMDH, which can be easily calculated in escal() in metafor package.

Second, it is very good that the authors used a multilevel model to account for the non-independence in the effect size and account for the multiple sources of heterogeneity. They used give random effects, publication identity (i.e., among-study variation), population identity, phylogeny and species identity (i.e., among-species variation not explained by phylogeny), as well as a residual variance term. The authors should use (log) likelihood test to examine whether population identity, phylogeny and species identity can improve the model quality (e.g., less AIC). For a complex model, it usually requires a large dataset. If population identity, phylogeny and species identity can not contribute to the model quality, a more parsimonious model is preferred.

Fourth, for the publication bias test, I do think they need to test the publication bias of lnCVR. Because statistical significance, rather than variance difference, drives publication bias. Therefore, theoretically lnCVR are assumed not affected by publication bias in the way mean difference effect sizes are. Please refer to the following papers for details:
Yang Y, Hillebrand H, Lagisz M, et al. Low statistical power and overestimated anthropogenic impacts, exacerbated by publication bias, dominate field studies in global change biology[J]. Global change biology, 2022, 28(3): 969-989.

Results
The current result section lacks an important aspect of a meta-analysis – the summary of your data. For example, how many moderators, how many levels of each? Please follow the reporting of meta-analysis properly. Please add this important section

When presenting results, the authors provide too many details in the results section. Please reduce them and make the results section succinct.

Discussion
The discussion section was well written.

Source

    © 2022 the Reviewer.

Content of review 2, reviewed on August 17, 2022

The authors have done a great job in revising the MS

Source

    © 2022 the Reviewer.

References

    Pablo, C., J., T. M., Alfredo, S., Yacob, H., J., B. C., Denis, R., Anne, C., M., D. D. 2022. A global meta-analysis reveals higher variation in breeding phenology in urban birds than in their non-urban neighbours. Ecology Letters.