Content of review 1, reviewed on July 01, 2024
Review of ECOG-07520
Trait-based ecology depends upon values for traits for entities (individuals of species or higher taxa), yet trait data are sparse. This manuscript reports investigations into a common method of imputing missing values of trait measurements for species.
Specifically, the paper uses and recommends leave-one-out cross validation and calculation of indices of discrepancy between imputed and held-out observations as indices of precision, and bias. This is important work as the uncritical use of large databases and gap filling procedures have the potential to lead to erroneous conclusions.
It considers two contexts: (a) analyses of the ensemble (e.g. trait-trait correlations, summary statistics and; (b) imputation of specific values for entities for some other use, e.g. calculation of community-weighted-means, traits as variables in explanatory or predictive statistical models. While these two distinct use cases are usefully differentiated, specifics are not part of this paper. That is, the impact of discrepancy on the inferences from downstream analyses is not examined.
The key message I got was that imputation can be fairly safely used for analyses of a whole dataset. But that use for specific prediction of values for species (or other entities) needs one to be more circumspect. How much so is not clear.
I see much of value in this paper, though I feel it fall short of it’s potential to be useful and influential. The main issues I have are the following, along with several specific comments.
a) Cross-validation is not new, it is absolutely a good method to be using here, and it is perhaps worth referring to some related work. The literature of Species Distribution Modelling has paid particular attention to how stratification might be used. Importantly, ecological datasets are heterogeneous and there is structure in those datasets that may inform cross-validation procedures. See Roberts, D. R., Bahn, V., Ciuti, S., Boyce, M. S., Elith, J., Guillera‐Arroita, G., ... & Dormann, C. F. (2017). Cross‐validation strategies for data with temporal, spatial, hierarchical, or phylogenetic structure. Ecography, 40(8), 913-929. Valavi, R., Elith, J., Lahoz-Monfort, J. J., & Guillera-Arroita, G. (2018). blockCV: An r package for generating spatially or environmentally separated folds for k-fold cross-validation of species distribution models. Biorxiv, 357798.
b) The presentation of the paper with several large tables is not greatly effective. Graphical presentation may help. E.g. imputed vs observed plots. Estimates and error bars. Are confidence limits the quantity of use here or should it be prediction intervals? In the case of estimating specific values the species trait matrix, I would suggest it should be predictive intervals , while for summaries of the whole data sets, CI’s are appropriate.
c) Substantive interpretation of the suggested indices. Are they independent, is there redundancy in them? Could you suggest bands of the scale of “good” acceptable with caution” , “poor, not to be trusted” or something similar?
d) Box 1 is underused, underspecified and introduced too late. Redundancy is a broad term and could usefully be dissected into e.g. within species records, individual based records, missingness.
Specific comments
Title
The title appears incomplete. it makes no sense to end with the verb “using”. Using what?
L18 redundancy in what sense?
Introduction
L43 Singular/plural agreement
L49 it is probably worth mentioning here that the BHPMF estimates a probability distribution of trait values, rather than a single value. Just to make the point clearly.
L60ff I found the review of what has been addressed in trait imputation for ecology a bit meagre, and so the justification for this paper (as distinct from other works ) is unclear. What did Joswig et al do and how is that different from what you propose?
Joswig, J. S., Kattge, J., Kraemer, G., Mahecha, M. D., Rüger, N., Schaepman, M. E., ... & Schuman, M. C. (2023). Imputing missing data in plant traits: A guide to improve gap‐filling. Global Ecology and Biogeography, 32(8), 1395-1408.
some other works:
Johnson TF, Isaac NJB, Paviolo A, González-Suárez M. Handling missing values in trait data. Global Ecol Biogeogr 2020; 30: 51–62. https://doi.org/10.1111/geb.13185
Palma, E., Vesk, P.A. & Catford, J.A. Building trait datasets: effect of methodological choice on a study of invasion. Oecologia 199, 919–935 (2022). https://doi.org/10.1007/s00442-022-05230-8
L62 It would be useful to end the Intro with a statement of Aims or research questions.
L68ff paraphrase to introduction.
GSPFF: really do we need this initialism?
L87. Importantly, not all these traits were collected by the original authors, by the same methods or same teams.
You do not say what you did when you have multiple measurements of a given trait for the same species. These are critical decisions dependent upon the final use for the trati data. (e.g. see the above refs).
L98 L98 Ok so this is the crux of your evaluation. you evaluate all species one by one. There is information in those different species, which ones are predicted well, and which poorly and how that relates to the structure of the data sets (e.g. phylogeny or position upon trait spectra.)
L103 this first step is not written as an analysis. it constrains the range. An analysis would record all instances of imputations falling outside the observed range.
L106 minor point, but in my opinion it makes more sense to treat your prediction (imputed value) as the input or independent value.
L109 perhaps a short argument for justifying what indices might be useful ahead of presenting the indices. E.g. a good imputation would be one where imputed values are close to the (withheld) observed values. Further, imputed values should be closer to the observed value than observed values are to the grand mean. Etc, or whatever you believe.
L114 by ‘element’ do you mean ‘entity’ as used previously?
L114 in the equation (unlabelled) you write O’, but in the text you define Ō bar. Is this the mean across all entities?
L116 this is good, but could be built upon to suggest further substantive interpretation.
L133 this linear model does not seem to be reported.
L141 it would be good the define and justify what you mean by “redundancy”
L168ff please do not fall into the fallacy of interpreting accept/reject hypothesis tests as measures of accuracy or effect. You have large sample size, though varied across your traits. Your statistical power will be large, but perhaps variable between traits. The estimates of those intercepts and slope are more meaningful (along with their uncertainty). Plotting these would be helpful.
L173ff use of the word ‘substantial’ would seem to benefit from some quantitative evaluation or graphical presentation. How large a departure matters? what is insubstantial?
L175 biased…in varied ways that are hard to generalise over.
L179 “species”. “GSPFF”—really is this needed?
L190 is this is drawing on the results of your linear modelling? If so, the presentation is insufficient. table, estimate, uncertainty and plots.
L212 height measured on an individual is a far different attribute from the traits such as Seed Mass and Stem specific Density.
L217-227 this para graph about the way you constructed the 3 Baraloto datasets and the impact is useful, but I found it hard to get in the first, or second reading. Sorry I have no useful suggestion, other than greater specificity throughout the paragraph.
L243 at last, a figure, but it is probably the least important result you have. Ok, knowing that the estimated uncertainty from the MC samples is not a useful indicator of the likely discrepancy is useful. But my point is all those tables are ineffective communication devices. They just serve to compactly state values.
L253 a word missing here. it doesn’t make sense.
L254 agreed, but you miss the opportunity to clarify what those uses are and how others have examined them
L265-276. this is the critical point of this paper. Imputing individual elements is problematic. it would be good to point to some other work examining how to construct a trait dataset. e.g. Palma et al. depending upon purpose of the analyses.
Box 1 this doesn’t seem to be referred to. (until very late)
redundancy is used in a broad sense. is there not value in specific forms of redundancy and estimating their relative contributions?
e.g overall completeness of the entity by trait matrix.
completeness of rows, completeness of traits. within species measurements. Phylogenetic structure (many species in a genus).
Table 4 what are the two versions of the individual and individual2 datasets?
Fig. 2 this is ok, but could be improved by considering a substantive interpretation of the indices, into good, caution, bad/avoid…
Table 1 do the last 2 columns record all imputed values falling outside the empiric range? say so clearly.
Table should report n. but I feel this would better be presented as graphs
Source
© 2024 the Reviewer.
Content of review 2, reviewed on November 18, 2024
Review of ECOG-07520.R2
I am reviewing this after revision. The work presents an analysis of trait imputation by various methods and its evaluation through leave-one-out analyses.
The work is thoughtful and incisive and presents new information.
The revision has addressed many issues with the earlier ms.
Notably a expanded set of imputation methods, greater literature review, greater emphasis on graphical display of results, improving communication.
There are a small number of responses I disagree with, but which may fall into differences of opinion.I include them mainly by way of explanation. Not because I believe they need addressing.
I feel the authors missed an opportunity to investigate residuals. It would be informative to analyses what explains those residuals. But that may be beyond the scope of the current ms.
I’m disappointed the author has chosen to disagree with both Reviews seeking some substantive interpretation of performance metrics, on the grounds that they are subjective. Science is subjective. You are best placed to guide interpretation. Alpha =0.05 is arbitrary; Fisher was asked what false positive rate he’d consider important. But I digress, I consider the author’s decision to just show the data, ok.
The response to the review the author restates a binary interpretation, of whether a method is biased (or not). The point is the larger that sample size the greater the power. The rejection of the null, offers little insight into HOW biased the method is. That is why the estimates (with uncertainty) of slope (bias) are more valuable.
Specific typos/minor points
Fig 5 is ok, but “our method should be used” is little use. Please replace with essence of method.
L209 is this “individul2”? please define.
L219 this sentence and the next needs rewriting. (suggest “least” for “less” and “all” for “al” and not sure what is meant by the latter part of the secnd sentence.
L240 Sentence needs inclusion of something like: “not one of the following measures of variability ….”
L286 remove “a”
L288 unclear meaning “what same property”?
Source
© 2024 the Reviewer.
References
Damian, G. L., Jesus, A., C., S. F., G., S. N., Boardman, K. N. J., Schwantes, M. B., R., B. T., Ferreira, d. L. R. A., Emilio, V., Esteban, a., Monteagudo, M. A., Flores, L. G. R., Manoel, d. S. R., Gerhard, B., Alejandro, A., Gonzalo, R., Hirma, R., Santos, P. N. C. d., S., M. P., Sabina, C. R., A., d. C. W. J., Mathias, D., Anthony, D. F., Hur, M. B., R., F. T., Yadvinder, M., L., P. O., David, G., Sandra, D. 2025. Use and misuse of trait imputation in ecology: the problem of using out-of-context imputed values. Ecography.
