Content of review 1, reviewed on November 13, 2024
MEE Attribution paper review
The paper is timely, and spans many areas and including many key references. The paper aims are timely and important to “Here, we identify the challenges and decisions involved in detecting and attributing biodiversity change and provide a guide for selecting suitable methods based on available data and specific research questions.” However, I don’t feel that in its current form the paper meets its goals.
However, I have a couple major issues with how the solutions and challenges are discussed; including an important technical critique in how the authors are using and describing casual discovery methods – I think the use is incorrect.
I have a few large comments and concerns:
1. Attribution in this context has not been defined in the paper. The idea of the driver and response and what we are trying to establish and why is not set up well or defined well. So, my largest comment is the paper needs a definition and clarity in the aims of biodiversity change driver attribution (and directional detection) and therefore clarity in methods for estimation versus discovery which is blurred in the paper (e.g. CCM is discovery; would you suggest using SEM for a single driver (or interactive driver) response model? SEM makes much stronger assumptions for casual inference that all paths are correctly specified. You end with a section on “To attribute or not to attribute” but really have dug into what you’re attributing. Instead there is a laundry list of estimation and various methods but it could be connected more to the context and goals at hand throughout the paper.
Related to the comment above, clarity in methods for estimation versus discovery which is blurred in the paper. CCM is a causal discovery method, not an estimation method or for identifying directional effects. It’s place on figure 2 is not correct, nor how it is discussed in the paper. See papers by Runge et al. Discovery is not appropriate for causal estimation of magnitudes of effects, which is how the authors implicitly characterize the tasks of attribution or even detecting if seeking directional effect. For instance, the paper describes detection as “ the focus is generally more on estimating the magnitude and rate of change rather than a binary classification of change or no change. At Line 244” Rather, using CCM would be a step earlier in a workflow to identify potential drivers (i.e. where you have causal discovery). The relationship between CCM, s-map, EDM, and their ability to differentiate causal effects from confounding ones and mediators, and the assumptions required for casual interpretation, are not well established yet to my knowledge. Granger Causality is predictive, not causal, and has nothing to do with the well-established frameworks of the Structural Causal Model and Potential Outcomes framework which are the predominant frameworks for causal inference and graph-based discovery in other fields (558-561). All mention of it should be cut from the paper in my opinion. I’d actually suggest putting 558-574 as a discovery method and removing any mention of Granger causality. It’s not causality as being discussed here, nor in the literature in other fields regarded as such.
Step 4: Evaluation and Interpretation is weak and missing the mark. It takes about why multimodel inference approaches aren’t good for causal inference, but not how causal models are evaluated. They are evaluated based on the strength, plausibility, and validity of the assumptions required for the data, study design, and context. Causal inference - attribution and detection of directional effects - requires assumptions. These are touched upon for each of the methods to some extent (e.g. did a great job for IV..) but should be emphasized more. They are critical for causal interpretation and required for any causal analysis.
How does the section on 755 on Classic ML interpretability methods relate?
What about measurement error which also introduces biases and challenges attribution and even directional detection?
In Section 1 on remote sensing, Can you connect opportunities with remote sensing back to the spatial and temporal resolution issues -- and for which drivers and biodiversity datasets each remote sensing product could be used for? A table of remote sensing products that can be used for data for drivers, and their spatiotemporal resolutions would be a super helpful contribution here. It would greatly enhance the usability of this work for the stated goal of “selecting suitable methods based on available data and specific research questions.” – for which the paper currently falls short of achieving. Lines 230-237 are too generic to be useful. Similarly, the discussion at lines 323-327 relates to the remote sensing section but it is split in multiple places (this is in the detection part) and the dots are not connected yet for the reader.
The separation of bias (is the estimate effect from the estimator true on average as the real causal effect?) vs inference (e.g. how are standard errors modeled? Do we have the power to detect effects?) are not clearly outlined in this paper. For example:
254-257: how do common approaches deal with clustering and serial correlation in the data, which no double effects standard errors, power, and inferences?
In general, the challenges in the data section are more emphasized than how we can overcome them and practical suggestions, which I thought would be the focus of the paper in equal measures. The data structure section at lines 199, for instance, largely just repeats the introduction. The body of the paper is repetitive and inadequate on the solutions. Reorganization to reduce redundancies could add space to go into the solutions more (as promised in the papers exciting aims that aren’t met yet): For example:
- Section 282 – largely repeats with the intro – yet solutions as mentioned here are inadequately described. ).” If you are going to say this paper is going into solutions, this is an area that would be exciting to expand. “Solutions to potential spatio-temporal biases due to data gaps include subsampling, weighting or imputation techniques (Bowler et al., 2024; Nakagawa & Freckleton, 2011).” And “mass, so bias correction methods need to be able to address this specific type of non-random bias (Johnson et al., 2021; Sandel et al., 2015; Schrodt et al., 2015”
- 316-317 – This is repetitive with the previous section on data. “ Biodiversity data are inherently noisy due to genuine fluctuations in biodiversity as well as sampling artifacts” – Can you make this section in the body of the paper on detection more about the detection methods and process?
- 347-349: This text mentions solutions here but says nothing about them! “Recently, frameworks have been developed to characterize non-linear trends
(Rigal et al., 2020) and abrupt shifts (Pélissié et al., 2024) using structured time-series, offering more consistent ways of characterizing the variability in trends.” Expand here. Can the discussions of CATEs be better linked to the aims – the drivers and response? In general, you have a laundry list of estimation approaches but don’t say when you would use one over another. IV estimates a local average treatment effect (LATE), as does RDD, versus matching which often estimates the ATT (average treatment effect on the treated) vs a CATE.. What is the relevant estimand for attribution studies in ecology? What is the interpretation of the output of CCM in this context?
Finally, I think the writing, clarity, and over-use of overgeneralizations can be improved throughout as described below in specific comments.
Specific comments:
10-11 I would say related to attribution, yes, but this statement is overly general and not true. There has been a lot of progress made in this area in recent years, including in many you cite in the paper like Butsic et al., 2017, Larsen et al 2019 Methods in Ecology and Evolution, Wauchope et al., 2022, etc. “However, practical guidance on data and model selection for causal inference in ecology, which deals with inherently complex systems, are still lacking.”
18 – no mention of measurement error
19 – the use of the term structured data is throughout the paper but its never defined until much later (lines 102) – also do you just mean standardized samples and longitudinal data?
28-29 not clear here how your use of terms and phrasing “constructing theoretical causal models a priori” versus “ full causal models” differ.
Line 33 – such as what tools?
52-53 – what are the attribution methods that you are saying ecologists are currently relying on? Bold claim with not a lot of back up here. Again, attribution in this context has not been defined in the paper.
57- this line doesn’t make sense as a challenge: “including the complexity of establishing a causal framework in ecology.” There are several widely used frameworks for causality in ecology and in general that are in use with the Structural Causal Model and the Potential outcomes framework both being used in ecology and in conjunction.
58-63: this is a laundry list of separate things, consider revising and breaking up into multiple sentences. This paragraph in general jump back and forth among ideas and could benefit from being reorganized for clarity and logic.
70: key point that is buried -- ‘defining a baseline for detecting change is challenging in itself.”
79 – “becoming louder” What do you mean here? “used more” , “adopted” ?
89-90 this is a nice wrap up sentence and made me think perhaps this third paragraph should go second, and this sentence provides a nice bridge to the challenges
102 – this type of data has a name and its longitudinal or panel data.
103 – not clear as written what you mean “and sometimes with a formal design for representative spatial sampling (“spatial structure” in Fig.”)
110 – also the fingerprint of colonialism and power dynamics in biodiversity monitoring data – see and cite Chapman, M. et al. (2024) Science.
114—“though the importance of each will depend on the type of change being detected and the driver and response in question” you have yet to introduce the ideas of driver and response in this context.
120 – what do you mean structured data, like longitudinal? Therefore, not totally clear what this means either “Semi-structured data, such as eBird (Kelling et al., 2019), combine elements of structured and unstructured data, allowing variable data collection protocols but with improved documentation and interoperability.”
183-184: could you provide some examples here to make this more concrete?
185-188: Why? Needs to be unpacked.
188-189: Why is climate mor straightforward than harvest? Data availability? Endogeniety?
189 – the promise of remote sensing is a new idea and new paragraph. Can you connect this back to the spatial and temporal resolution issues ? And for which drivers? A table of remote sensing products that can be used for data for drivers, and their spatiotemporal resolutions would be a super helpful contribution here.
247- 250: What do you mean by top down vs bottom up approaches to detection? These are not defined or explained clearly here.
254-257: how do common approaches deal with clustering and serial correlation in the data, which no double effects standard errors, power, and inferences?
Section 282 – largely repeats with the intro – yet solutions as mentioned here are inadequately described “Solutions to potential spatio-temporal biases due
301 to data gaps include subsampling, weighting or imputation techniques (Bowler et al., 2024;
302 Nakagawa & Freckleton, 2011).” And “mass, so bias correction
308 methods need to be able to address this specific type of non-random bias (Johnson et al., 2021;
309 Sandel et al., 2015; Schrodt et al., 2015).” If you are going to say this paper is going into solutions, then do it.
316-317 – This is repeative with the previous section on data. Biodiversity data are inherently noisy due to genuine fluctuations in biodiversity as well as
317 sampling artifacts –
Paragraph at 402 – lots of jargon introduced and not defined and explained (exogeneity, collider, mediator, etc) until later. Maybe they could be defined as mentioned ?
714-718 – this feels like a rogue paragraph not connected to the narrative.
719 – no transition back to the data… again, 719-723 seems like a non-sequitur, but an important piece of guidance for the field.
Line 855 – this isn’t true. It’s not based on the best model in the predictive R2 sense.
Source
© 2024 the Reviewer.
References
Franziska, S., Miriam, B., Joaquim, E., E., B. D., Colin, F., Pierre, G., Romain, G., Matthias, G., S., M. I., Naia, M., Luca, S., Mickael, H., Gabrielle, M., Emmanuelle, P., Sara, S., Marianne, T., Grant, V., Cyrille, V., Wilfried, T. 2025. Advancing causal inference in ecology: Pathways for biodiversity change detection and attribution. Methods in Ecology and Evolution.
