Content of review 1, reviewed on July 08, 2024
The authors present an interesting new approach on how to aggregate labels for training ML models for the identification of plants based on crowdsourced data that can be noisy due to the open-ended sampling strategy that leads to those labels. Aggregating these labels operates in a space that has trade-offs between the labeling accuracy and the amount of data retained (i.e. very accurate and stringent label aggregation might lead to very precise identifications at a large loss of observations and vice versa).
The authors provide a detailed description of their approach (including open source code) and do a rigorous job in comparing their strategy to existing alternatives. Overall, I think there are just some minor improvements that I would suggest to make this manuscript a bit more accessible:
I found the methodology section somewhat hard to follow at times, mainly due to the large amount of detailed algorithmic descriptions and notations, which made me flip back and forth between the actual descriptions and definitions. Due to a large part that is probably on me not being as familiar with the field, but I think other readers might also benefit from a small toy example of how the score calculation is done (along the lines of "assume 3 users, doing X observations, then…). I appreciate though that this might be outside potential word limits, in that case I would encourage thinking about whether it could fit into an appendix/supplement.
The text surrounding Figure 5 talks about D(expert), D(multiple votes) and D(disagreement), but the figure itself seems to only present the latter two, if space allows it would be nice to have D(expert) as a panel, to allow more easily comparing these (unless I misunderstood and D(expert) is the Fig 5A, but the percentages don't seem to fit that).
In its current form the qualitative exploration around Figure 7 is very brief. As a reader I wonder how representative those examples are, how frequent those causes of confusion are etc. In its current state it feels like this might be a better fit for the discussion, alternatively one could expand a bit on these sources of invalid observations.
Lastly, this is more a question of my own curiosity. Metadata, such as location etc. are being suggested as potentially being of further use for the aggregation. One other question that came to my mind was: Would it be useful to also take into account the taxonomic distance? I could see some contributors being specialists in a given taxonomic field, which could show up as mainly providing annotations from a small subset of the plant tree of life. Annotations made within that field might have a higher trust than one-off contributions for taxa in which participants are less experienced as seen in the data. Did you ever experiment with this, or think about that? :-)
Source
© 2024 the Reviewer.
Content of review 2, reviewed on September 18, 2024
I thank the authors for their comprehensive responses to the queries raised in the review - very kindly even to the ones that were mostly about my own curiosity! Overall, I think this manuscript is now in a good state for publication.
The authors did a great job at improving the few points I raised and I think that the accessibility for people less familiar with the field and clarity of the manuscript has now improved quite a bit.
I also much appreciated the detailed sharing of the code resources.
Source
© 2024 the Reviewer.
References
Tanguy, L., Antoine, A., Benjamin, C., Jean-Christophe, L., Mathias, C., Herve, G., Joseph, S., Pierre, B., Alexis, J. 2026. Cooperative learning of Pl@ntNet's Artificial Intelligence algorithm: How does it work and how can we improve it?. Methods in Ecology and Evolution.
