Content of review 1, reviewed on January 26, 2022

This manuscript addresses a key limitation and problematic in the current state of art of machine learning classification methods in Ecology, namely the practice of using one unique classifier trained on labelled training data and applying it to unlabelled dataset with different and diverse types of habitats. This work aims to resolve a problematic of utmost relevance in the field of population and community monitoring using bioacoustic methods. It proposes an original pipeline and framework with important checkpoints to reduce classification errors and improve both classifiers and ultimately population estimates. I expect this work to be highly appreciated by the bioacoustic and ecoacoustic community.

I congratulate the authors for a fine, and well-written manuscript which offers a simple and elegant method to address an important problem in bioacoustic studies. The results, interpretation of the results and the discussion were clear and accurately interpreted. I have suggested mostly only minor suggestions. This research will certainly help the growing scientific community that uses machine learning classifiers to reduce heterogeneity of errors.

Abstract
Well done on a concise abstract.

Minor suggestions:

l.19 avoiding influencing animal behaviour
I suggest you replace this sentence with "reducing anthropogenic monitoring disturbances"

l.21 " developed or implemented."
developed or implemented in Ecology.

l.29 "Contextual data included the predicted presence of the target species at similar times, the predicted presence of other species, or acoustically derived environmental data."
I suggest you remove this sentence, which I think would best fit into your methods.

Introduction

This is a well written introduction. It highlights the limitations and issues of current automated classification methods and challenges of using field datasets. I have made one major suggestion and a few minor suggestions. Please consider the major suggestion but feel free to ignore it if you do not think it is necessary.

Major suggestions:

There is an important overlap between paragraphs l.68 and l.90. I understand that you have tried to separate paragraphs based on "issues with datasets" versus "heterogenous error", but because both are tightly intertwined and interdependent, if possible, consider merging the paragraphs at l.68 and l.90 to reduce repetition.

Minor suggestions:

l.54 "However, outside of chiropterology very few studies have used automated classification to answer applied ecological questions and especially for multi-species
classification over large audio datasets in tropical forest ecosystems."

This is a general statement, some scientists would disagree with it. Automated classification has been extensively used by scientists from the ornithology and cetology community for many years and has resulted in rich literature. I think what you mean to emphasize here is that automated classification has not been applied to tropical forest ecosystems.

l.57. This is a very clear paragraph with great choice of literature citations that contributed to impressive advancements in the field of automated classification using machine learning.

l.68 This paragraph lists some of the many issues scientists can encounter when trying to create classifiers. Since one of the main question of your research is "Can an algorithm trained on data from one place be used elsewhere", I wonder if you could address a little more the challenges that heterogeneous habitats can generate for automated classification of species. For example, when one species exists in various types of soundscapes and habitats, many characteristics of these habitats matter and make accurate classification harder (e.g. vegetated dense forest versus open field; overlapping with heterospecific species in one habitat but not the other.)

l.75. "Additionally, The use"

A small typo (capitalized letter) here need fixing.

l.89. "or the error varies spatially in a space-for-time swap experimental design"
I recommend to add citations here

l.101. This is a great idea to have incorporated this useful table here. It highlights some of the main terminology and definition used in event detection.

l.109. This is true and a limitation of many studies, and I am glad you are offering a solution to this with your work

l.128. Excellent proposition.

Materials and Methods

l.136. (latitude ~ -3.046, longitude -54.947 WGS 84)
This is a strange way of reporting GPS points, please follow international conventions to report the location with proper units (e.g. degrees, then minutes, then seconds)

l.149. Sentence is a little strange here, need some fixing

l.154 same as l.136

l.156 Automated and manual classification

l.161. I would recommend moving this information earlier at l.157, because, the first sentence reads strangely as Tadarida is also the Genus name of some bats. Further, it is a relatively new package/toolbox, so I think it would be best to introduce it right away and describe what it does, etc.

l.164. Could you clarify what this hysteresis function works?

l.171. "manual labelling of detected sound events" by Tadarida
When reading the flow of the methodology, I was confused by the order of the methodological steps. It was unclear to me whether you first used Tadarida to make a rought auto-detection using frequency threshold (0.2 kHz and 4.2 kHz) and then manually labelled the detected sounds. I think this is the case, but another possibility is that you could have started right away from manual labelling, so I think it is important to clarifying this. One suggestion would be to replace "first" at l.171 by "next" and add "by Tadarida" at the end of the sentence. Also, make it clear that "Training Dataset 1" is the manually labelled dataset and not the initial "frequency threshold dataset". I suggest you add another sentence to clarify this instead of (Training Dataset 1).

l.174 (OCM)
Write the full name of acronyms the first time you present it. I now realize these are collaborators' name initials. Consider removing them. But if you insist in keeping them, make sure you clarify e.g. labeler OCM. Also, either use the initials of collaborators throughout the methods or leave them all out for consistency.

l.190. "300 sound types were identified" for the seven nocturnal species
Please could you clarify how you defined sound types. And why did you simplified then to 59 sound types. The reasoning here is missing.

l.195. A small typo with the ()

l.196 "15 second sound files"
Consider changing this throughout

l.199. Tadarida see SOM (citation)
Also, provide full name of SOM

l.200 This pipeline summarizes well all the steps readers should consider taking to reduce heterogeneity of error. Great job!

l.204. "we followed"

l.238 Knight et al. {missing date}

l.272 Acoustically derived environmental data
Maybe here is the place to add "Contextual data included the predicted
presence of the target species at similar times, the predicted presence of other species, or
acoustically derived environmental data."

Results
The results are clear, accurately reported and easy to follow. I had only minor suggestions.

l.393 X axis
Please write CV in full

l.411 Legend
Please "threshold" underneath Contextual versus Tadarida-Cv versus Tadarida-No Cv

l.435
Given most of your legend was placed below your figure, moving your legend underneath will be consistent with other figures

Discussion

Major comments:

The discussion was well written and every aspects of the results were discussed and addressed. However, the discussion was a little repetitive with the results. I think it is fine the way it is now, but consider presenting more examples of literature that relied on a single classifier and that reported biases in their population estimates. This will reemphasize the importance of your work. The discussion is a place where you can make your novel method really shine to encourage readers to follow its framework and avoid these biases.

Minor comments:

l.443. "with a random forest based Tadarida"
with a random forest based on a Tadarida

l.447. "and not to rely solely on"
and that do not to rely solely on

l. 478 "The second method to reduce precision error heterogeneity is a secondary contextual classifier."
The second method used to reduce precision error heterogeneity was a secondary contextual classifier.

l. 479. "In contrast, this requires considerably"
In contrast to other methods, this required

l.482 space before bracket

l.512. This was one of my thought earlier, I'm glad you are addressing this in your discussion

Source

    © 2022 the Reviewer.

Content of review 2, reviewed on May 09, 2022

Dear authors,

Thank you for the elaborate comments and edits that helped clarify areas of your work that needed some improvement. I am pleased with the edits and clarifications the authors have made. I have no further comments on your work. Great work!

Source

    © 2022 the Reviewer.