Content of review 1, reviewed on March 13, 2023
The paper describes an automatic photo-identification method applicable to different species. The topic of the paper is very interesting, but I have several concerns on this paper and on the method the authors propose. It seems that the proposed method collects pieces of algorithms developed by other researchers, with no clear justification for employing them (i.e. see my comments on pseudo-labelling). The novelty of the paper is not clear.
The authors present a seemingly working system, but without the ability to extract general rules and consideration on its performance. In particular, they even assert that the number of training images does not influence the performance of the proposed method, without finding a reasonable reason.
In the submitted paper, the authors ignore the state of the art of modern literature on the automated cetacean photo-identification. They only cite finFindR by Thompson et al. 2021 and few other methods with a superficial discussion. The authors should consider the following papers:
R. Maglietta, R. Carlucci, C. Fanizza and G. Dimauro, "Machine Learning and Image Processing Methods for Cetacean Photo Identification: A Systematic Review," in IEEE Access, vol. 10, pp. 80195-80207, 2022
Maglietta, R., Renò, V., Cipriano, G. et al. DolFin: an innovative digital platform for studying Risso’s dolphins in the Northern Ionian Sea (North-eastern Central Mediterranean). Sci Rep 8, 17185 (2018)
R. Maglietta et al., "Convolutional Neural Networks for Risso’s Dolphins Identification," in IEEE Access, vol. 8, pp. 80195-80206, 2020
R. Maglietta et al.,“ARIANNA: A novel deep learning-based system for fin contours analysis in individual recognition of dolphins”, Intelligent Systems with Applications, 18, 2023
Which are the benefits of their method compared with the state of the art of the automated photo-id algorithm?
Moreover, they wrote: “we introduce a multi–species approach to photo–id“, but I have not understand why the proposed approach is multi-specie approach. The dataset is surely multi-species dataset. In my opinion, this dataset has a great value for the scientific community providing a valuable information on those species. I have not understood where the proposed model allows species to share information, and how they can conclude that it outperforms a single species model if there is no comparison study with other methods. They wrote (rows 67-69) this, and they highlighted that their model works well for species with few training images, but no comparison with other methods was evaluated. The proposed approach provides individuals recognitions separately for each specie (see fig. 7). I suggest the authors to clarify this point.
The added value is the multi-specie dataset provided by the competition; I am not sure that it can be considered a novelty of the paper. Please clarify.
Other major revisions follow.
Abstract
Row 5: “However, the models that have facilitated this change, i.e., convolutional neural networks, need many training images to generalize well. As a result, they have often been developed for individual species that meet this threshold. These single-species methods may underperform, as they ignore potential similarities in photo- identification among species.”
There are some cetacean photo-id algorithms, already presented in the literature, that use statistical methods different from deep learning (see Maglietta et al. Scientific Reports 2018), which do not need many images to perform well. I think that the sentence “These single-species methods may underperform” is too strong, considering that no comparison among different approaches has been performed.
Please clarify what you mean with the sentence: “These single- species methods may underperform, as they ignore potential similarities in photo-identification among species.”
Introduction
Row 52-55: Firstly, overfitting is a problem of machine learning in general, and not only of deep learning. The sentence “That is, they will struggle to identify individuals in images outside of the training set unless they are trained with many diverse images” is misleading. Surely, deep learning requires huge amount of data to be trained, but overfitting can arise even if large training set are available. In fact, a robust methodology must be implemented to avoid overfitting, and several methods are available to mitigate it, that is the focus of the problem. Please, clarify it.
Row 56: Could it be that researchers have focused on developing automated systems for individual species because they do not have enough data from more than one species? Collecting sighting data is not an easy job, as well known.
Row 68: “Thus, a multi–species model that allows species to share information might outperform a single species model, particularly for species with few training images.”
There is no evidence of this matter, this is an opinion of the authors that must be proved.
Row 74-79: Concepts already said at the beginning of the introduction are repeated here (i.e. labor cost). Please, avoid these repetitions.
Row 85: “The dataset was assembled for a data science competition that challenged teams to recognize cetaceans from images of their dorsal/lateral side“. I understand that the effort of collecting data has not been done by authors, but it has been done within a competition. If it is correct, please clarify it. Acquiring and collecting data require a great effort and many costs. The reader should know if the dataset is an original contribution of the paper.
Row 88-90: I have appreciated the validation on the two additional catalogues. It is very useful to verify the generalization capability of the proposed model.
Materials and Methods
Rows 97-100: Authors should comment that the problem of recognizing new individuals during automated photo-id of cetaceans has been already discussed and solved in R. Maglietta et al., "Convolutional Neural Networks for Risso’s Dolphins Identification," in IEEE Access, vol. 8, pp. 80195-80206, 2020, doi: 10.1109/ACCESS.2020.2990427.
Row 105: “other applications would likely look similar”: which application? Please, insert references.
Rows 105-107: “The model’s backbone and neck are commonplace for image classification relative to the classification heads. Thus, we discuss the latter in greater detail.” Please insert references to demonstrate the commonplace of model’s backbone and neck.
Row 113: Khan et al. 2020 is not listed in References section.
Row 113-115: The choice of the type of neural network seems to be arbitrary here. The number of parameters is surely a relevant point. But a full evaluation of classification performances and computational costs of different neural networks should be considered by authors, in particular considering the “flurry of research” on this subject, as correctly highlighted by them.
Rows 117-120: Please, clarify this sentence.
Rows: what is i?? please specify it. Which is the range of value of xi?
Rows 164-202 refers to the published papers Deng et al 2019 and 2020. The description is quite confusing, and contains some useless details, such as the primordial definition of scalar product (rows 166-168). I suggest shortening the description, opportunely recalling the Deng’s papers. If a reader is interested in the Arcface methods can read the papers, where also the code has been made available.
Rows 206-217: As already said before, these rows refer to Ha et al. 2020 paper, where authors discuss dynamic margins for ArcFace loss. In general, in my opinion more than four pages to discuss already published results without any novelty are not useful.
Moreover, which is exactly the novelty of the presented method and, more in general, which is the novelty of the paper?
Row 246: I suppose that authors refer to ‘supplement-catalog-details.pdf’ file. Please, insert a row number in the table. Are the catalogues 40? Are the species 25? Please insert a detailed caption for the table in this file. I suggest describing better, with much more details, the meaning of each item in this table. Where are the 51033 training images? Please specify it, adding a row at the end of the table. Which is the number of individuals in each catalogue?
Rows 248-249: I do not understand where class imbalance is shown in figure. There is an unbalancing among the number of training and test examples. Moreover, the number of species listed in fig, 7 is 24. Why isn’t it 25??
Row 250-251: Authors wrote: “The 51,033 training images contained 15,587 identities. Of these, 9,258 (59%) had only one training image, while 14,237 (91%) had five or fewer.” It is misleading: 59% of what? I understand that it is the 59% of 15587. Similarly, 91% of what? if I have undestrand well, 91% of images had a number of images in the range from 5 to 1.
Rows 254-258: This description is very superficial. “several competitors”: who are the competitors?? Which models have been trained? Finally, we have now some descriptions about the preprocessing illustrated in figure 2. Which are the differences between the two models based on YOLOv5 and which is the methodology implemented?? A description of the algorithm devoted to the fin cropping should be added. I suggest inserting a dedicated section.
Rows 267-272: Which is the dataset size after data augmentation?
Table 1: please insert in the table caption a description of the arguments.
Rows 289-290: Which is the need of introducing pseudo labeling? You have already done data augmentation; can you prove that pseudo labelling is necessary?? If pseudo labelling is not used what happens?
Rows 326-327: “The precision varied among species (Fig. 7), and did not correlate with the number of training images or test images. “ This is very anomalous. How can you explain it??
Row 346-349: I am very surprised by this sentence! I suggest authors to illustrate this “no consistent relationship“. How have they tested this?? Just as example: Figure 8 seems to be their demonstration of this sentence. Now, I have a question. In fig. 8 the plot of train images per ID vs MAP is illustrated. But the authors wrote that “(91%) of identities had from 5 or fewer training images. So, in my opinion you have no identities with a higher number of images for training your model!!! That’s the main point. Are you sure that you can evaluate the relationship between MAPs and the numbers of training images per identity?? I am not so sure. Please clarify this point. Moreover, are you sure that you can train a machine learning based-algorithm using only one train image per ID?? How does Data augmentation impact on it?
Fig 7: I suggest illustrating error bars of MAP for each species.
Source
© 2023 the Reviewer.
References
T., P. P., Ted, C., Kenshin, A., Taiki, Y., Walter, R., Ken, S., Addison, H., M., O. E., B., A. J., Erin, A., Aline, A., W., B. R., Charla, B., Elsa, C., John, C., Julio, C., L., C. E., Amina, C., J., C. B., Enrico, C., Jens, C., W., D. J., A., F. E., Holly, F., Kiirsten, F., Trish, F., Wally, F., Barbara, G. V., Tilen, G., Marie, H., R., J. D., L., K. E., D., M. S., L., M. T., Liah, M., Catherine, M., Robert, M., Anastasia, M., N., O. D., C., P. H., H., R. M., J., R. W., Caroline, R., Renato, R., Salvatore, S., Stephanie, S., Beatriz, T., G., T. L., R., T. J., Cameron, T., Reny, T. M., R., W. C., Rebecca, W., Randall, W., M., Y. K., R., Z. J., Lars, B. 2023. A deep learning approach to photo-identification demonstrates high performance on two dozen cetacean species. Methods in Ecology and Evolution.
