Content of review 1, reviewed on June 27, 2023

This is a good, clearly set-out study that uses ZooMS to cross-check a previously-developed GMM identification protocol for sheep vs. goat teeth, demonstrating that it works on a fairly diverse set of ancient material roughly as well as it works on modern sheep and goats. This is a a very worthwhile test, that should improve confidence in the method among archaeologists and perhaps lead to increased uptake in preference to either molecular or qualitative morphological methods. Definitely worthy of publication and suitable for RSOS.

My main concern is that the paper does not engage at all with the question of variation with age and the impact (or lack thereof) of tooth wear on the GMM method: the word 'wear' does not appear in the manuscript, and 'age' only appears in the context of chronological periods. It is only noted that all reference M3s (and presumably also archaeological teeth?) were fully erupted. This seems a remarkable omission given that the primary reason zooarchaeologists will go to extra lengths to identify mandibles to species is to build age-at-death profiles based on occlusal wear. For one thing it's a missed opportunity to highlight the value of the study, so I would recommend stressing in the introduction the importance of mandibles and especially M3s for age-at-death analysis (that said, one could argue that the ability to identify mandibles with M3s but not those without risks biasing age studies against the youngest age categories... but that is not so much the present authors' problem!).

More importantly, you really need to tackle the worry that many readers will have that age (or more precisely wear stage, which is of course correlated with age) might result in shape change and hence affect identification reliability. I realise that you tackled this in your previous paper (Jeanjean et al. 2022, Sorting the flock, JAS), but this at least needs to be noted for the benefit of readers of the present paper who haven't also read that one. Moreover, the previous study showed limited age-related variation rather than none. It's notable that 1-2yr specimens (i.e. Payne stage D) seem to have been excluded from the the M3 analysis (or at least from the diagnostic plots for age) in the previous paper - which makes intuitive sense since these are by definition erupted but unworn and hence have a very different occlusal appearance - but a few are included in the reference dataset here (SI table 1). Were these actually used in the analysis? And perhaps more importantly were any of the archaeological specimens in this age category? What about older mandibles, in the 8-10yr (or 8+) range - i.e. Payne stage I - which showed some divergence from younger teeth in the first study?

Minimally I think you need to (a) recap your previous findings on age/wear by way of reassurance to a sceptical reader and (b) state in the main text the range of M3 wear states included in both the reference and archaeological datasets for the present study. Payne mandible stages are fine to this end, since they are based primarily on M3 where it's present. If at all possible, however, it would be good for transparency to include Payne stage information for all the archaeological specimens in SI table 2 (and to express the 'age' information in SI table 1 in terms of wear too, to match). I would personally have been interested in checking the wear states of the 5 mis-IDed specimens to see if there was any pattern (e.g. are they mostly unusually young or unusually old) - something that would be worth noting in the main text if so, or indeed if not - but that's simply not possible to do from the data supplied. Indeed, it would be great if SI Fig 1 included specimen identifiers in one way or another, so that a sceptical or just curious reader could link the confidence of the IDs to the archaeological information, again ideally including wear stage.

I was also a little concerned about the potential for lateral wear on the mesial end of the M3 impacting on landmarks 1 and 2, though of course this is rarer for M3 than for M1 or M2. Did this occur at all, in reference or archaeological specimens? If so, how was it dealt with? I assume the lack of semi-landmarks between these points is for exactly this reason? Not necessarily a problem, but I think worth considering in the text, even if only to state that it isn't a problem...

Minor comments, by line:

86-88: discrete criteria on the mandibles generally aren't very well trusted, so I think the advantage of in-situ teeth is more about having multiple teeth to work from, rather than the presence of the bone per se.

141: would be better grammar to put "belonged" after "specimens".

Figs 3 and 4: the colours used here are really not very clear. In particular the two "green" (I see this as yellow, but perhaps that's just me) colours are so close as to be almost indistinguishable. Please adjust to make them more distinct. The two blues aren't so bed.

227: would be clearer to put the "both" between "for" and "modern"

334: "milk purpose" doesn't really work as a phrase. "for the purpose of milk"?

343: "in more depth" would be better grammar here.

SI fig 1: could you put the black stars at whichever end of the boxplot a specimen was ultimately morphologically IDed as, i.e. at the top for sheep IDs and at the bottom for goat IDs? This would make it easier for a reader to see at a glance that mixed GMM results went in both directions but that there does seem to be some asymmetry, with 10 sheep having IDs straddling both species (including the 5 that were actually misIDed as goats) but only 2 goats being in this situation.
Thinking about it, this asymmetry could really do with a note in the main text too.

Source

    © 2023 the Reviewer.

References

    Marine, J., Krista, M., Silvia, V., Ariadna, N., Renate, S., Miquel, P. P., Sergio, J., Claude, G., Faiza, T., Rania, R., Cyprien, M., Allowen, E. 2023. ZooMS confirms geometric morphometrics species identification of ancient sheep and goat. Royal Society Open Science.