Content of review 1, reviewed on October 06, 2020
Re-review of “Pre-introduction introgression contributes to parallel differentiation and contrasting hybridisation outcomes between invasive and native marine mussels”
I appreciate the authors diligence at attempting to address reviewers concerns. I think this is a good resource paper that is worth being published, but I don’t think the authors fully addressed my concerns about lncRNA population structure or introgression amounts and FST. I’ve detailed why below.
-lncRNA snp quality:
I acknowledge that my suggestion of retaining the most possible alignment locations was based on experience working with genomic references but I think there is a difference between removing highly similar sequences because you aren’t sure if they’re isoforms vs removing most of the possible alignment targets. I would be concerned that so few of the lncRNA SNPs are found when aligning to the total transcriptome set. What happened to the rest of them? With regard to the population structure in 25 lncRNA SNPs, it’s not fair to compare that PCA with the PCA of ~17K SNPs. With too few bi-allelic sites you wouldn’t see population structure, regardless of whether you could with a larger number of sites. I wouldn’t require this, but it would be more convincing if you compared the same number of SNPs, say permuted sets of the 17K, and calculated some measure of cluster separation. With that you could say where the lncRNA SNPs fall in a spectrum of regular SNPs in terms of population structure.
All that being said, I appreciate the greater talk about the possibility of genotyping error.
-MaxFST
I’m happy with the authors response to my concerns.
-Figure 4 and allelic deviation
I don’t fully understand the relationship between D and FST and the verbal arguments from the authors have not convinced me. I agree with the authors take that it wouldn’t work to focus on outliers, but I am still concerned that the slope they’re reporting and discussing is a feature of the stat, rather than the system. If you limit to loci where the admixed allele frequency is between the parental allele frequency, I don’t see how there couldn’t be smaller absolute values of D at lower FST values because at lower FST values parental allele frequencies are closer together. It’s also complicated because D itself relies on admixture proportion which relies on relative allele frequencies, so how that all plays out is unclear. I would recommend some sort of simulation. Something where you start at the parental allele frequencies, make an admixed population with a bottleneck and then see what the slope looks like for this stat to see how often the slope is non-zero. It’d be even better if there was an hybrid zone model with selection on some loci, but that seems too much to ask for.
Maybe there is a paper they can cite that has done something like? I will cede to the editor’s decision on this, but currently I’m not convinced by this analysis. I also noticed that the authors cite Simon et al. 2019 for D, but that exact stat isn’t used in Simon et al (at least not explicitly). They use bgc which is conceptually quite similar. I’m not sure why the authors didn’t use bgc, which fits each locus to a hybrid zone model and has a more robust background for how the stat behaves.
Source
© 2020 the Reviewer.
