Content of review 1, reviewed on January 30, 2020
The authors sequenced three strains of Drechmeria coniospora using ONT, and assembled them with CANU. By examining the differences among new assemblies and previous reported assemblies at both structure and base levels, they found 1) long reads helped to assemble complex genomic regions, 2) chimeric reads led mis-assemblies, 3) homopolymer errors in final polished assemblies made the poor gene prediction. They also suggested to be careful to use the reference genomes assembled with pure ONT reads. I agree with most of those findings and suggestions. However, I think the main criticisms raised by the authors on ONT were already known to community. Also, the authors should continue to polish their assembly using NGS reads to explore the genome plasticity.
Obviously, without polishing using NGS reads, pure ONT assemblies cannot be used for gene prediction. But I am curious about the two very long but chimeric reads which respectively caused mis-assemblies in Swe1 and Swe3. There might be many other chimeric reads but didn't bring mis-assembly in CANU. Or if uses other long reads assemblers, will chimeric reads bring mis-assemblies on Drechmeria coniospora?
Declaration of competing interests Please complete a declaration of competing interests, considering the following questions: Have you in the past five years received reimbursements, fees, funding, or salary from an organisation that may in any way gain or lose financially from the publication of this manuscript, either now or in the future? Do you hold any stocks or shares in an organisation that may in any way gain or lose financially from the publication of this manuscript, either now or in the future? Do you hold or are you currently applying for any patents relating to the content of the manuscript? Have you received reimbursements, fees, funding, or salary from an organization that holds or has applied for patents relating to the content of the manuscript? Do you have any other financial competing interests? Do you have any non-financial competing interests in relation to this paper? If you can answer no to all of the above, write 'I declare that I have no competing interests' below. If your reply is yes to any, please give details below.
I declare that I have no competing interests.
I agree to the open peer review policy of the journal. I understand that my name will be included on my report to the authors and, if the manuscript is accepted for publication, my named report including any attachments I upload will be posted on the website along with the authors' responses. I agree for my report to be made available under an Open Access Creative Commons CC-BY license (http://creativecommons.org/licenses/by/4.0/). I understand that any comments which I do not wish to be included in my named report can be included as confidential comments to the editors, which will not be published. I agree to the open peer review policy of the journal.
Authors' response to reviews: We thank the reviewers for their constructive criticisms. In the light of their comments, we have undertaken a substantial amount of work, resulting in 4 new supplementary figures (Figs S4, S5, S6, S8), and revision of 2 others (Fig. 4; Fig S1), as well as extensive changes to the text in the light of their remarks.
The principal criticism was the absence of accurate short-read data to produce high-quality genomes. We have addressed this issue by generating these datasets and have made this data as well as the polished whole genome sequences publicly available. The sequences have allowed us to validate fully our conclusions regarding the current state of ONT-only genome assemblies (see revised Figure 4, and new Supplementary Table 1). The polishing we have done has allowed us to address with more confidence the question of plasticity at the level of the DNA sequence, and we have included a general comparison of the polished genome sequences in the revised manuscript. A deeper exploration of the differences between the strains would involve generating a curated gene prediction for each genome and undertaking a detailed cataloguing of SNPs and indels. This is a project in itself, and although we are now embarking on this, we believe that it is outside the scope of the current manuscript. With regards the overall comparisons, we now include a BUSCO analysis of the different assembly steps in the main text (new Table 1). As we point out, this kind of analysis should be treated with caution. In contrast to our careful examination of 305 mono-exonic genes, the BUSCO scores of the assemblies after long-read or short -read polishing were very similar since frame-shifts due to homopolymer length errors are not necessary picked up by BUSCO. We note in passing that the proportion of correct genes in the assemblies before polishing (Figure 4) is in reality even lower than originally presented. Due to an error in the data parsing script, we had used assemblies from a preliminary polishing test instead of assemblies without polishing. The results show that current polishing tools do a very good, but far from perfect job of error correction when using just nanopore reads. As one of the reviewers pointed out, ONT-only genomes merit being flagged in sequence databases.
Otherwise, Reviewer 1 was curious about the two long chimeric reads that caused mis-assemblies in Swe1 and Swe3, and more generally chimeric reads. We now show that chimeric reads are more likely to be trimmed by Canu, and to a greater extent, compared to non-chimeric reads (new Supplementary Figure 4). As requested by both reviewers, we also tried a different assembler, Flye. Although it is less sensitive to the presence of chimeric reads (new Supplementary Figure 6A), we detected other anomalies with this tool (new Supplementary Figure 6B-D). Our observations support the reviewers’ suggestion that using more than one assembler can help determine correct genome structure.
With regards Reviewer 2’s comments:
- For the issue of short read polishing, see above. The different references have been included.
- We now address the question of the quality of PacBio-only genomes in the main text.
- We have expanded the section regarding the DNA extraction protocol and added information about read coverage (new Supplementary Table 1).
- Concerning the loss of sequence accuracy after polishing, we addressed the question of whether the affected regions were repeat regions. Of the two examples of inconsistencies introduced by long-read polishing (Figure 5), the first was a region close to a nuclear insert of a fragment of the mitochondrial genome (also referred to as numts sequences), which in part is duplicated in tandem (new Supplementary Figure 8A-B). We conclude that the combination of this duplication plus the fact that reads from mitochondrial DNA can align in the region, led polishing tools to introduce erroneous sequence. The second case was about a chunk of ca. 10 kb omitted by Canu in the assembly of Swe3, but strongly supported by the reads. We had shown that polishing tools introduced a sequence with the correct length but with an erroneous composition. As presented in the new Supplementary Figure 8 C-D, there is indeed a small repeated sequence in the neighbourhood of the 10 kb chunk.
- As requested we undertook a close examination of the rearrangement break points, but failed to detect any notable signature of transposable elements or repeated sequences in the neighbourhood (revised Supplementary Methods).
- See comments about Flye, above
Source
© 2020 the Reviewer (CC BY 4.0).
Content of review 2, reviewed on April 27, 2020
Overall, the authors have improved the manuscript by performing genome polish using NGS data and compared the assemblies generated by different assemblers. However, the new manuscript still highlights the unsatisfying accuracy of ONT-only assembly, and the improvement of genome quality after polishing using ONT and NGS reads, which has been shown in many different works (Schmidt et. al.,2017, the Plant Cell; Belser et.al., 2018, Nature Plant; Michael et.al., 2018,Nature communications ). In the light of previous works on ONT assembly, the novelty of this finding should be toned down.
One very interesting but often overlooked problem in assembly using long read is the mis-assembly derived from chimeric reads. It would be very novel and informative to deeply explore the assembly of the chimeric reads observed in this study, that is why I suggested the experiment using other assemblers to deeply explore this issue. The authors did provide Supplementary Figure 4 and 5 to show the basic stats. However, I don't think "The use of other genome assembly tools, and the comparison of assembly discrepancies is an additional method to produce high confidence genomes" in L204-L206 is enough to produce accurate genome structure. Because users will be confused when assemblers give inconsistent structures, they are expecting how to do with conflict in the absence of reference genome. Besides, the authors reported the assembly generated using flye but ignored the mis-assembly caused by chimeric reads in it.
Declaration of competing interests Please complete a declaration of competing interests, considering the following questions: Have you in the past five years received reimbursements, fees, funding, or salary from an organisation that may in any way gain or lose financially from the publication of this manuscript, either now or in the future? Do you hold any stocks or shares in an organisation that may in any way gain or lose financially from the publication of this manuscript, either now or in the future? Do you hold or are you currently applying for any patents relating to the content of the manuscript? Have you received reimbursements, fees, funding, or salary from an organization that holds or has applied for patents relating to the content of the manuscript? Do you have any other financial competing interests? Do you have any non-financial competing interests in relation to this paper? If you can answer no to all of the above, write 'I declare that I have no competing interests' below. If your reply is yes to any, please give details below.
I declare that I have no competing interests.
I agree to the open peer review policy of the journal. I understand that my name will be included on my report to the authors and, if the manuscript is accepted for publication, my named report including any attachments I upload will be posted on the website along with the authors' responses. I agree for my report to be made available under an Open Access Creative Commons CC-BY license (http://creativecommons.org/licenses/by/4.0/). I understand that any comments which I do not wish to be included in my named report can be included as confidential comments to the editors, which will not be published. I agree to the open peer review policy of the journal.
Authors' response to reviews:(https://drive.google.com/file/d/11TghMmTEJ-uEbgkokf5NOSZc1t2tz8bj/view?usp=sharing)
Source
© 2020 the Reviewer (CC BY 4.0).
References
Damien, C., Jan, P., Jerome, R., Guillaume, B., Vladimir, B., J., E. J. Long-read only assembly of Drechmeria coniospora genomes reveals widespread chromosome plasticity and illustrates the limitations of current nanopore methods. GigaScience.
