Content of review 1, reviewed on February 09, 2020

Reviewer comments for GIGA-D-19-00433

The manuscript by Courtine and co-workers entitled "Long-read only assembly of Drechmeria coniospora genomes reveals widespread chromosome plasticity and illustrates the limitations of current nanopore methods" addresses an important research topic of broad interest to the readership of GigaScience and beyond.
Large advances in assembly algorithms and NGS technologies, in particular third generation long read sequencing technologies from Pacific Biosciences and Oxford Nanopore Technologies, have led to the fact that the number of genome assemblies reported is growing exponentially (including thousands of prokaryotic genomes). Given that many publications have advocated the use of a combination of long read sequences (known to have issues with homopolymer stretches) with highly accurate short read sequences, I was surprised that this was not done in this study. Nevertheless, one relevant take home message from the study (Discussion page 8) is that a certain number of genome assemblies that only rely on ONT data are being uploaded in public data repositories; ideally, such assemblies are flagged at the NCBI or users are at least made aware of the implications.

Comments 1. The major criticism with the paper and its take home message as presented is the question why the authors did not include higher quality short read data right from the beginning? For a superb, state of the art core facility like the EMBL, this would seem straight-forward. For studies on complex, repeat-rich prokaryotes, the advantages of very long ONT reads to help to correctly place long, nearly-identical repeat sequences of 40 kb and above has been reported (Schmid et al, Nucleic Acids Research 2018; PMID: 30137508); this study and reports on eukaryotes e.g. from the Loman group (Jain et al Nat Biotechnology 2018¸ PMID: 29431738) have used short read (or PacBio) data to achieve very high quality genome assemblies of both prokaryotes and eukaryotes.

  1. In my view, the message that 3rd generation long read sequencing from the ONT platform when used alone is not up to par to assemble complex eukaryotic genomes needs to be put a bit more into perspective. On page 8 (2nd paragraph) it is mentioned that ONT only assemblies are being added to public sequence repositories (52 eukaryotes, 99 bacteria). I agree that these should be flagged. However, other studies should be cited that have reported eukaryotic genome assemblies and that have relied on additional higher quality short read (e.g. the key study from Jain et al, Nature Biotechnology 2018 (PMID: 29431738)). Would the authors conclude that assemblies based on PacBio data alone also do not merit inclusion into databases? What are the factors that undermine the utility of long read ONT data when used alone the most, the existence of long homopolymer stretches, or the existence of long near-identical repeats that need to be treated carefully in order not to lower the assembly quality (see also point 4).

  2. The ONT platform is very sensitive to the quality of the genomic DNA extracted. Therefore, the method section needs to specify in more detail which DNA extraction protocol has been used (currently, the reader is referred to REF #14). Similarly, I was missing in the result section the coverage of the genome assembly that was achieved by long ONT reads.

  3. One key point is the fact that polishing can reduce the accuracy in certain regions in the assembly (see top paragraph on page 6). Are the regions where this was observed repeat regions? This aspect has been noted in a study on long-read based genome assemblies from metagenome samples (Somerville et al BMC Microbiology 2019 (PMID:31238873): "Moreover, great care needs to be taken when contigs are polished individually, since this can lead to the erroneous removal of true, natural sequence diversity due to cross mapping of reads in repeat regions (e.g., repeated sequences such as 16S rRNA operons, insertion sequences/transposases)". It would be valuable to detail this result section a bit further and explore/mention this for the examples shown in Figure 5.

  4. One analysis that would have perfectly fit for the message of this paper was the analysis of the borders of the regions of the sites of reorganization between the Dan1 and Dan2 genomes (see page 8 1st paragraph; this is mentioned as an outlook). Can't these be run through a repeat analysis or against families of common repeat elements?

  5. Canu was the only assembly algorithm explored. I wonder whether other assemblers that work well with error-prone reads and near identical repeat sequences (e.g. Kolmogorv et al, Nat Biotechnol 2019; PMID:30936562) would be helpful in this respect.

Minor Figure S1 -> the axis is not labeled (kb) Page 5; remove empty spaces after citation [9]. Methods (page 10); adapt… then adaptor were then trimmed -> then adaptors were trimmed

Declaration of competing interests Please complete a declaration of competing interests, considering the following questions: Have you in the past five years received reimbursements, fees, funding, or salary from an organisation that may in any way gain or lose financially from the publication of this manuscript, either now or in the future? Do you hold any stocks or shares in an organisation that may in any way gain or lose financially from the publication of this manuscript, either now or in the future? Do you hold or are you currently applying for any patents relating to the content of the manuscript? Have you received reimbursements, fees, funding, or salary from an organization that holds or has applied for patents relating to the content of the manuscript? Do you have any other financial competing interests? Do you have any non-financial competing interests in relation to this paper? If you can answer no to all of the above, write 'I declare that I have no competing interests' below. If your reply is yes to any, please give details below.

I declare that I have no competing interests.

I agree to the open peer review policy of the journal. I understand that my name will be included on my report to the authors and, if the manuscript is accepted for publication, my named report including any attachments I upload will be posted on the website along with the authors' responses. I agree for my report to be made available under an Open Access Creative Commons CC-BY license (http://creativecommons.org/licenses/by/4.0/). I understand that any comments which I do not wish to be included in my named report can be included as confidential comments to the editors, which will not be published. I agree to the open peer review policy of the journal.

Authors' response to reviews: We thank the reviewers for their constructive criticisms. In the light of their comments, we have undertaken a substantial amount of work, resulting in 4 new supplementary figures (Figs S4, S5, S6, S8), and revision of 2 others (Fig. 4; Fig S1), as well as extensive changes to the text in the light of their remarks.

The principal criticism was the absence of accurate short-read data to produce high-quality genomes. We have addressed this issue by generating these datasets and have made this data as well as the polished whole genome sequences publicly available. The sequences have allowed us to validate fully our conclusions regarding the current state of ONT-only genome assemblies (see revised Figure 4, and new Supplementary Table 1). The polishing we have done has allowed us to address with more confidence the question of plasticity at the level of the DNA sequence, and we have included a general comparison of the polished genome sequences in the revised manuscript. A deeper exploration of the differences between the strains would involve generating a curated gene prediction for each genome and undertaking a detailed cataloguing of SNPs and indels. This is a project in itself, and although we are now embarking on this, we believe that it is outside the scope of the current manuscript. With regards the overall comparisons, we now include a BUSCO analysis of the different assembly steps in the main text (new Table 1). As we point out, this kind of analysis should be treated with caution. In contrast to our careful examination of 305 mono-exonic genes, the BUSCO scores of the assemblies after long-read or short -read polishing were very similar since frame-shifts due to homopolymer length errors are not necessary picked up by BUSCO. We note in passing that the proportion of correct genes in the assemblies before polishing (Figure 4) is in reality even lower than originally presented. Due to an error in the data parsing script, we had used assemblies from a preliminary polishing test instead of assemblies without polishing. The results show that current polishing tools do a very good, but far from perfect job of error correction when using just nanopore reads. As one of the reviewers pointed out, ONT-only genomes merit being flagged in sequence databases.

Otherwise, Reviewer 1 was curious about the two long chimeric reads that caused mis-assemblies in Swe1 and Swe3, and more generally chimeric reads. We now show that chimeric reads are more likely to be trimmed by Canu, and to a greater extent, compared to non-chimeric reads (new Supplementary Figure 4). As requested by both reviewers, we also tried a different assembler, Flye. Although it is less sensitive to the presence of chimeric reads (new Supplementary Figure 6A), we detected other anomalies with this tool (new Supplementary Figure 6B-D). Our observations support the reviewers’ suggestion that using more than one assembler can help determine correct genome structure.

With regards Reviewer 2’s comments:

  1. For the issue of short read polishing, see above. The different references have been included.
  2. We now address the question of the quality of PacBio-only genomes in the main text.
  3. We have expanded the section regarding the DNA extraction protocol and added information about read coverage (new Supplementary Table 1).
  4. Concerning the loss of sequence accuracy after polishing, we addressed the question of whether the affected regions were repeat regions. Of the two examples of inconsistencies introduced by long-read polishing (Figure 5), the first was a region close to a nuclear insert of a fragment of the mitochondrial genome (also referred to as numts sequences), which in part is duplicated in tandem (new Supplementary Figure 8A-B). We conclude that the combination of this duplication plus the fact that reads from mitochondrial DNA can align in the region, led polishing tools to introduce erroneous sequence. The second case was about a chunk of ca. 10 kb omitted by Canu in the assembly of Swe3, but strongly supported by the reads. We had shown that polishing tools introduced a sequence with the correct length but with an erroneous composition. As presented in the new Supplementary Figure 8 C-D, there is indeed a small repeated sequence in the neighbourhood of the 10 kb chunk.
  5. As requested we undertook a close examination of the rearrangement break points, but failed to detect any notable signature of transposable elements or repeated sequences in the neighbourhood (revised Supplementary Methods).
  6. See comments about Flye, above

Source

    © 2020 the Reviewer (CC BY 4.0).

Content of review 2, reviewed on May 08, 2020

Reviewer comments for GIGA-D-19-00433R1 The authors have done a very good job at addressing the various reviewer comments in a revised manuscript. More specifically, they have added a substantial amount of new analyses that now more thoroughly support the claims made (e.g., showing in the new Table 1 how different polishing steps affect the quality of a set of USCOs, including information on fragmented and missing genes and mentioning that these olishing steps have to be treated with caution in repeat regions; exploring other genome assemblers (here Flye) ans using some manual curation efforts to create a final assembly of higher quality; further exploring mis-assembled regions for the presence of certain classes of repeats and clarifying the origin of a long chimeric read that lead to a certain assembly artefact) and that are in line with observations made by other groups in the field. Indeed, one of the key messages is that caution has to be exercised when dealing with ONT long-red only assemblies (i.e. without any polishing using more accurate short read data) in genome sequence repositories. I support publication of the revised manuscript.

Declaration of competing interests Please complete a declaration of competing interests, considering the following questions: Have you in the past five years received reimbursements, fees, funding, or salary from an organisation that may in any way gain or lose financially from the publication of this manuscript, either now or in the future? Do you hold any stocks or shares in an organisation that may in any way gain or lose financially from the publication of this manuscript, either now or in the future? Do you hold or are you currently applying for any patents relating to the content of the manuscript? Have you received reimbursements, fees, funding, or salary from an organization that holds or has applied for patents relating to the content of the manuscript? Do you have any other financial competing interests? Do you have any non-financial competing interests in relation to this paper? If you can answer no to all of the above, write 'I declare that I have no competing interests' below. If your reply is yes to any, please give details below.

I declare that I have no competing interests.

I agree to the open peer review policy of the journal. I understand that my name will be included on my report to the authors and, if the manuscript is accepted for publication, my named report including any attachments I upload will be posted on the website along with the authors' responses. I agree for my report to be made available under an Open Access Creative Commons CC-BY license (http://creativecommons.org/licenses/by/4.0/). I understand that any comments which I do not wish to be included in my named report can be included as confidential comments to the editors, which will not be published. I agree to the open peer review policy of the journal.

Authors' response to reviews:(https://drive.google.com/file/d/11TghMmTEJ-uEbgkokf5NOSZc1t2tz8bj/view?usp=sharing)

Source

    © 2020 the Reviewer (CC BY 4.0).

References

    Damien, C., Jan, P., Jerome, R., Guillaume, B., Vladimir, B., J., E. J. Long-read only assembly of Drechmeria coniospora genomes reveals widespread chromosome plasticity and illustrates the limitations of current nanopore methods. GigaScience.