3 ms·
What if this is just an artifact of the DNA sequence assembly. They soil they collect probably contains DNA from many organisms. The sequence assembly algorithm
by bno1 5y ago
What if this is just an artifact of the DNA sequence assembly. They soil they collect probably contains DNA from many organisms. The sequence assembly algorithms are likely looking for patterns in noise [1].
[1] https://en.m.wikipedia.org/wiki/Apophenia https://en.m.wikipedia.org/wiki/Apophenia
- kasperset 5y agoThis could be the case. I remember this controversy few years ago about Tardigrade genome. The problem was that other paper said that "foreign" dna was the probable reason behind cryptobiosis in Tardigrade. This particular paper found that it was mostly likely a contamination. Koutsovoulos, G., Kumar, S., Laetsch, D. R., Stevens, L., Daub, J., Conlon, C., Maroon, H., Thomas, F., Aboobaker, A. A., & Blaxter, M. (2016). No evidence for extensive horizontal gene transfer in the genome of the tardigrade Hypsibius dujardini. Proceedings of the National Academy of Sciences, 113(18), 5053–5058. https://doi.org/10.1073/pnas.1600338113 https://doi.org/10.1073/pnas.1600338113
- noname123 5y agoYes. This type of error could be easily determined. If it's an contamination or an error, it will have very low read coverage (typical sequencing project be 50-200x meaning if you realign and pile up the reads back to the main assembly, you get a normal distribution centered around 50-200 reads supporting a particular base or region. If you have a very low coverage region or a sharp drop off in coverage, then it's most likely an error. Metagenomics assembly also do additional binning prior to assembly based on the abundance and GC content of the cluster reads to separate out the different taxas of the sample. How well the read clusters are distanced by this huerisirc is another measure of quality of assembly.
- f6v 5y ago> If it's an contamination or an error, it will have very low read coverage (typical sequencing project be 50-200x meaning if you realign and pile up the reads back to the main assembly, you get a normal distribution centered around 50-200 reads supporting a particular base or region. If you have a very low coverage region or a sharp drop off in coverage, then it's most likely an error. Wouldn’t sequence-specific biases(capture efficiency, amplification) result in distortions?
- andrewon 5y agoWas thinking the same and wondering how did they obtain the million-base long sequences. Turned out they used short read sequencing with 150 or 250 bp reads and computationally assembled the long reads [1]. While this is a traditionally valid method, the newer long read sequencing technologies such as Oxford Nanopore or PacBio would be an more appropriate and direct method. [1] https://www.biorxiv.org/content/10.1101/2021.07.10.451761v1.full.pdf https://www.biorxiv.org/content/10.1101/2021.07.10.451761v1....
- f6v 5y agoLong reads are still not as widely used since the error rates are much higher than in shotgun approach.
- tdido 5y agoActually, for the use case of genome assemblies you can compensate for the error rate with depth of coverage. Long reads are already the state-of-the-art for assembling genomes and are allowing us to get information that is simply invisible using short reads. https://www.biorxiv.org/content/10.1101/2021.05.26.445798v1 https://www.biorxiv.org/content/10.1101/2021.05.26.445798v1
- f6v 5y agoIt's definitely getting better, but errors have to be corrected computationally, and it still seems like a big challenge.
- londons_explore 5y agoOne would think a hybrid approach will bring the best results...? Long reads help with alignment and arrangement, and short reads eliminate small errors of just a few base pairs.
- ejstronge 5y agoPacBio reads don't extend to the million base length and are a considerably more expensive approach than short read sequencing if assembly is happening anyway. ONT is a great tool but the goal of this report isn't necessarily to show that these molecules are 1 million bases in length vs showing that they represent novel DNA sequences. In summary, ONT and PacBio are neither more appropriate or necessarily more direct methods here.
- greazy 5y agoVery unlikely, the algos used to assembly genomes can make errors but not to such an extent.
- f6v 5y agoI bet you can calculate a probability that random short reads from a bunch of organisms can be assembled in a single 1 megabase-long sequence. And that probability is going to be low.