4 ms·
Is this an accuracy or precision issue? I am imagining that if you actually have access to the device, you could do as many runs as you want, getting to arbitra
by sannee 6y ago
Is this an accuracy or precision issue? I am imagining that if you actually have access to the device, you could do as many runs as you want, getting to arbitrarily low error rates.
- dnautics 6y agodon't know why this was downvoted. If I'm not mistaken, there is generally a high error rate per pore fundamentally because it's a single molecule experiment. These get averaged out, but may be difficult to align as it might not necessarily be a straightforward averaging. There are also segments that are fundamentally generally difficult to sequence correctly (single nucleide runs, not even a super high n) that will probably never get satisfyingly resolved no matter how many times you sequence.
- brofallon 6y agoThis is a common misconception - "averaging out" errors only works if the errors are pretty rare at any given site. This is true for some types of errors & sequencing technologies, but not universally true. Some types of DNA sequences (most notably homopolymers and other simple repeats) are very difficult to sequence correctly, and X% of the reads there will be incorrect. If X>20% of so, then it may look like real germline variation no matter how many reads are sequenced
- koeng 6y agoThe errors are non-random. That's why they use machine learning to figure out those errors. You could, of course, also just do traditional statistics on sequences that you want to sequence all the time. I've done that with plasmids before, and it works pretty good. I think there are a few papers on it too.
- searune 6y ago> The errors are non-random. Could you elaborate / give an example? Are the errors deterministic? Is it like ISI (Inter-Symbol Interference[1]) in signal processing, where some symbols interfere with the reception of the next symbol(s)? Are there short range errors (one letter) or long continuous errors? [1] https://en.wikipedia.org/wiki/Intersymbol_interference https://en.wikipedia.org/wiki/Intersymbol_interference
- koeng 6y agohttps://gist.github.com/Koeng101/abc674e1acd575646748afcbcc74698d https://gist.github.com/Koeng101/abc674e1acd575646748afcbcc7... There is a real example I ran a few months ago. How to read it is here https://en.m.wikipedia.org/wiki/Pileup_format https://en.m.wikipedia.org/wiki/Pileup_format Positions like 172 have errors more often than not because the basecaller is wrong sometimes (note: this is from a sequence verified sample). The errors come up more often in some sequences than they do in others. I’m not really sure about symbol processing, but if you have any beginner resources for that I’d appreciate them!
- marsdentech 6y agoIt's a complicated issue; I tend to think of the error component of any one MinION observation as being a function of the k-mer in the pore at the time (i.e. the subject of the observation) and, with some decaying dependence, the sequences (i.e. in both directions) that extend out from either side of the target k-mer. You might say that MinION error is a function of the target k-mer and its immediate environment. It gets even messier when you try to imagine the form of that function; for one, it's not _completely_ good enough to remain in sequence space alone: among other things, the "shape" (i.e. the conformation) of that (DNA or RNA) molecule around the target k-mer will influence how the shape of the pore will change in response to the target k-mer, which, in turn, will influence the observed current signal (i.e. manifest as a deviation from the "expected" or "ideal" current signal for that k-mer!). As I understand it, Nanopore don't spend too much time actually modelling k-mer-in-pore dwell-mechanics; instead their best base callers use machine learning to generalise across the swathes of available sequencing data for known targets (and give really quite impressive results, all things considered).