4 ms·
> The errors are non-random. Could you elaborate / give an example? Are the errors deterministic? Is it like ISI (Inter-Symbol Interference[1]) in signal proce
by searune 6y ago
> The errors are non-random.
Could you elaborate / give an example? Are the errors deterministic? Is it like ISI (Inter-Symbol Interference[1]) in signal processing, where some symbols interfere with the reception of the next symbol(s)? Are there short range errors (one letter) or long continuous errors?
[1] https://en.wikipedia.org/wiki/Intersymbol_interference https://en.wikipedia.org/wiki/Intersymbol_interference
- koeng 6y agohttps://gist.github.com/Koeng101/abc674e1acd575646748afcbcc74698d https://gist.github.com/Koeng101/abc674e1acd575646748afcbcc7... There is a real example I ran a few months ago. How to read it is here https://en.m.wikipedia.org/wiki/Pileup_format https://en.m.wikipedia.org/wiki/Pileup_format Positions like 172 have errors more often than not because the basecaller is wrong sometimes (note: this is from a sequence verified sample). The errors come up more often in some sequences than they do in others. I’m not really sure about symbol processing, but if you have any beginner resources for that I’d appreciate them!
- marsdentech 6y agoIt's a complicated issue; I tend to think of the error component of any one MinION observation as being a function of the k-mer in the pore at the time (i.e. the subject of the observation) and, with some decaying dependence, the sequences (i.e. in both directions) that extend out from either side of the target k-mer. You might say that MinION error is a function of the target k-mer and its immediate environment. It gets even messier when you try to imagine the form of that function; for one, it's not _completely_ good enough to remain in sequence space alone: among other things, the "shape" (i.e. the conformation) of that (DNA or RNA) molecule around the target k-mer will influence how the shape of the pore will change in response to the target k-mer, which, in turn, will influence the observed current signal (i.e. manifest as a deviation from the "expected" or "ideal" current signal for that k-mer!). As I understand it, Nanopore don't spend too much time actually modelling k-mer-in-pore dwell-mechanics; instead their best base callers use machine learning to generalise across the swathes of available sequencing data for known targets (and give really quite impressive results, all things considered).