3 ms·
Thanks for detailed reply! This really puts things into perspective. From the NLP side of the data munching isle, it sounds like a lot more complicated variati
by maga 10y ago
Thanks for detailed reply! This really puts things into perspective.
From the NLP side of the data munching isle, it sounds like a lot more complicated variations on language processing problems.
Comparing reads to each other is like approximate string matching (or fuzzy matching) we use to account for spelling errors in words, but genome chunks are longer and you don't have dictionaries to check against.
And finding the best arrangement of those reads is akin to language modeling where we assign probabilities to word sequences which later can be used to predict the most likely sequence. In case of genomes, though, with so little data, no standard "words", and the error rates in reads, it's like trying to put back a shredded book of in illiterate author only having few other books of illiterate authors as reference.
- dekhn 10y agomy introduction to DNA sequence analysis was "the linguistics of DNA" by David Searlers: https://www.jstor.org/stable/29774782?seq=1#page_scan_tab_contents https://www.jstor.org/stable/29774782?seq=1#page_scan_tab_co... it applies the chomsky hierarchy to sequence analysis, there is in fact a ton of interesting literature around graphical models like HMMs, see https://www.amazon.com/Biological-Sequence-Analysis-Probabilistic-Proteins/dp/0521629713 https://www.amazon.com/Biological-Sequence-Analysis-Probabil... for more details.