3 ms·
It's mostly not about computational power. The stitching together that jfarlow mentions is part of the "secondary" analysis where the raw genome data must be pu
by JangoSteve 10y ago
It's mostly not about computational power. The stitching together that jfarlow mentions is part of the "secondary" analysis where the raw genome data must be put together, but that's mostly a solved problem, as there are plenty of gold-standard open-source libraries that employ statistically complex calculations to align raw data to the current version of the human reference genome. That's part of what takes less than a day along with the primary sequencing. It's constantly being improved, but most labs would not consider this an issue that keeps them up at night, as what we currently have works reasonably well.
The part that takes a long time (i.e. the "bioinformatics bottleneck" I referred to), is that once the sequencing data is stitched together, you end up with a ton of variants, and you don't know which (if any) are clinically significant.
Imagine that each nucleotide in your genome is a marble, and that the entirety of your sequenced genome is a 1-story building filled with marbles (that's how many nucleotides are in your genome), and that each one is supposed to be a specific color out of 4 possible colors. Now imagine that 10,000 (or more) of those marbles are the wrong color.
Primary analysis (putting your sample into a machine and essentially getting back a list of what color all your marbles are), as well as secondary analysis (i.e. the process of sorting your marbles into the correct order so that you can actually tell _which_ marbles specifically are the wrong color) together are what cost less than $1000 and take less than a day.
The real problem is that you may find that 10,000 (or more) of your marbles are the "wrong" color, but 9,900 of them make absolutely no difference in a clinical sense. To be clear, a marble that's the wrong color is a "mutation", or "variant". Maybe this mutation makes my eyes slightly bluer, or my finger nails a little harder.
In other words, the part that takes a long time is actually going through each mutation and figuring out which one (out of the 10,000 variants) is clinically significant or relevant to the disease/symptoms you seem to have, and then figuring out if there's a known treatment for that particular root cause. Currently, this is done by employing MD PhDs to look through each patient's sequenced data, and then cross-referencing that with the millions of published studies to see what has ever been seen before, or is known to be associated with some disease.
And it can take a human hours to a day to do this per patient. So, the number of patients times the amount of time it takes per patient, divided by the number of MD PhDs a lab can hire to do this, is what leads to the backlog.
So, actually yes, it is computational power... but it's human computational power.
- jfarlow 10y agoI'm curious if you think your primary sequence data is strong enough to support confident automated lookups. Do you computationally prioritize the mutations prior to human annotation? If your primary sequence is good enough it should be pretty straightforward to look for coding sequences, mutations which create truncations, mutations which change amino acid charge, mutations in proteins known to be oncogenes, etc.
- JangoSteve 10y agoThis is currently done; it's part of the filtering and annotation process that happens at the tale end of the secondary analysis. Everything I'm talking about comes after that. This is part of why the human interpretation part takes hours to a day, instead of a day to several days. However, it's far from perfect and can always be improved. That's one of the things we're trying to help with, is to do further pre-processing on the knowledge-base of all genomic information contained in all published literature, to try to drive the current hours-to-a-day timeframe down to minutes-to-an-hour, or just minutes. One issue though, is that it's a constant tug-of-war optimizing between false-positives (giving you too many variants to manually review) and false-negatives (removing variants due to some data threshold which may have been clinically significant).