4 ms·
> It takes less than a day and costs less than $1000 to sequence your exome (and even your whole genome), but the backlog for analysis of the sequencing results
by bozoUser 10y ago
> It takes less than a day and costs less than $1000 to sequence your exome (and even your whole genome), but the backlog for analysis of the sequencing results in labs can be 9 months or more.
Layman qs. what is stopping the labs to quickly analyze the genome? Computational power or few labs doing this kind of work?
- jfarlow 10y agoIn general it's a (computationally) hard problem to restitch a genome together. Even today, when you 'get your genome sequenced' you are not getting a full read-through of your entire genome's data. Imagine you want to reconstruct the data on two RAIDs that are mostly, but importantly not exactly, mirrors of each other. Each RAID has 23 drives. Each drive has ~1Gb or so of data. And much of the data is not only mirrored between the two RAIDs, but is also mirrored between the 23 drives - and many of that mirrored data is 'off by 1' in very important ways (both 'must', and 'must not' scenarios). Further some of the data contains very long sections of highly repetitive data. And some of the data is mechanically biased to be harder to read than others. You must now reconstruct those two RAIDs with single-bit accuracy - as a single bit-flip in certain sections determines whether or not you get cancer. The data you are given to do the reconstruction is a 200Gb single column CSV file with each row being 12 bytes of data. Go.
- bozoUser 10y agohmm I don`t think so I fathom the complete complexity of the process but with so many powerful GPU`s out there, is there a possibility of reconstruction in a matter of days if not hours?
- JangoSteve 10y agoYes, it depends on the size of the sequencing panel that was done (i.e. a targeted panel for a specific gene or set of genes versus whole exome versus whole genome). But even for whole exome, you're talking a few hours or faster depending on hardware. In addition to our company, Genomenon (which seeks to speed up the interpretation time required to analyze the data _after_ it's been computationally aligned and annotated), I'm also friends with another startup down the street, called Parabricks, which seeks to speed up this alignment process (aka secondary analysis) even further.
- virtuabhi 10y agoThe steps in the genomic analysis pipelines are not always embarrassingly parallel. The degree of parallelization cannot be increased to an arbitary number. In addition, the parallel executions can give slightly different results when compared with the serial output, so we need "safe" data-partitioning schemes and rigorous error control. If you are further interested in parallelization schemes for genomic pipelines, please have a look at our paper on the strengths and limitations of big data technology for genomic analysis (published last week) - https://people.cs.umass.edu/~aroy/sigmod17-roy.pdf https://people.cs.umass.edu/~aroy/sigmod17-roy.pdf
- j6m8 10y agoA little column A, a little column B. Something worth thinking about: Processing on HIPAA-compliant resources is expensive, and as with any medical procedure, that financing/insurance is complicated to bill for (or at least, takes a long time... which makes it more expensive to run, etc).
- JangoSteve 10y agoIt's mostly not about computational power. The stitching together that jfarlow mentions is part of the "secondary" analysis where the raw genome data must be put together, but that's mostly a solved problem, as there are plenty of gold-standard open-source libraries that employ statistically complex calculations to align raw data to the current version of the human reference genome. That's part of what takes less than a day along with the primary sequencing. It's constantly being improved, but most labs would not consider this an issue that keeps them up at night, as what we currently have works reasonably well. The part that takes a long time (i.e. the "bioinformatics bottleneck" I referred to), is that once the sequencing data is stitched together, you end up with a ton of variants, and you don't know which (if any) are clinically significant. Imagine that each nucleotide in your genome is a marble, and that the entirety of your sequenced genome is a 1-story building filled with marbles (that's how many nucleotides are in your genome), and that each one is supposed to be a specific color out of 4 possible colors. Now imagine that 10,000 (or more) of those marbles are the wrong color. Primary analysis (putting your sample into a machine and essentially getting back a list of what color all your marbles are), as well as secondary analysis (i.e. the process of sorting your marbles into the correct order so that you can actually tell _which_ marbles specifically are the wrong color) together are what cost less than $1000 and take less than a day. The real problem is that you may find that 10,000 (or more) of your marbles are the "wrong" color, but 9,900 of them make absolutely no difference in a clinical sense. To be clear, a marble that's the wrong color is a "mutation", or "variant". Maybe this mutation makes my eyes slightly bluer, or my finger nails a little harder. In other words, the part that takes a long time is actually going through each mutation and figuring out which one (out of the 10,000 variants) is clinically significant or relevant to the disease/symptoms you seem to have, and then figuring out if there's a known treatment for that particular root cause. Currently, this is done by employing MD PhDs to look through each patient's sequenced data, and then cross-referencing that with the millions of published studies to see what has ever been seen before, or is known to be associated with some disease. And it can take a human hours to a day to do this per patient. So, the number of patients times the amount of time it takes per patient, divided by the number of MD PhDs a lab can hire to do this, is what leads to the backlog. So, actually yes, it is computational power... but it's human computational power.
- 10y ago