5 ms·
I've worked in the area of clinical genomics using whole genome sequencing. Your statement is unfortunately untrue though it perhaps could be if tools were bett
by micro_cam 11y ago
I've worked in the area of clinical genomics using whole genome sequencing. Your statement is unfortunately untrue though it perhaps could be if tools were better written.
While it is easy enough to analyze a single genome on your laptop most current popular analytical tools simply fall over when you start looking at hundreds of genomes on even a large server. Even basic stuff like combining multiple genomes into one file with consistent naming of variants can take entirely ridiculous multi terabyte amounts of ram because the tools to do so just weren't written with this scale in mind.
Most of these tools could (and should) be rewritten to do things without loading the whole data set in memory and work natively in a cluster of commodity clusters. There is some resistance to this of course because scientists prefer to use the published and established methods and often feel new methods need to be published and peer reviewed etc.
Until new tools are written and widely adopted to a large shared memory machine is a bandaid many hospitals and research seam eager to adopt.
- 1971genocide 11y agoMulti Terabytes of RAM ? If it was any other field in computer science people would be really critical of your methodology. The size of the human genome is 21 MB. If you are trying to find the co-ordinate of every cancer cell in a human body then sure, You need a lot of RAM. But the output of the collective field of cancer research doesn't seem to be there yet. So why do you need so much RAM ? Usually when your problem becomes NP-hard. You switch to simpler models. Have you checked the search space for all simpler models ? Or are you sticking to complex models since it helps you publish papers ? You also need to understand that hardware only gets you so far, running a cluster has its own costs - network latency. Most often than not, better techniques are required, rather than than say the tremendous improvement in computational power is not good enough.
- toomuchtodo 11y ago> The size of the human genome is 21 MB. No. > In the real world, right off the genome sequencer: ~200 gigabytes > As a variant file, with just the list of mutations: ~125 megabytes > What this means is that we’d all better brace ourselves for a major flood of genomic data. The 1000 genomes project data, for example, is now available in the AWS cloud and consists of >200 terabytes for the 1700 participants. As the cost of whole genome sequencing continues to drop, bigger and bigger sequencing studies are being rolled out. Just think about the storage requirements of this 10K Autism Genome project, or the UK’s 100k Genome project….. or even.. gasp.. this Million Human Genomes project. The computational demands are staggering, and the big question is: Can data analysis keep up, and what will we learn from this flood of A’s, T’s, G’s and C’s….? https://medium.com/precision-medicine/how-big-is-the-human-genome-e90caa3409b0 https://medium.com/precision-medicine/how-big-is-the-human-g...
- sgt101 11y agoAlso the world of genomics has done fantastic work on compression and if you can compress it further you will probably win a decent award with a ceremony and free booze.
- jarvist 11y agoScientific computing requires a lot of memory, and a lot of computer time. I think it's fair to say that the underlying libraries (LAPACK,ScaLAPACK, Intel's MKL) are the most intensively optimised code in the world. Most of the non trivial algorithms are polynomial in both time and memory. I suspect this Press Release is hinting at a next-generation (cheap, fast) DNA sequencing method. These are derived from Shotgun Sequencing methods, were hundreds of gigabytes of random base pair sequences are reassembled to a coherent genome. The next-generation methods realise cost savings by an even more lossy method of reading smaller fragments of the genome, with much greater computational demands to reassemble.
- a8da6b0c91d 11y agoWhat hard evidence is there that genomics is relevant to cancer treatment, as proven by survival rates? Color me skeptical.
- tom_b 11y agoHmmm? Genomic breast cancer subtypes that each respond to different chemotherapies? http://www.nature.com/nature/journal/v490/n7418/full/nature11412.html http://www.nature.com/nature/journal/v490/n7418/full/nature1...
- a8da6b0c91d 11y agoThat doesn't really say anything about proven treatment efficacy.
- tom_b 11y agoSome information is hidden away in supplemental table 6, which points out candidate drugs to affect different biological pathways for different mutations. You could also skim http://www.nature.com/nature/journal/v406/n6797/full/406747a0.html http://www.nature.com/nature/journal/v406/n6797/full/406747a... for more information about genomic classification of breast cancer. From a treatment prespective, I would say that just glancing at http://ww5.komen.org/BreastCancer/SubtypesofBreastCancer.html http://ww5.komen.org/BreastCancer/SubtypesofBreastCancer.htm... would provide information on treatment decisions generally made by finding appropriate subtype classifications. I think that it is pretty clear that genomic sequencing of patient normal and tumor tissue to find mutations is going to be standard-of-care sooner rather than later, but it is fair to point out that genomic sequencing is not currently standard-of-care. However, I know of studies currently underway that look at variant calls and the possibility of taking action on those calls in ways the involve specifically adding those results back into the patient medical record. I am struggling a bit with how to phrase this, but I don't think you can argue against (1) different subtypes of breast cancer are separate diseases and can be classified by genomic sequencing and (2) treatments for these separate diseases are different and have different efficacies.
- 11y ago
- dunk010 11y agoYes indeed. And new tools are being written - see the Adam project for an interesting example: https://github.com/bigdatagenomics/adam https://github.com/bigdatagenomics/adam and the associated variant caller Avocado: https://github.com/bigdatagenomics/avocado https://github.com/bigdatagenomics/avocado. Others are also trying to get the old tools working on Hadoop, for instance Halvade: https://github.com/ddcap/halvade/wiki/Halvade-Manual https://github.com/ddcap/halvade/wiki/Halvade-Manual, Hadoop-BAM https://github.com/HadoopGenomics/Hadoop-BAM https://github.com/HadoopGenomics/Hadoop-BAM, SeqPig: http://seqpig.sourceforge.net/ http://seqpig.sourceforge.net/, and the guys at BioBankCloud: https://github.com/biobankcloud https://github.com/biobankcloud. It's going to take quite a while for this stuff to get fleshed out, and for researchers to adopt it. But the sheer weight of data is going to force things in the Hadoop direction eventually. It is inevitable.