7 ms·
Big data in genomics: The $1k genome has arrived
- epistasis 11y agoThough that $1000 produces 100GB of data, after processing there's probably only about ~100MB of features left for machine learning, at most. And most of that will be incidental. With enough data we can hope to have a better filter between signal and noise. Until now, biology has had a huge problem that most big data settings don't: far far more features than labels. With enough patients' data, the matrix will become squarish, but that's a long time from now, still.
- A_Beer_Clinked 11y agoI found this link: https://medium.com/precision-medicine/how-big-is-the-human-genome-e90caa3409b0 https://medium.com/precision-medicine/how-big-is-the-human-g... In summary: > 1. In a perfect world (just your 3 billion letters): ~700 megabytes > 2. In the real world, right off the genome sequencer: ~200 gigabytes > 3. As a variant file, with just the list of mutations: ~125 megabytes
- TheEzEzz 11y agoIf there is only .1% variation then we should be able to get a diff down to ~1MB with some cleverness.
- epistasis 11y agoThat sounds quite feasible, though it hasn't really been worth the effort until we have quite a few more genomes. And typically extra information about the variant (is it in a gene, does it change a protein, etc.) so that extra lookups aren't required during a scan. There are typically 4-6 million variants discovered through this method of genome sequencing in a normal genome. A simple variant consists of a genome coordinate at ~32 bits (one of 3.2e9), and the change from the reference, which is a x,y index into {A, C, G, T}^2, at ~4 bits. Typically the coordinates are spaced on average ~1k bases apart, so the coordinate could probably be squeezed into ~15bits with clever encoding. So a naive encoding of this information gets to 27MB, and that could probably be shoved down into 10MB if coordinates are deltas from the previous, rather than absolute. 1MB seems feasible, but with diminishing returns computationally.
- BKPetkov 11y agoInteresting - wouldn't you need to have access to "the rest" of the genomes that you are comparing against? In other words, wouldn't you need to keep that ~100 GB from the newly sequenced genome in temporary storage while comparing against the rest of the database stored somewhere in the cloud, before then condensing the new genome into a variant file?
- epistasis 11y agoTypically we don't look at other genomes while we find the variants in an individual genome. Each genome is analyzed against the "reference" human genome, which is an average of 10 individuals. This forms the coordinate basis that is shared for everyone else. Pretty much all genomic data uses a reference genome as the basis. This is versioned, and has a bug tracker, etc., for various regions that have been difficult to assemble. The flow is: 1. BCL (scans of the glass slide) 2. FASTQ (individual short reads and quality scores, unsorted and in random order) 3. BAM (individual short reads aligned to the reference genome) 4. VCF (the "diff" vs. the reference genome) All of this can be done with <10GB of reference data and code, where the reference data is the current human genome, a burrows-wheeler transform of the human genome, gene locations, and dbSNP (the database of common human variation).
- BKPetkov 11y agoHow long does it take (and with what computational bandwidth) to produce a 125MB variant file from 200GB raw sequence data?
- adenadel 11y agoDepending on the pipeline you use and the compute resources available you could have a full workflow done in anywhere from several hours to a couple days. Illumina BaseSpace is free (for now) and has some example data sets with a bunch of canned pipelines for analysis if you're interested in trying it for yourself. https://basespace.illumina.com/ https://basespace.illumina.com/
- jghn 11y agoYou're not going to VCF on a whole genome in several hours.
- gbveiuwlbiu 11y agoCould you please elaborate on this?
- adenadel 11y agoWith particular hardware and software you can. Edico Dragen claims speeds for bcl -> vcf of 20 minutes [1]. With Microsoft Research's snap aligner and 450GB of memory you can get whole genome alignment in ~30 minutes and then variant calling can be done in a couple hours. 1. http://www.edicogenome.com/dragen/dragen-gp/ http://www.edicogenome.com/dragen/dragen-gp/
- arca_vorago 11y agoI've seen 200GB runs take 4 days, I've seen runs take 3 hours. Depends on your computing structure but more importantly is your IO. High CPU core counts and high speed storage access make a big difference, as does distributing the computational workload.
- deleted 11y ago
- bayesianhorse 11y agoIt's hard to quantify how much data is left after "analysis". An assembled genome sequence, at least a certain quality, is actually useful beyond machine learning. For example loss-of-function mutations can be detected without resorting to other genomes beyond the reference. Inversions, translocations and copy number variations can also give clues to illness.
- jerven 11y agoIf for 50% of the genes (human) we have no idea of function, then its hard to determine at this point what is a loss-of-function mutation in the first place. Source: I am the lead web developer for uniprot.org, and I deal with lack of real info daily as does everyone in the Swiss-Prot team.
- gbveiuwlbiu 11y agoApart from data privacy and formal restrictions on sharing, what would you say the reason for this lack of public data is? Am I right to say that the bandwidth and storage needed to upload whole genomes is prohibitive?
- jerven 11y agoWe collectively just don't know. Not even in the cutting edge literature. Humanity, only has a basic understanding of what the biochemical role of many proteins is, and due too that also a similar idea of what the genes role is in the rather complicated system of making a functioning human. Bandwidth and storage are infrastructure issues that could be solved with enough money. In biology we are still lacking good technical solutions to do the chemistry and even something "simple" like getting a protein crystalized so we can determine its 3D structure is not trivial or cheap yet. Fast genome and RNA sequencing are massive improvements and really help. But the basic understanding of what all those genes, regulatory and other parts do is still relatively primitive. We often "complain" that we wish that people stopped sequencing and went back to do more biochemistry for functional characterization instead ;)
- 11y ago
- noname123 11y ago>Though that $1000 produces 100GB of data Not sure if you are referring to the raw sequence data coming from the sequencers, nowadays the standard practice is to align a specific sample's genome (e.g., an individual's genome) to the reference genome (the Human Genome Project) and store only your variations against that reference genome, stored as a BAM (binary alignment) file. Furthermore, for most clinical cases, if you only care about the current known SNP (single nucleotide polymorphism in coding genes), you can generate a VCF file at those known SNP sites, compressing the data further. This is the way most population genetics is done by comparing the VCF files of cohorts, not doing analysis of single genomes of patients one by one.
- adenadel 11y agoYou've got a bit of misinformation in here. He's correct about the raw sequence data for a 30X whole human genome being 100GB of data. The standard practice is to align the sample to the reference, but we store alignments in the BAM file, not variations. Some of the mismatches in the alignments are due to noise in the sequencer. We run the BAM through a variant caller which outputs a VCF which contains the variants. You do not typically genotype at only known SNP sites when generating a VCF file (although you may restrict by a region). If you were only interested in certain SNP sites you would be better off running a microarray.
- jerven 11y agoOf course its a VCF file which we hope contains the variants ;( the tech is still fickle and so is some of the biology... Soon we are going to go towards variant graphs and VCFs will disappear again, slowly (IMHO).
- searine 11y agoThe bottleneck in genomics hasn't been cost since about 2012. The chokepoint is analysis. Most biologists don't know how to program, and most programmers don't know biological context, neither know statistics well. I work with so many scientists whose only thought is to sequence first and ask questions later. Usually all the real work ends up falling on the shoulders of one skilled researcher while the rest look on like some unionized road crew. It's only going to get worse, but the good news is, if you are one of the biologists who can program and use statistics then you're in good shape. There is already so much idle data out there already, that you'll never have to spend a dime on sequencing.
- adenadel 11y agoPeople say this all the time, but with some of the most common applications of high throughput sequencing there are very good canned solutions (using open source software) that you can pay for. DNAnexus, Seven Bridges, and Illumina BaseSpace all provide cloud storage and analysis. Unless you are doing a custom prep for your sequencing one of these probably has an analysis solution for you.
- BKPetkov 11y agoIs the time required to upload data to the cloud ever a problem with these solutions? Of course, it depends on what you are trying to do, but suppose you were working with thousands of genomes?
- adenadel 11y agoSure, all the time. Network and I/O are the biggest blockers for sequence analysis. For any organization that is working with thousands of genomes they probably have their own compute resources. I know of at least one organization who is currently sending thousands of genomes to the cloud for analysis, so it's certainly feasible to some extent.
- gbveiuwlbiu 11y ago
- mjpuser 11y agoA quick search for public genome data led me here http://www.completegenomics.com/public-data/69-Genomes/ http://www.completegenomics.com/public-data/69-Genomes/. I wonder if there will be a time when you can search millions through a rich UI, comparing your own genome with others, etc.
- aheilbut 11y agoHave a look at ExAC, which has 60k exomes and a pretty snazzy UI: http://exac.broadinstitute.org http://exac.broadinstitute.org
- jhull 11y agoFor more genomic data check out www.solvebio.com You'll need to register for a (free) API key to access public data, although we're removing that requirement in the next few days (I work here.)
- bayesianhorse 11y agoI think the currenntly more interesting application in sequencing is pathogen detection. These have smaller and simpler genomes (mostly), and tracking them and their features improves epidemiology and the choice of treatments.
- noname123 11y agonextflu.org and wwarn.org and ebola.nextfu.org NextFlu takes global samples uploaded to a global flu genomic sample database and construct a phylogeny tree to monitor how influenza evolve and I think the authors behind the project wants to make predictions as to which cohort of influenza variation will become dominant. Wwarn.org I believes tracks the emergence of Artesminin resistance (front-line drug in malaria treatment) in malaria in SE Asia and tries to map it out on GIS to inform public health officials further from SE Asia how it is spreading to their region (India, Africa where current Artesminin resistance gene is only 5% while Artesminin resistance is already the dominant wild type in SE Asia) and whether to modify front-line treatment protocol.
- damurdock 11y agoPathogen identification is indeed a very exciting application for NGS. In case you're interested, here[0] is a paper about a tool called SURPI (Sequence-based Ultra-Rapid Pathogen Identification) which was designed for that purpose. Also, here[1] is a case report from the NEJM where SURPI was used to diagnose a patient with Neuroleptospirosis, which allowed him to be treated quickly and eventually recover. SURPI isn't the only horse in this game, of course, but I've worked with it before so it immediately came to mind. [0]: "A cloud-compatible bioinformatics pipeline for ultrarapid pathogen identification from next-generation sequencing of clinical samples" http://genome.cshlp.org/content/24/7/1180.long http://genome.cshlp.org/content/24/7/1180.long [1]: "Actionable Diagnosis of Neuroleptospirosis by Next-Generation Sequencing" http://www.nejm.org/doi/full/10.1056/NEJMoa1401268 http://www.nejm.org/doi/full/10.1056/NEJMoa1401268
- tridint 11y agoThe clinical work they're doing is great, but the code is problematic. Its a bunch of Perl and Python duct taped together with shell scripts. From the github repo: Shell 84.3% Perl 8.9% Python 6.3% C 0.5% Check out the source https://github.com/chiulab/surpi https://github.com/chiulab/surpi
- mschuster91 11y agoWhat I find most worrying with mass-collecting fully sequenced DNA (or DNA at all) is that it is inevitable that law enforcement, military, spy agencies or politicians will want access to the data. Given enough sequenced DNA and a free-fall of sequencing costs, it might be feasible in the future to find out who exactly took a dump in public just due to DNA. And well, the Israelis are already doing this with dog dumps (http://uk.reuters.com/article/2008/09/16/uk-israel-dogs-idUKLG37942520080916 http://uk.reuters.com/article/2008/09/16/uk-israel-dogs-idUK...), so applying the same tech to humans is not far away. Fucking scary if you ask me.
- deleted 11y ago[deleted]