3 ms·
I'm a product manager. Side project is building a genome data ingestion pipeline with my dad. We're both really passionate about giving people a faster and chea
by rfc 11y ago
I'm a product manager. Side project is building a genome data ingestion pipeline with my dad. We're both really passionate about giving people a faster and cheaper solution to processing genome data.
Current project is to allow researchers to send fully sequenced genomes to our server, align the sequenced genome to a referenced genome, then store the data in a performant way that can be queried against. Right now, its either too technical for the researchers or too expensive. I have dreams of making it something big but time is a limited resource as well as finding people who A) know how to build massive data platforms at scale and B) know enough about genomics to build the platform.
We're coming along well so far. The alignment of sequenced genomes to referenced genomes if done (although not scale ready yet). Currently working on/learning the data storage side and what to do there.
Shameless plug: if anyones interested in our project, we could use some help.
- aryamaan 11y agoSounds interesting.
- rfc 11y agoThanks. We're pretty excited about it. There's lots of big challenges to solve, especially on the data from. Full 30x coverage of a human genome is (uncompressed) around 180gb. So the core problem we're trying to solve right now is: 1) Compressing the files locally and sending via FTP to server 2) Aligning 180gb files against the reference genome (180gb also) in less than 5 minutes 3) Storing that data in a performant way so that you can do comparative genomics (compare multiple genomes against each other) 4) Do the above with 100's-1,000's of genomes a day. It's pretty rewarding. Apart from the cool data side, it's feels good to know that building something like this could really help accelerate critical research in finding diseases.
- adenadel 11y agoIs anyone using your pipeline? There are lots of fully developed commercial pipelines for human WGS (Illumina BaseSpace, DNAnexus, Seven Bridges, etc.) Depending on the hardware you have available bowtie2 might not be the best aligner. What are you using for variant calling? If you're using the GATK Haplotype caller you're going to have issues with commercial licensing.
- rfc 11y agoNo one is using it yet. Still just a side project :) Seven Bridge & DNAnexus probably have the most complete/thorough system (imo). The challenge though is doing it fast, at scale, and user friendly enough. Seven Bridges is cost effective and can definitely scale but it's not necessarily the fastest solution on the market. While these guys offer the full end-to-end solution, we're just trying to focus specifically on accelerating the pipeline. Maybe in the future we'll get to the full circle, but that's not our focus. We use a variation of bowtie2 that allows us to scale well. We're able to align a 30x genome right now in about 8 minutes and are doing a couple more tests that might get us down into the <5 min range. The goal is to be able to do multiple at once in <5min with our portion of the pipeline taking less than 10 min total. There are universities, such as Harvard, that are trying to align 1,000s of these a month but the current providers can't keep up. So, we're seeing if we can provide something that can help them out.
- adenadel 11y agoDon't worry, I work for a company in the space and fully understand the challenges :) I have a really hard time believing that you can align a 30x genome in 8 minutes because the I/O time is longer than that, and bowtie doesn't use one of the faster algorithms for alignment (usually the kmer hashing methods are faster when you have machines with enough memory. Microsoft's snap and Illumina's isaac are two examples). What kind of hardware are you benchmarking on?
- rfc 11y agoCheck out NVBowtie. We don't use standard CPUs to do the alignment. The benchmarks were ran on 2 Nvidia K80s which were in the same blade. To scale it, you need Infiniband between the blades (found this out the hard way). The genomes are loaded into memory to reduce read times. Just to be clear, we're not reading this directly off of the sequencers. Our assumption is the sequenced data is already stored in which we load the data onto the cluster. I'm not necessarily the technical one of our group unfortunately but, if you're interested, I'd be interested in picking your brain.
- noname123 11y agoOut of curiosity, what genome assembler are you using for your project? (I'm having to build a genome assembly pipeline with a reference assembly as well for a bioinformatics class I'm auditing at a local college, not a professional in this field either). The papers I am reading and trying to replicate uses SMALT.
- rfc 11y agoAssuming you're referring to the alignment side of things. We use a variation the Bowtie2 algorithm that allows us to align multiple genomes at once to the same reference genome.
- noname123 11y agoThanks for your reply, rfc. In my Bioinformatics class, we went through Bowtie algorithm funny enough last week (the vague details I still remember are the funny way it compresses fragments as rotations and then goes onto transform the rotations). Gl on your project.