4 ms·
No one is using it yet. Still just a side project :) Seven Bridge & DNAnexus probably have the most complete/thorough system (imo). The challenge though is doi
by rfc 11y ago
No one is using it yet. Still just a side project :)
Seven Bridge & DNAnexus probably have the most complete/thorough system (imo). The challenge though is doing it fast, at scale, and user friendly enough. Seven Bridges is cost effective and can definitely scale but it's not necessarily the fastest solution on the market. While these guys offer the full end-to-end solution, we're just trying to focus specifically on accelerating the pipeline. Maybe in the future we'll get to the full circle, but that's not our focus.
We use a variation of bowtie2 that allows us to scale well. We're able to align a 30x genome right now in about 8 minutes and are doing a couple more tests that might get us down into the <5 min range. The goal is to be able to do multiple at once in <5min with our portion of the pipeline taking less than 10 min total.
There are universities, such as Harvard, that are trying to align 1,000s of these a month but the current providers can't keep up. So, we're seeing if we can provide something that can help them out.
- adenadel 11y agoDon't worry, I work for a company in the space and fully understand the challenges :) I have a really hard time believing that you can align a 30x genome in 8 minutes because the I/O time is longer than that, and bowtie doesn't use one of the faster algorithms for alignment (usually the kmer hashing methods are faster when you have machines with enough memory. Microsoft's snap and Illumina's isaac are two examples). What kind of hardware are you benchmarking on?
- rfc 11y agoCheck out NVBowtie. We don't use standard CPUs to do the alignment. The benchmarks were ran on 2 Nvidia K80s which were in the same blade. To scale it, you need Infiniband between the blades (found this out the hard way). The genomes are loaded into memory to reduce read times. Just to be clear, we're not reading this directly off of the sequencers. Our assumption is the sequenced data is already stored in which we load the data onto the cluster. I'm not necessarily the technical one of our group unfortunately but, if you're interested, I'd be interested in picking your brain.
- adenadel 11y agoAh, gpus are cheating... :) What I meant about i/o is that if you're reading the data off of disc your read/write time is probably longer than 8 minutes. With SSDs you can obviously go faster. One problem with scaling is that you need a bunch of machines that have the gpus available (although if you can get the whole pipeline under 30min you wouldn't need many machines). The cloud genomics companies mostly use AWS and Google cloud for their compute and I don't know what sorts of non CPU compute resources are available. You would probably be interested in looking into Edico Genomics' Dragen FPGA. Is there a way to PM on HN?
- rfc 11y agoGet in touch with me here: http://bit.ly/1OX84Nb http://bit.ly/1OX84Nb Would love to talk more :)