4 ms·
Check out NVBowtie. We don't use standard CPUs to do the alignment. The benchmarks were ran on 2 Nvidia K80s which were in the same blade. To scale it, you need
by rfc 11y ago
Check out NVBowtie. We don't use standard CPUs to do the alignment. The benchmarks were ran on 2 Nvidia K80s which were in the same blade. To scale it, you need Infiniband between the blades (found this out the hard way). The genomes are loaded into memory to reduce read times.
Just to be clear, we're not reading this directly off of the sequencers. Our assumption is the sequenced data is already stored in which we load the data onto the cluster.
I'm not necessarily the technical one of our group unfortunately but, if you're interested, I'd be interested in picking your brain.
- adenadel 11y agoAh, gpus are cheating... :) What I meant about i/o is that if you're reading the data off of disc your read/write time is probably longer than 8 minutes. With SSDs you can obviously go faster. One problem with scaling is that you need a bunch of machines that have the gpus available (although if you can get the whole pipeline under 30min you wouldn't need many machines). The cloud genomics companies mostly use AWS and Google cloud for their compute and I don't know what sorts of non CPU compute resources are available. You would probably be interested in looking into Edico Genomics' Dragen FPGA. Is there a way to PM on HN?
- rfc 11y agoGet in touch with me here: http://bit.ly/1OX84Nb http://bit.ly/1OX84Nb Would love to talk more :)