4 ms·
I notice that the article fails to mention how long it would take to extract all 700 terabytes of data... Assuming 5.5 petabits stored with 1 base pair represe
by tarice 14y ago
I notice that the article fails to mention how long it would take to extract all 700 terabytes of data...
Assuming 5.5 petabits stored with 1 base pair representing 1 bit, we can extrapolate the time required to extract the data based off the time taken to sequence the human genome (3 billion base pairs).
5.5 petabits / 3 billion bits ~= 2 million, so theoretically it should take 2 million times longer to sequence the original.
3 years ago, there was an Ars Technica article about how it now only takes 1 month to sequence a human genome[1]; the article now claims that microfluidic chips can perform the same task in hours.
Assuming 2 hours (low end) to sequence the human genome:
2 hours * 2 million = 4 million hours = 456 years, give or take a few years.
So, maybe not so great for storing enormous amounts of data. But if you want to store 1 GB, it would only take ~6 hours. Not too bad.
[1]http://arstechnica.com/science/2009/08/human-genome-completed-using-one-machine-for-four-weeks/ http://arstechnica.com/science/2009/08/human-genome-complete...
- scarmig 14y agoHmm. DNA is fairly easy to duplicate though, right? Wouldn't that allow an exponential speedup?
- tarice 14y agoI assumed that the microfluidic chip speed listed would include parallel processing. Even if it doesn't, you'd still need something like 5,000 experiments in parallel for it to take less than a month...
- JunkDNA 14y agoYes, the microfluidics used today make use of this for reading large numbers of small segments of DNA in parallel. The current "gold standard" for DNA sequencing (manufactured by Illumina) uses millions of tiny fragments of DNA which are read optically as DNA sequence is extended.
- epistasis 14y agoI think that DNA, if ever used in practice, will be primarily used for long-term, archival storage of data that's rarely accessed, if ever. However there are a couple of things that alleviate the problems with reading the DNA. 1) DNA sequencing technology is currently advancing much faster than silicon technology, so give it time and it's likely that it will catch up with current hard disk reading speed at comparable sizes and volumes. 2) DNA's self-hybridizing nature makes it easy to pull out blocks with specific addresses (if you wait for the hybridization). So if you include address labels in the DNA as you write it out, you can probably pull it out in chunks of kilobase to megabase at a time. 3) As the other commenter pointed out, this is extremely easy to parallelize. So if you want to go twice as fast, divide the sample in half put half in each machine. Dilute and pipette as necessary.
- skosuri 14y agowe talk about a potential petabyte storage mechanism in the supplement: paper http://db.tt/ZDoDJZeD http://db.tt/ZDoDJZeD supplement http://db.tt/elIqsy72 http://db.tt/elIqsy72 tldr; we are 6-8 orders of magnitude away from doing petabytes routinely; that said costs of sequencing/synthesis have seen such drops over the last decade or so. there are many barriers though for that continuing for another decade