4 ms·
> Typically, a DNA sequencing machine that’s processing the entire genome of a human will generate tens to hundreds of gigabytes of data. When stored, the cumul
by ziga 8y ago
> Typically, a DNA sequencing machine that’s processing the entire genome of a human will generate tens to hundreds of gigabytes of data. When stored, the cumulative data of millions of genomes will occupy dozens of exabytes.
100GB * 1M = 100 petabytes, not exabytes.
It's also worth quantifying the cost of storing this data. Storing 100GB on Amazon Glacier costs ~$5/year, which is still a small fraction of the total cost of whole genome sequencing.
- tofof 8y agoBut researchers don't need to simply backup the data - that is to say write the data once and have it sitting there as a recovery option. That's what Glacier is good for. We need to be able to interact and compute with the data, search it, compare it, etc. We need to be able to practically store several hundred individuals' genomes for even a modest (n=100) GWAS study. And as the article explains, it's not simply a question of storage, but also of compressing in a way that we can still quickly do computations on the data without having to resort to just completely uncompressing everything. I'm going to actually run the numbers to show why Glacier would be an exceptionally poor option, but I'll disclaim up front: the poor fit will stem from the fact that Glacier is optimal for store-forever read-never workflows, which is not even close to what our workflow will be. Let's actually do the math on that modest n=100 study. At 150 GB per individual (see evandijk70's comment), you seem to be thinking 'ok, 15 TB[1], $5/GB/year = $750, no problem'. Now let's actually add the cost of not just the size of the data stored, but the retrievals and data transfers. Let's even assume we're willing to wait the 6 hours (an entire workday!) for a standard Glacier retrieval. If we're extremely conservative about how we use the data, we can possibly get by with only 10 retrievals of each genome: 10 x 15 TB = 150000 gigabytes retrieved, at $0.01 = $1500. Plus the data transfer charge, $0.09 for the first 10000 GB, $0.085 for the next 40000 GB, $0.07 for the remainder: $11300 [2]. So for a small study, we're talking about not $750, but $13,550 using your proposal of Amazon Glacier. And then the real kicker - if we actually tried to do this, moving 165 TB (1 write, 10 reads) of data around would take more than 5 straight months of 24/7 uploading or downloading at 100 Mbps. And at $13,550, that's literally more than half a biology grad student's salary here at the University of Illinois. For two of these studies, you could just instead hire a third scientist and pay them to do nothing but drive harddrives around. Obviously that's ridiculous and in no way a realistic solution to anything. Admittedly, today, small GWAS studies are probably closer to n=25, with 30x coverage rather than 100x. But lopping a zero off the costs and transfer times doesn't change anything - it would still be an absurd use of Glacier. But the example should hopefully be eye-opening to the importance of being able to effectively manage this volume of data. The example is the output of a single scientist working for perhaps a month on the actual sequencing, and a couple more months in preparation to get a really good/narrow selection of subjects, etc. So for a 10 person lab working at high efficiency, you can imagine how quickly the data could stack up. All our data operations -- search, compute on, transfer, and store -- would be improved dramatically if we had compression schemes that approached the efficiency of, say, HEVC or AV1. As an comparison - uncompressed video, 1920x816 pixels (standard cinema 2.35:1 widescreen), at 23.976 fps, for 1 hr 45 minutes, in 24 bit color = 1920x816 x 23.976 x (60x60+60x45) x 24/8 = 700 GB of data. HEVC at reasonable settings will chop that to about... 2 GB. That sort of 350-fold compression is what this article is talking about hoping to achieve, which is a far cry from the 20-fold you can get from gzipping a genome. [1] An individual archive on Glacier, by the way, is limited to 40 terabytes, so we're already pushing against that boundary. [2] The limit on data transfers per month is 500 TB, so with our single study at 150 we're already pushing against that boundary as well.
- ziga 8y agoThe article mentioned storing data for a decades, which implies a backup use-case with infrequent access (of course, data only would be deposited in Glacier after the initial analysis is performed). If you expect to retrieve the data frequently, that's clearly not the right storage tier to use. S3 Infrequent Access tier is ~3x the cost, which still supports my point about the relative cost compared to total cost of sequencing. To address some of your other points: the 40TB limit per archive is not a limit on the amount of data your can store in Glacier. And assuming 100MBps throughput is implying you'd use a single node to analyze the data, which does not make sense at this scale.