3 ms·
This article totally failed to justify why compression algorithms are so important for this application. How big is the human genome, how much can we compress i
by TD-Linux 8y ago
This article totally failed to justify why compression algorithms are so important for this application. How big is the human genome, how much can we compress it by, and why does this matter?
I took a look and zipped FASTA files are about 800MB. Large, sure, but people are streaming Netflix movies far bigger than this every night.
It's really hard for me to see the MPEG-G effort as anything but a money grab for compression patents and licensing.
- mjburgess 8y ago1 Human = 100 GB 500,000 Humans is a small sample of the UK population one company is looking to create. All this is in the article.
- TD-Linux 8y agoThe article doesn't say that 1 human = 100GB anywhere. It does say a reference genome is ~1GB, which sounds similar to a zipped FASTA. 500k humans would make this 400TB of data. Really not that much, especially compared to the $500 million dollars it would take to sequence all those humans at $1,000 each.
- tofof 8y ago> The article doesn't say that 1 human = 100GB anywhere. FTFA: "Typically, a DNA sequencing machine that’s processing the entire genome of a human will generate tens to hundreds of gigabytes of data." Plural tens = 2to3 x 10 = 20-30. Plural hundreds = 2to3 x 100 = 200-300. Median[20,30,200,300] ~= 100 GB.
- DoctorOetker 8y agoHis point still stands that in that 100GB storage costs substantially less than sequencing @ $1000 Why does whom need the original pre-calling measurements? I get that the reads will be randomly distributed. Let's pretend that for a given position in the genome, the number of reads covering it are poisson distributed: i.e. the number of reads is chosen such that with sufficient probability each region is probably covered by one or more reads. This means the peak of the poisson distribution is much higher. So most regions will have over-redundant amount of coverage compared to the aimed for minimum global coverage. So the bulk of the data are actual repetitions of a single underlying sequence in the genome. Why can't these regions covered much more than minimum coverage regions, be called with sufficient certainty to discard the raw measurements? then only the low coverage regions may need raw data. For most research purpouses it seems the raw measurements would be unnecessary? To the extent they are necessary we are admitting that we can't affordably sequence genomes yet (if you want I can sequence yours with my dice).
- epicureanideal 8y agoIt seems to be more like 1.5GB rather than 100 GB, and someone else commented that compressed is around 800 MB. https://bitesizebio.com/8378/how-much-information-is-stored-in-the-human-genome/ https://bitesizebio.com/8378/how-much-information-is-stored-... Even at 1.5 GB, this doesn't seem like an insurmountable problem at the current cost of storage. 750,000 GB = 750 TB, and you can buy an 8 TB hard drive for $150. So for $15,000 you can store 500,000 human genomes. If $15,000 is too expensive for a company, that company is doing something very wrong, when any salary will be many multiples of that.
- evandijk70 8y agoThe raw sequencing data is 100 GB, and that's what needs to be stored, as the algorithms deducing the sequence are not perfect and under continuous development. Moreover, storing a genome in one location on a consumer hard drive is asking for trouble. In practice the genomes are stored backed up in multiple locations and in RAID systems on enterprise hard drives, so your estimate is of by about a factor of 1,000.
- sofaofthedamned 8y agoWhen working on a genome project I saw files up to 150gb, and as I understand it these are compressed csv files.its a lot of storage, especially if you unrl this and index it further for performance reasons. Also the more you compress, the longer the processing time. Some jobs took days, so there was a trade-off. I'm not a genome expert btw, just a devops nerd on a related project.
- mmt 8y ago> 750,000 GB = 750 TB, and you can buy an 8 TB hard drive for $150. So for $15,000 you can store 500,000 human genomes. As others have mentioned [1], the raw sequencing data (still) has to be stored, so 1.5GB is off by about 30x. Even at 6:1 compression, that's still 5x. More importantly, though, the cost of storage, especially at scale, is significantly higher than just the price of bare drives. Backblaze famously complained about this [2] in 2009. Even their particularly cheap, particularly low-performance solution claimed to add almost 45% on top of the cost of the bare drives. A higher-performancd storage system (using, say, standard SAS expander backplanes instead of their non-standard SATA multipliers) has, in my experience, added 70% on top of the bare drives (and that's being frugal, which seems remarkably rare, even among startups). Your $15k is now $128k, minimum. That's just purchase cost, and, while operating cost isn't huge for that little storage, it's not zero. The cost gets much, much worse, if the orgnization can't/won't hire someone like me or won't allow that someone to implement such a frugal DIY assembly (which is the vast majority of organizations today). That leaves "enterprise" storage or cloud, which are both approximately as expensive. I recently heard an estimate of half a million dollars per petabyte for NetApp, which would translate to $1,875,000 here. [1] Well detailed by https://news.ycombinator.com/item?id=17821523 https://news.ycombinator.com/item?id=17821523 [2] famous on HN, anyway. https://www.backblaze.com/blog/petabytes-on-a-budget-how-to-build-cheap-cloud-storage/ https://www.backblaze.com/blog/petabytes-on-a-budget-how-to-...