3 ms·
The article doesn't say that 1 human = 100GB anywhere. It does say a reference genome is ~1GB, which sounds similar to a zipped FASTA. 500k humans would make t
by TD-Linux 8y ago
The article doesn't say that 1 human = 100GB anywhere. It does say a reference genome is ~1GB, which sounds similar to a zipped FASTA.
500k humans would make this 400TB of data. Really not that much, especially compared to the $500 million dollars it would take to sequence all those humans at $1,000 each.
- tofof 8y ago> The article doesn't say that 1 human = 100GB anywhere. FTFA: "Typically, a DNA sequencing machine that’s processing the entire genome of a human will generate tens to hundreds of gigabytes of data." Plural tens = 2to3 x 10 = 20-30. Plural hundreds = 2to3 x 100 = 200-300. Median[20,30,200,300] ~= 100 GB.
- DoctorOetker 8y agoHis point still stands that in that 100GB storage costs substantially less than sequencing @ $1000 Why does whom need the original pre-calling measurements? I get that the reads will be randomly distributed. Let's pretend that for a given position in the genome, the number of reads covering it are poisson distributed: i.e. the number of reads is chosen such that with sufficient probability each region is probably covered by one or more reads. This means the peak of the poisson distribution is much higher. So most regions will have over-redundant amount of coverage compared to the aimed for minimum global coverage. So the bulk of the data are actual repetitions of a single underlying sequence in the genome. Why can't these regions covered much more than minimum coverage regions, be called with sufficient certainty to discard the raw measurements? then only the low coverage regions may need raw data. For most research purpouses it seems the raw measurements would be unnecessary? To the extent they are necessary we are admitting that we can't affordably sequence genomes yet (if you want I can sequence yours with my dice).