5 ms·
They're using T and G for a 1, and A and C for a 0; why not double the density and get two bits from each letter? T = 00 G = 01 A = 10 C = 11 for exam
by conanite 14y ago
They're using T and G for a 1, and A and C for a 0; why not double the density and get two bits from each letter?
T = 00
G = 01
A = 10
C = 11
for example.
- seiji 14y agoBase pairs. They don't occur individually.
- lupatus 14y agoIANA DNA expert, but they still seem to be ordered [1]. So a possible scheme to increase data density could be: AT = 00 TA = 01 CG = 10 GC = 11 The trick would be to always correctly identify which is the left and which is the right strand. I don't know if that is possible in practice though. [1] http://en.wikipedia.org/wiki/Base_pair#Examples http://en.wikipedia.org/wiki/Base_pair#Examples
- ajross 14y agoRibosomes seem to manage just fine. :) You just encode a big marker (making sure it's not a palindrome-paired version of itself!) as a header. If you see that, it's a correct order. If not, it's not.
- lupatus 14y agoThis header idea is great because then you only need to keep one strand and can toss the other, potentially quadrupling the amount of data storage (I'm assuming you can keep single strands of DNA stable). [Left strand] A = 00 T = 01 C = 10 G = 11 [Right strand] T = 00 A = 01 G = 10 C = 11 Anyone know these guys at Harvard, b/c this might be a way to put, at most, 2800 terabytes in a gram? (I don't know how long the header sequences would have to be).
- bhickey 14y agoI suspect that George's lab used a sensible encoding scheme. They're fairly sharp. I can't find a copy of the original article, but it's certainly more informative than some science journalism fluff piece.
- sp332 14y agoThey might not have used a "coding" scheme at all, if they were interested in characterizing the frequency and types of errors.
- sliverstorm 14y agoI don't think you can simply toss one of the strands. DNA is so compact because of the way it coils, and you likely lose that if you only have one strand.
- kolinko 14y agoIt's possible to have single stranded DNA, but you'd have problems with error correction. Let's say DNA breaks, or some errors appear in the code. Thanks to the double stranded structure it's "quite easy" to repair the code. Besides that, it's not the density which is a problem right now, but the access speed. The amount of data in DNA is so immense that doubling the density won't give any practical improvements for decades to come - if ever. Having said that, if I'm not mistaken, some viruses are encoded by single stranded DNA & ssRNA. I'm not sure, but the density might be the reason for that.
- algorias 14y agothat's irrelevant since you have 2 strands that are mirror copies of each other. just prefix your data with a single A, if it's read as a T, just invert the bits of the rest of the strand.
- InclinedPlane 14y agoI'm not sure exactly why they decided on that encoding, I suspect there is some technical reason such as error tolerance.
- skosuri 14y agoit's to avoid sequence features which are problematic like high GC content
- InclinedPlane 14y agoAh, interesting. In natural DNA this is achieved by using a more complex encoding scheme (3 base pairs -> one amino acid), combined with the vast majority of DNA not encoding genes directly.
- skosuri 14y agowell, the genetic code has presumably evolved to evolve; it's unclear that it's evolved to preserve information
- skosuri 14y agowe didn't because we wanted to avoid particular sequence features that are difficult to synthesize and sequence. we probably could have gotten away with something like 1.8 bits per base, but we were already doing fine on density, so we thought a 2x hit wouldn't be that bad.
- deleted 14y ago[deleted]
- skosuri 14y agoi should clarify, because of extra sequence and the address barcode, we are technically only 0.6 bits/base