4 ms·
Agreed. I think maybe the value of the paper is that non-computer scientist biologists will read it and gain awareness. To computer scientists, it's pretty ob
by lph 9y ago
Agreed. I think maybe the value of the paper is that non-computer scientist biologists will read it and gain awareness.
To computer scientists, it's pretty obvious that DNA used as input to a program could be maliciously crafted.
- dboreham 9y agoI'm not sure that is obvious because as a lay-Biologist the input data is just strings of ... um 2-bit ? symbols with no known encoding format (we haven't discovered the "JPEG" of DNA yet). Is quite surprising to me that anyone could screw up parsing that kind of data such that the process could be taken over. Now if DNA contained an encoded turing-complete language that the lab ends up executing... But I think we're some way off from achieving that.
- mattkrause 9y agoThe parsing is actually fine. The issue is that they forced a buffer to overflow--they put N bytes of data into a bucket that was can only hold n bytes. The extra N-n bytes were carefully chosen so that they would do something "interesting" when they spill out of the bucket and into the program's code (there are ways of avoiding this, like the No-eXcecute bit, which I guess were ignored/turned off for this demo). In some ways, this is easier to do with unstructured data. If there was some specific format, the overflowed data would have to fit into that format AND do something malicious.
- dboreham 9y agoThis makes no sense to me as someone who has written plenty of protocol code. The DNA data is essentially random bits. It has no packet size.
- mattkrause 9y agoAssuming I understand the paper correctly.... 1. The DNA is sequenced. 2. Sequence data are compressed with a version of fqzcomp which has been modified to use/overflow a static buffer. 3. Something about this particular DNA sequence overflows a buffer in fqzcomp 4. The overflow contains some kind of exploit. The paper is a little vague about where they inserted the exploit into fzcomp(https://fqzcomp.sourceforge.io/ https://fqzcomp.sourceforge.io/), but it does process the DNA sequence a block at a time. It tracks the frequency distribution of nucleotides so it can model them for compression. I hope the exploit was something really clever, like finding a DNA sequence that puts the encoder in a weird state, but it's also possible they just set BLK_SIZE to 1 or something like that.