5 ms·
Building an invalid string in Python 2.x
- drunkpotato 13y agoThat's really cool! Character encoding issues is something we wrestle with all the time, and it is surprisingly hard to reason about all the ways supposedly "string" data are handled in the course of a typical workflow. I cringe; I hadn't even considered bugs in the encoding and decoding process itself.
- hannibal5 13y agoHow many open source libraries and programs there exist that can actually work correctly with full Unicode with all kinks involved? I know none. Emacs seems to do best job, but I have not investigated it much. I know one proprietary library for Unicode that can be used for search (from multiple different sources of UTF strings), indexing etc. and claims to support full Unicode that deals with all things involved, including directionality, surrogates, control chars etc. Their internal representation of stings for string processing is vector of displayed characters objects (not code points). Using UTF-* encoding as internal representation works only for simple string processing for subset of Unicode.
- Beltiras 13y agoI'm not seeing any meaningful exploits coming from this. You can maybe send a request that will fail but I can't see any sort of injection taking place.
- dlitz 13y agoIt managed to insert invalid Unicode into a SQLite database, causing a subsequent SELECT to fail. That's at least a DoS attack.
- icebraining 13y agoYes, but only if you're decoding user input as UTF-7, which would be insane.
- doki_pen 13y agoWhat if you were scraping a webpage and it reported its encoding as UTF-7?
- mattdeboard 13y ago...which is, in fact, exactly how this bug was exposed.
- gsnedders 13y agoModern browsers don't support UTF-7 any more after a number of XSS attacks relying on inserting UTF-7 encoded script elements which then cause the document to be sniffed as UTF-7. The only place UTF-7 is still widely used is in email clients.
- est 13y ago> only if you're decoding user input as UTF-7 Hmm, may I ask what makes utf8 won't produce U+DEADBEEF? Or something remotely like that? Edit: '\xfb\x9b\xbb\xaf'.decode('utf8') UnicodeDecodeError: 'utf8' codec can't decode byte 0xfb in position 0: invalid start byte
- gsnedders 13y agoIt's a bug in the UTF-7 decoder that yields an invalid codepoint (outwith of the Unicode codespace) and isn't checked anywhere.
- rspeer 13y agoWhen UTF-8 was first defined, they didn't know how big the Unicode range was going to be, so they defined it as a 1-6 byte encoding that could encode any 32-bit codepoint. When Unicode was deemed to end at U+10FFFF (because that's the largest value that UTF-16 can encode), UTF-8 was revised to be a 1-4 byte encoding that ends in the same place. Python clearly implements UTF-8 in a way that uses at most four bytes per codepoint (why support five and six byte sequences if they'll never be used?). I think what we're seeing in '\xfb\x9b\xbb\xaf' is four bytes out of a six byte sequence.
- excitom 13y agoWhen working as an AIX kernel program in 1985, I set registers to a unique value so it would be easy spot code that tried to use an uninitialized value. My choice: 0xdeadbeef. Good to see that constant is still in use.
- estebank 13y agoWhenever I find myself having to change a MAC address, I end up using DEADBEEFCAFE. I hope I never forget about changing them back and end up having to debug two different machines with the same MAC (which has actually happened to me in the wild, with two machines coming out of factory with the same MAC, talk about bad luck and shitty quality control).
- nknighthb 13y agoI've seen duplicate MACs twice in the last few years, on two different lines of embedded/consumer electronics boards from two different factories. There was a kind of Abbot & Costello routine that went on the first time, when a Taiwanese colleague with limited English reported the problem to me.
- pdpi 13y agoThere's a fair few of those out in the wild. The magic number that identifies a java .class file is 0xCAFEBABE.
- arethuza 13y agoHere's a list: http://en.wikipedia.org/wiki/Hexspeak http://en.wikipedia.org/wiki/Hexspeak
- dahart 13y agoHere's another list. :P < /usr/share/dict/words perl -ne 'print if m/^[abcdefilzsbtgo]*$/ && m/^........$/;' | perl -ne 'print if !(m/i/ && m/l/);' | tr 'ilzstgo' '1125790' | tr '[:lower:]' '[:upper:]' | perl -ne 'print "0x$_"' | column
- brokentone 13y agoThis reminds me of Godel's incompleteness theorem - which I'll poorly present as: Any system that is sufficiently complex and complete will contain legal assertions that will disprove or destroy the system. (Those that do not are not complete). http://en.wikipedia.org/wiki/G%C3%B6del's_incompleteness_theorems http://en.wikipedia.org/wiki/G%C3%B6del's_incompleteness_the... http://www.amazon.com/G%C3%B6del-Escher-Bach-Eternal-Golden/dp/0465026567 http://www.amazon.com/G%C3%B6del-Escher-Bach-Eternal-Golden/...
- kbd 13y ago> ... will contain legal assertions... Well, the comments state that to make this happen he had to "[exploit] a bug in the UTF-7 decoder". So, not legal assertions.
- brokentone 13y agoIt parses, thus it is legal.
- jerf 13y agoNeither throwing an exception nor having a perfectly-deterministic buggy behavior is what Godel was referring to. This shouldn't remind you of anything related to the incompleteness theorem, because it's completely unrelated.
- anaphor 13y agoNot completely unrelated: https://en.wikipedia.org/wiki/G%C3%B6del_numbering https://en.wikipedia.org/wiki/G%C3%B6del_numbering but what he/she was talking about is unrelated.
- sp332 13y agoI don't want to be condescending, but that isn't what the theorem says. (I'm not even sure it's true.) Incompleteness means there is a true statement, that cannot be proved true inside the system.
- twoodfin 13y agoI'm not a Python geek, but I found the C implementation for unicode strings in CPython really interesting code reading: http://hg.python.org/cpython/file/tip/Objects/unicodeobject.c http://hg.python.org/cpython/file/tip/Objects/unicodeobject.... CPython supports several internal representations from one to four bytes per character to optimize for space and performance. There's also a nifty sort of Bloom filter for quick discrimination of strings that might contain characters of interest.
- gsnedders 13y agoThat's new in Python 3.3, which doesn't fall foul of that bug. This is PEP 393 (http://www.python.org/dev/peps/pep-0393/ http://www.python.org/dev/peps/pep-0393/) if you want more reading about it.
- mzs 13y agoHere's the bug (utf-7 decoder) so you don't have to login to github: http://bugs.python.org/issue19279 http://bugs.python.org/issue19279