4 ms·
Eh, there's been far worse and their justification for it kinda-almost makes sense. They were trying to avoid 00 in the strings at all costs. It's simple enough
by TkTech 11y ago
Eh, there's been far worse and their justification for it kinda-almost makes sense. They were trying to avoid 00 in the strings at all costs. It's simple enough to encode and decode: https://github.com/TkTech/Jawa/blob/master/jawa/util/utf.py#L12 https://github.com/TkTech/Jawa/blob/master/jawa/util/utf.py#...
- KMag 11y agoAvoiding nulls wasn't the messed up part. Encoding UTF-16 surrogate pairs separately as pseudo-UTF-8 was the weird part. Some codepoints that would be 3 bytes in UTF-8 or 4 bytes in UTF-16 wind up as 6 bytes in Java's pseudo-UTF-8.
- TkTech 11y agoThe encoding for the surrogate pairs is a spec (not a standard) called CESU-8[1]. The modified UTF-8 in the ClassFiles is really just CESU-8 with an exception for U+0000. 1: http://www.unicode.org/reports/tr26/ http://www.unicode.org/reports/tr26/
- Dylan16807 11y agoAnd that spec is basically "We screwed up the UTF-16 support and wrote a UCS-2 to UTF-8 converter instead. Um, well, it works, don't touch it."