3 ms·
But processing UTF8 is still without a doubt more complex; you're not going to weasel your way out of that fact, no matter how many cases you can think of where
by andreasgonewild 9y ago
But processing UTF8 is still without a doubt more complex; you're not going to weasel your way out of that fact, no matter how many cases you can think of where they are comparable. Why can't several alternatives be allowed to coexist and complement each other? Why does everything have to be UTF8, or JavaScript, or Rust, or Go or whatever?
- aurelian15 9y agoCan you elaborate? I didn't say that you should use UTF-8 (that's just what I prefer personally), but my point was that you should never make any assumption about a Unicode string without consulting the corresponding Unicode tables and essentially have to treat strings as "opaque sequence of something" anyways. May as well be a byte sequence. Regarding your last point, I'm totally with you (if I understood you correctly). Of course applications should support multiple input/output encodings, but as a programmer you have to decide on some internal representation. That being said, I really don't see how processing UTF-8 is significantly more complex than processing, say, UTF-16. In both cases you need to handle continuation units for the extraction of Unicode code points.
- userbinator 9y agoThat being said, I really don't see how processing UTF-8 is significantly more complex than processing, say, UTF-16. In both cases you need to handle continuation units for the extraction of Unicode code points. UTF-8 has 4 valid cases, one for each length, and many more invalid cases for each length (2-byte sequence missing trail byte, 3-byte sequence missing 1 trail byte, 3-byte sequence missing 2 trail bytes, 4-byte sequence missing 3 trail bytes, 4-byte sequence missing 2 trail bytes, ..., overlongs, UTF-8'd surrogates, overflow, etc.) Differences between implementations' treatment of error cases have lead to some security concerns; see https://hsivonen.fi/broken-utf-8/ https://hsivonen.fi/broken-utf-8/ and discussion at https://news.ycombinator.com/item?id=14451822 https://news.ycombinator.com/item?id=14451822 for an example. UTF-16 has two valid cases (one or two code units) and two error cases (lead surrogate not followed by trail surrogate, lone trail surrogate). It's more like a DBCS, except each code unit is 2 instead of 1 byte.
- johncolanduoni 9y agoProcessing UTF8 is very barely more complex, and mostly for whoever writes your language's String implementation. Once you have to worry about whether a code point is more than one unit, there's not much difference between 1-2 units and 1-4 units. Ideally you should be working with grapheme clusters anyway since that's the only way to have a hope of not butchering things (even non-normalized Latin text may contain multi-codepoint letters), but most languages don't give you a good way to deal with them so that's difficult in practice. With UTF-8 you'll at least have a shot at noticing that you're not handling multi-unit codepoints well, while with UTF-16 you won't notice unless you test Chinese or a more off the beaten path language.
- mikeash 9y agoI think that last part is important, and it's a big reason why I dislike UTF-16. It's much harder to notice an inadequate implementation when you're using UTF-16. Note that most of Chinese is still in the BMP so even then it will probably work fine. You'll get failures on more obscure Chinese characters, most emoji, and really obscure scripts like linear B and cuneiform.
- millstone 9y agoUTF-8 makes it more obvious that you're mishandling multi-unit code points, BUT it introduces its own issues, specifically invalid code units and non-shortest forms. These issues represent security vulnerabilities which have been successfully exploited, and are impossible-by-design with UTF-16.