5 ms·
As much as I enjoy making fun of JavaScript, it seems more likely to me that the reason JavaScript uses UTF16 internally is the same as for most other languages
by andreasgonewild 9y ago
As much as I enjoy making fun of JavaScript, it seems more likely to me that the reason JavaScript uses UTF16 internally is the same as for most other languages that support Unicode; it's more efficient and convenient to process. UTF8 has variable character boundaries, which means that indexing/counting requires decoding char by char; but it works wonders as an exchange format since it's compact and mostly any language can deal with it.
- mikeash 9y agoAll Unicode encodings require intelligent indexing. JavaScript uses UTF-16 because that (or rather its predecessor UCS-2) was the standard when it was being created. Same reason Java and Apple's Objective-C frameworks use it.
- andreasgonewild 9y agoNot to the same extent as UTF8. It's not like UTF8 invalidated all other encodings; they still fill a purpose and UTF16 seems to still be a popular choice for internal processing, despite the misguided push to use UTF8 for everything.
- mikeash 9y agoWhat's the difference? Both UTF-8 and UTF-16 are variable-length encodings where careless mutation with integer indexes can produce invalid results. UTF-8 is 1-4 bytes per code point, whereas UTF-16 is only 1-2 code units per code point, but that doesn't really make it easier. And proper handling really requires detecting grapheme cluster boundaries, which is the same difficulty regardless of whether you use UTF-8, UTF-16, or UTF-32.
- andreasgonewild 9y agoSo you're looking at up to four times as many chars per code point to take into account, but you still claim it's mostly the same thing. Good luck with the lobbying then, I think we're going to have to agree to disagree on this one.
- mikeash 9y agoCan you describe an operation (other than "count the number of UTF-16 code units") which is easier to code for UTF-16 than UTF-8?
- aurelian15 9y agoWell, even counting the number of code units is straight forward for UTF-8: while (*c) count += ((*(c++) & 0xC0) == 0x80) ? 0 : 1; See https://stackoverflow.com/questions/9356169/utf-8-continuation-bytes https://stackoverflow.com/questions/9356169/utf-8-continuati... for more details.
- mikeash 9y agoThat's counting code points, not code units. A code point is the Unicode "character" number. A code unit is the smallest unit used by an encoding, such as a byte in UTF-8 or two bytes in UTF-16. Counting the number of UTF-8 code units in a UTF-8 string is of course trivial. Counting the number of UTF-16 code units in a UTF-8 strings would take more work. But there's probably no reason you'd want to compute that anyway.
- millstone 9y agoYes, the most important operation: string validation! UTF-16's validation concerns are: 1. Broken surrogate pairs, which is mostly benign. 2. Byte-order confusion. While UTF-8 has: 1. Invalid code points, for example, code points for surrogate halves. 2. Invalid code units, such as 0xFF. 3. Non-shortest forms, where a character may be encoded multiple ways. 4. Representation of NUL, and potential for confusion with APIs that expect null-terminated strings. In practice the UTF-8 issues have caused much more serious vulnerabilities.
- mikeash 9y ago#4 is double-counting, since that's a special case of #3. In any case, these are all concerns for a decoder, but not for an API, which is what we're discussing here. In fact, the original comment I replied to up there was advocating the opposite: UTF-16 internally, and UTF-8 for interchange!
- userbinator 9y agoUTF-8 is 1-4 bytes per code point, whereas UTF-16 is only 1-2 code units per code point, but that doesn't really make it easier As someone who has actually written UTF-8/UTF-16 conversion code, I can immediately tell you which one is far easier to implement: UTF-16. The number of valid cases is basically halved, and the number of error cases in UTF-16 is a fraction of those in UTF-8. Put another way, there are plenty more invalid UTF-8 sequences than invalid UTF-16 sequences.
- johncolanduoni 9y agoThis whole argument is about whether a language's built in string support should use UTF-8. There is no way that you could build a JavaScript engine where you'd even notice the additional complexity of UTF-8 compared to everything else you have to get right (and yes I've had to handle UTF-8 decoding directly before). Plus I'm pretty sure all the major JavaScript engines (V8 for sure) already know how to handle UTF-8 since that's the encoding most scripts come from.
- mikeash 9y agoI agree that UTF-16 is slightly simpler to parse, but compared to everything else you need for Unicode-aware string processing, both are completely trivial. In any case, the discussion here is the appropriate string API, and the relative difficulty of working with those. Exposing UTF-8 versus UTF-16 changes essentially nothing: in both cases you need to either deal with non-integer indexes or deal with integer indexes where not all values are valid. Good string APIs are hard. Most Unicode-aware languages pick one particular encoding and then toss the programmer in the deep end with it. The only language I've seen get it vaguely correct is Swift. (I'm sure there are others, but it's definitely not common.) Swift strings provide multiple views, so you can work with UTF-8, UTF-16, UTF-32, or grapheme clusters, as you need. It doesn't allow using integer indexes directly, so you have to confront the fact that indexing is actually non-trivial. Swift 3 requires using views, and Swift 4 makes the String type itself a sequence of grapheme clusters, which is usually the correct answer to the question of "what unit do you want to work with?"
- jcranmer 9y ago
- YSFEJ4SWJUVU6 9y agoI don't know what angle you are coming from, but by far and large most uses of UTF-16 are entirely due to legacy reasons – APIs and languages made before Unicode extended beyond what fits in 16 bits. Back then, the trade-off made sense, and was popular too, because variable-sized encodings were much more rare. These days UTF-16 really is the worst of both worlds, but we are stuck with what we have.
- aurelian15 9y agoThis. Even if you would represent Unicode strings as 32-bit codewords, it would still require careful processing, e.g. to extract single characters from the string. For example, the single character "Ï" can both be represented as the single codeword "u+CF" and the codeword sequence "u+308 u+49". Due to these complications I favour the old method of just treating strings as byte sequences (and possibly enforcing UTF-8 as encoding, since it is a strict superset of ASCII) and to use specialized functions for Unicode processing.
- andreasgonewild 9y agoBut processing UTF8 is still without a doubt more complex; you're not going to weasel your way out of that fact, no matter how many cases you can think of where they are comparable. Why can't several alternatives be allowed to coexist and complement each other? Why does everything have to be UTF8, or JavaScript, or Rust, or Go or whatever?
- aurelian15 9y agoCan you elaborate? I didn't say that you should use UTF-8 (that's just what I prefer personally), but my point was that you should never make any assumption about a Unicode string without consulting the corresponding Unicode tables and essentially have to treat strings as "opaque sequence of something" anyways. May as well be a byte sequence. Regarding your last point, I'm totally with you (if I understood you correctly). Of course applications should support multiple input/output encodings, but as a programmer you have to decide on some internal representation. That being said, I really don't see how processing UTF-8 is significantly more complex than processing, say, UTF-16. In both cases you need to handle continuation units for the extraction of Unicode code points.
- userbinator 9y agoThat being said, I really don't see how processing UTF-8 is significantly more complex than processing, say, UTF-16. In both cases you need to handle continuation units for the extraction of Unicode code points. UTF-8 has 4 valid cases, one for each length, and many more invalid cases for each length (2-byte sequence missing trail byte, 3-byte sequence missing 1 trail byte, 3-byte sequence missing 2 trail bytes, 4-byte sequence missing 3 trail bytes, 4-byte sequence missing 2 trail bytes, ..., overlongs, UTF-8'd surrogates, overflow, etc.) Differences between implementations' treatment of error cases have lead to some security concerns; see https://hsivonen.fi/broken-utf-8/ https://hsivonen.fi/broken-utf-8/ and discussion at https://news.ycombinator.com/item?id=14451822 https://news.ycombinator.com/item?id=14451822 for an example. UTF-16 has two valid cases (one or two code units) and two error cases (lead surrogate not followed by trail surrogate, lone trail surrogate). It's more like a DBCS, except each code unit is 2 instead of 1 byte.
- pjc50 9y agoAnd Windows ("#define _UNICODE", which actually makes everything use UCS-2)
- pornel 9y agoUTF-16 is a variable-width encoding. Thanks to surrogate pairs some code points take 2 bytes, and some take 4. Even if you use UCS-2 (the 2-byte "UTF-16" infamous for mangling emoji) constant-time indexing of code points still doesn't give you constant time access to what humans would call characters ("grapheme clusters" in Unicode), because of decomposed characters with modifiers and joiners. Languages that use UTF-16 mostly use it only because they're older than UTF-8 (or are tied to a platform older than UTF-8).
- millstone 9y agoUTF-16 is a variable width encoding that is also more convenient to process compared to UTF-8. For example you don't have to deal with invalid code units, non-shortest forms, etc. I think you're right that older languages use UTF-16 and newer ones use UTF-8. But it also seems empirically true that UTF-16 languages do better at grappling with Unicode's subtleties, compared to UTF-8. The temptation of UTF8-is-bsically-C-strings is hard to ignore.
- duskwuff 9y ago> For example you don't have to deal with invalid code units, non-shortest forms, etc. You have to deal with unpaired surrogates, though. And just because that wasn't annoying enough, any UTF-8 sequence which contains an encoded surrogate (which is technically invalid, but not prohibited by most implementations) is impossible to encode as UTF-16.
- duskwuff 9y agoJavascript doesn't specify anything about what character encoding is used "internally". Indeed, as the first comment on the article points out, some implementations (like V8) internally use ASCII to store strings that contain only ASCII characters. What Javascript does is use UTF-16 semantics for Unicode strings. The reason why it does this is simple: when those methods were implemented in the mid-1990s, UTF-16 was largely synonymous with Unicode. No characters beyond U+FFFF were defined until the release of Unicode 3.1 in March 2001.