3 ms·
So you're looking at up to four times as many chars per code point to take into account, but you still claim it's mostly the same thing. Good luck with the lobb
by andreasgonewild 9y ago
So you're looking at up to four times as many chars per code point to take into account, but you still claim it's mostly the same thing. Good luck with the lobbying then, I think we're going to have to agree to disagree on this one.
- mikeash 9y agoCan you describe an operation (other than "count the number of UTF-16 code units") which is easier to code for UTF-16 than UTF-8?
- aurelian15 9y agoWell, even counting the number of code units is straight forward for UTF-8: while (*c) count += ((*(c++) & 0xC0) == 0x80) ? 0 : 1; See https://stackoverflow.com/questions/9356169/utf-8-continuation-bytes https://stackoverflow.com/questions/9356169/utf-8-continuati... for more details.
- mikeash 9y agoThat's counting code points, not code units. A code point is the Unicode "character" number. A code unit is the smallest unit used by an encoding, such as a byte in UTF-8 or two bytes in UTF-16. Counting the number of UTF-8 code units in a UTF-8 string is of course trivial. Counting the number of UTF-16 code units in a UTF-8 strings would take more work. But there's probably no reason you'd want to compute that anyway.
- millstone 9y agoYes, the most important operation: string validation! UTF-16's validation concerns are: 1. Broken surrogate pairs, which is mostly benign. 2. Byte-order confusion. While UTF-8 has: 1. Invalid code points, for example, code points for surrogate halves. 2. Invalid code units, such as 0xFF. 3. Non-shortest forms, where a character may be encoded multiple ways. 4. Representation of NUL, and potential for confusion with APIs that expect null-terminated strings. In practice the UTF-8 issues have caused much more serious vulnerabilities.
- mikeash 9y ago#4 is double-counting, since that's a special case of #3. In any case, these are all concerns for a decoder, but not for an API, which is what we're discussing here. In fact, the original comment I replied to up there was advocating the opposite: UTF-16 internally, and UTF-8 for interchange!