4 ms·
Some languages that had the misfortune to be developed while people thought UCS-2 would be enough define string length as number of UTF-16 code units, which is
by wolf550e 5y ago
Some languages that had the misfortune to be developed while people thought UCS-2 would be enough define string length as number of UTF-16 code units, which is not even on the list because it's not a useful number.
- garmaine 5y agoYeah, and most languages/libraries which predate Unicode define string length as the number of bytes (#1). That's probably the most common interpretation, actually. But most new implementations count code point, in my experience. I believe this is also the Unicode recommendation when context doesn't determine a different algorithm to read. Been a while since I read those recommendations though, so I could be wrong.
- zinekeller 5y ago> I believe this is also the Unicode recommendation when context doesn't determine a different algorithm to read. Except that emojis are universally two "characters", even those that are encoded as several codepoints. Also, non-composite Korean jamo versus composited jamo.
- garmaine 5y agoLike this: “:)” ? Japanese kana also count as two characters. Which they largely are when romanized, on average. Korean isn’t identical but the information density is approximately the same. Good enough to approximate as such and have a consistent rule.
- chrismorgan 5y agoMy experience is that most new stuff is deliberately choosing and even exposing UTF-8 these days, counting in code units (which is there equivalent to bytes). I’d say that not much counts in code points, and in fact that it’s unequivocally bad to do so, posing a fairly significant compatibility and security hazard (it’s a smuggling vector) if handled even slightly incautiously. Python is the only thing I can think of that actually does count in code points: because its strings are not even potentially ill-formed Unicode strings, but sequences of Unicode code points ('\ud83d\ude41' and '\U001f641' are both allowed, but different—though concerning the security hazard, it protects you in most places, by making encoding explicit and having UTF codecs decline to encode surrogates). String representation, both internal and public, is a thing Python 3 royally blundered on, and although they fixed as much as they could somewhere around 3.4, they can’t fix most of it without breaking compatibility. JavaScript is a fair example of one that mostly counts in UTF-16 code units instead (though strings can be malformed Unicode, containing unmatched pairs). Take a U+1F641 and you get '\ud83d\ude41', and most techniques of looking at the string will look at it by UTF-16 code units—but I’d maintain that it’s incorrect to call it code points, because U+D83D has no distinct identity in the string like it can in Python, and other techniques of looking at the string will prevent you from seeing such surrogates. It would have been better for Python to have real Unicode strings (that is, exclude the surrogate range) and counted in scalar values instead. Better still to have gone all in on UTF-8 rather than optimising for code point access which costs you a lot of performance in the general case while speeding up something that roughly no one should be doing anyway. (I firmly believe that UTF-16 is the worst thing to ever happen to Unicode. Surrogates are a menace that were introduced for what hindsight shows extremely clearly were bad reasons, and which I think should have been fairly obviously bad reasons even when they standardised it, though UTF-8 did come about two years too late when you consider development pipelines.)
- int_19h 5y agoIt's not just about optimizing Unicode access. It's also because system libraries on many common platforms (Windows, macOS) use UTF-16, so if you always store in UTF-8 internally, you have to convert back and forth every time you cross that boundary.
- chrismorgan 5y agoMost languages that seem to use UTF-16 code units are actually mixed-representation ASCII/UTF-16, because UTF-16’s memory cost is too high. I think all major browsers and JavaScript engines are (though the Servo project has shown UTF-16 isn’t necessary, coining and using WTF-8), and Swift was until they migrated to pure UTF-8 in 5.0 (though whether it was from the first release, I don’t know—Swift has significantly changed its string representation several times; https://www.swift.org/blog/utf8-string/ https://www.swift.org/blog/utf8-string/ gives details and figures). Python is mixed-representation ASCII/UTF-16/UTF-32! So in practice, a very significant fraction of the seemingly-UTF–16 places were already needing to allocate and reencode for UTF-16 API calls. UTF-16 is definitely on the way out, a legacy matter to be avoided in anything new. I can’t comment on macOS UTF-16ness, but if you’re targeting recent Windows (version 1903 onwards for best results, I think) you can use UTF-8 everywhere: Microsoft has backed away from the UTF-16 -W functions, and now actively recommends using code page 65001 (UTF-8) and the -A functions <https://docs.microsoft.com/en-us/windows/apps/design/globalizing/use-utf8-code-page https://docs.microsoft.com/en-us/windows/apps/design/globali...>—two full decades after I think they should have done it, but at least they’re doing it now. Not sure how much programming language library code may have migrated yet, since in most cases the -W paths may still be needed for older platforms (I’m not sure at what point code page 65001 was dependable; I know it was pretty terrible in Command Prompt in Windows 7, but I’m not sure what was at fault there, whether cmd.exe, conhost.exe or the Console API Kernel32 functions). Remember also that just about everything outside your programming language and some system and GUI libraries will be using ASCII or UTF-8, including almost all network or disk I/O, so if you use UTF-16 internally you may well need to do at least as much reencoding to UTF-8 as you would have the other way round. Certainly it varies by your use case, but the current consensus that I’ve seen is very strongly in favour of using UTF-8 internally instead of mixed representations, and fairly strongly in favour of using UTF-8 instead of UTF-16 as the internal representation, even if you’ll be interacting with lots of UTF-16 stuff.