3 ms·
> With utf-16 you can get O(1) string operations at a cost of doing them on code points rather than graphemes. No you cannot. UTF-16 is a variable-length encod
by msl 8y ago
> With utf-16 you can get O(1) string operations at a cost of doing them on code points rather than graphemes.
No you cannot. UTF-16 is a variable-length encoding (up to two 16-bit code units per code point).
> I think a lot of people totally discount the enormous cost that Latin-1 users are paying for CJK (etc) support they don’t use.
If we limit ourselves to Latin 1, UTF-16 requires 100% more memory than UTF-8. How often do you need to randomly access a long string anyway?
- bradleyjg 8y agoRight, code units not code points. My mistake. Nonetheless that’s what you get back from e.g. java’s charAt.
- ramshorns 8y agoYou can do O(1) string operations in UTF-8 too, if you do them at the code unit level. It's just as wrong, but it's more obvious that it's wrong because it only works for ASCII instead of only working for BMP.
- WorldMaker 8y agoBringing things full circle in this thread, this is absolutely why it is a big deal that emoji have become so popular and that people care about them, and that they are mostly homed in the Astral Plane. Now there's a giant corpus of UTF-16 data people are interacting with daily that absolutely makes it clear that you can't treat UTF-16 like UCS-2, and if you are still doing bad string operations in 2018 you have fewer excuses and more unhappy users ("why is my emoji broken?!").
- bradleyjg 8y agoSame thing I said above except even more so. What’s the most used language that can’t be represented with the BMP? Bengali is one possibility but AFAIK there’s a widely used Arabic form in use as well as the traditional script.