7 ms·
We still have about 85% of codepoint space unused. Hopefully, by the time it becomes a problem, UTF-16 will be long dead
by delamon 8d ago
We still have about 85% of codepoint space unused. Hopefully, by the time it becomes a problem, UTF-16 will be long dead
- colejohnson66 8d agoBut by then, the 4-byte limit of UTF-8 will itself have ossified. Even today, reverting back to the 6-byte limit is nigh impossible.
- Razengan 8d agoBy then we will have quaternary quantum computers and FTL circuits where the information appears request it before you
- nasso_dev 8d agoi hope so too, but UTF-16 being used by languages such as java and javascript makes me fear it might be here to stay.... i hope im wrong
- deleted 8d ago[deleted]
- hnlmorg 8d agoThe number of glyphs available by adding additional bytes drops exponentially because each subsequent byte has one less bit available. So I think if we ever were in a situation where > 1 million code points isn’t enough, then we should look at an entirely new way to serialise those code points.
- delamon 8d agoI don't quite get it. 5-byte utf-8 encoding gets extra 5 bits compared to 4 byte, and 6-byte gets extra 10 bits. If you were thinking about bits in leading byte, then yes, you are losing one bit for every extra trailing byte, but you also get 6 bits from it. So adding a byte gives you extra 5 bits.
- hnlmorg 8d agoYeah, you’re right. I might have attempted to do mental arithmetic before coffee…
- 7bit 8d agoUtf-16 is famously used by Windows for everything important as well.
- adornKey 8d agoAnd UTF-8 isn't even fully compatible with windows UTF-16 - UTF8 can't encode a lot of truncated windows UTF-16 filenames.. You need WTF-8 for that. https://artoria2e5.github.io/XB18030/ https://artoria2e5.github.io/XB18030/ It seems when designing Unicode most energy went into emoji. And there was nothing left for fancy things like fixed-length string buffers. The only explaination why UTF8 Buffers aren't compatible with UTF16 Buffers... is a really strong emoji...
- colejohnson66 8d agoThankfully, this is changing. Win32's 'A' ANSI/ASCII APIs now support UTF-8 if your app declares such a wish. https://stackoverflow.com/a/69181417/1350209 https://stackoverflow.com/a/69181417/1350209
- flohofwoe 7d agoThe internal string encoding of a programming language doesn't matter as long as it supports UTF-8 at the boundaries. E.g. the text encoding standard on the web is clearly UTF-8, even though JS strings may be internally stored as UTF-16 (or any other encoding). Same on macOS/iOS btw: AFAIK NSString is internally UTF-16, but I've never seen a UTF-16 text file on macOS, it's all UTF-8 (unless the file originated on Windows of course).
- account42 7d agoThe internal string encoding and its limitations does leak into the APIs.