3 ms·
> […] why limit yourself to 24-bit numbers? For compatibility with UTF-16: o Restricted the range of characters to 0000-10FFFF (the UTF-16 accessi
by throw0101a 15d ago
> […] why limit yourself to 24-bit numbers?
For compatibility with UTF-16:
o Restricted the range of characters to 0000-10FFFF (the UTF-16
accessible range).
* https://datatracker.ietf.org/doc/html/rfc3629#section-12 https://datatracker.ietf.org/doc/html/rfc3629#section-12
* https://en.wikipedia.org/wiki/UTF-16 https://en.wikipedia.org/wiki/UTF-16
The original spec had 31 bits (the UTF-32/UCS-4 range):
* https://datatracker.ietf.org/doc/html/rfc2279 https://datatracker.ietf.org/doc/html/rfc2279
* https://en.wikipedia.org/wiki/UTF-32 https://en.wikipedia.org/wiki/UTF-32
- orangeboats 14d agoWe really ought to deprecate UTF-16 someday. The fact that it pretends to be a fixed-length encoding has caused all sorts of bugs over the years, with many people assuming n(UTF-16 codepoints) == n(characters) which breaks when the string contains non-BMP characters. And also, for personal aesthetic reasons I hate that it limits the Unicode codepoint range to an awkward non-power-of-two number (now there are 0x110000 codepoints in total). UTF-8 and UTF-32's 2^31 feels much more natural.