3 ms·
> UTF-8000 is in no way endorsed by or representative of the Unicode Consortium. Not until they decide to expand the emoji range, allocate space for all past a
by sph 12d ago
> UTF-8000 is in no way endorsed by or representative of the Unicode Consortium.
Not until they decide to expand the emoji range, allocate space for all past and future fictional languages, as well as birdsong and dog barks.
Someone at the consortium is rubbing their hands with glee with all the newfound space.
But honestly, cool hack! If you invent a method to encode large numbers into bytes, why limit yourself to 24-bit numbers?
- flohofwoe 12d ago> ...24-bit numbers? Technically current UTF-8 only goes up to 21 bits (that's the current UNICODE range), for the encoding itself that is an arbitrary limit though, with the 'single lead byte' method of traditional UTF-8 it could go up to 36 bits "payload".
- throw0101a 12d ago> […] why limit yourself to 24-bit numbers? For compatibility with UTF-16: o Restricted the range of characters to 0000-10FFFF (the UTF-16 accessible range). * https://datatracker.ietf.org/doc/html/rfc3629#section-12 https://datatracker.ietf.org/doc/html/rfc3629#section-12 * https://en.wikipedia.org/wiki/UTF-16 https://en.wikipedia.org/wiki/UTF-16 The original spec had 31 bits (the UTF-32/UCS-4 range): * https://datatracker.ietf.org/doc/html/rfc2279 https://datatracker.ietf.org/doc/html/rfc2279 * https://en.wikipedia.org/wiki/UTF-32 https://en.wikipedia.org/wiki/UTF-32
- orangeboats 11d agoWe really ought to deprecate UTF-16 someday. The fact that it pretends to be a fixed-length encoding has caused all sorts of bugs over the years, with many people assuming n(UTF-16 codepoints) == n(characters) which breaks when the string contains non-BMP characters. And also, for personal aesthetic reasons I hate that it limits the Unicode codepoint range to an awkward non-power-of-two number (now there are 0x110000 codepoints in total). UTF-8 and UTF-32's 2^31 feels much more natural.
- sharktheone 12d agoI think that wouldn't change much. They would just make use of more grapheme clusters. For emojies they already make heavy use of the Zero-Width-Joiner. So a woman firefighter is the woman emoji + ZWJ + fire engine. Sure the UTF-8000 approach is much better encoding size wise.
- mitxela 12d agoI wonder how they're going to encode a female fire engine in the future.
- Dylan16807 12d agoNo worries, that would use female sign, not woman.
- sharktheone 12d agoI don't understand why you are saying this. That just seems a bit inappropriate
- mitxela 12d agoIt's a joke based on the construction of emojis? Female plus fire engine obviously denotes a female fire engine but has apparently been repurposed as a female firefighter instead
- TeMPOraL 12d agoI predict eventual convergence between UTF-whatever and most popular tokenizer for whatever LLM escapes to become world-ruling AGI. I mean, if someone's seriously going to try encoding birdsong and dog barks, at this point they're basically reinventing tokens for multi-modal language models.