2 ms·
> UNICODE codepoints where the UTF-8 encoding lengths are the same for lowercase and uppercase. This very same property got me to post this [1], which sent me
by Rendello 22d ago
> UNICODE codepoints where the UTF-8 encoding lengths are the same for lowercase and uppercase.
This very same property got me to post this [1], which sent me down the rabbithole of learning about Unicode in earnest and building my Unicode tool. Which may have an initial release some time this millennia... maybe.
1. https://news.ycombinator.com/item?id=42014045 https://news.ycombinator.com/item?id=42014045
- skrellm 21d agoYeah, UNICODE messed this up, really badly. Sometimes there's an offset (like with Latin, +/- 32), sometimes lowercase and uppercase is interleaved (eg. Latin extended), and sometimes they are in totally different blocks simply because they forgot to add both letter cases at once... (the distance of the codepoints affects UTF-8 encoding difference the most). I've also paid attention to optimize the most common case where both UTF-8 encodings' first bytes are the same (that's 40508 pairs out of 40549). But no escape, it must handle the remaining 41 pairs specially in a slower code path, which would not be needed at all should UNICODE guys did their homework better.
- Rendello 21d agoOne of my favourite ones in the post I linked is the "ff" ligature. It uppercases to "FF", meaning in UTF-8 it goes from one encoded character to two, and from 3 bytes total to 2 bytes total. Lots to read about here for those interested (and that's without getting into `Casefold`, `NFKC_Casefold`, simple vs complex case mappings, the CLDR, etc.): https://www.unicode.org/versions/Unicode17.0.0/core-spec/chapter-5/#G21180 https://www.unicode.org/versions/Unicode17.0.0/core-spec/cha...