4 ms·
The current scheme is extensible to 7x6=42 bits (which will probably never be needed). The advantage of the current scheme is that when you read the first byte
by stkdump 5y ago
The current scheme is extensible to 7x6=42 bits (which will probably never be needed). The advantage of the current scheme is that when you read the first byte you know how long the code point is in memory and you have less branching dependencies, i.e. better performance.
EDIT: another huge advantage is that lexicographical comparison/sorting is trivial (usually the ascii version of the code can be reused without modification).
- coldpie 5y ago> The current scheme is extensible to 7x6=42 bits (which will probably never be needed). I have printed this out and inserted it into my safe deposit box, so my children's children's children can take it out and have a laugh.
- account42 5y agoYou can actually extend it indefinitely, 42 bits is just the farthest you can go by trividally extending the bit patterns with the existing logic.
- stkdump 5y agoUnicode 13 uses 143859 code points. 21 bits can encode more than 10x of that. And there are not that many languages left to add. We are already debating elvish. I am not sure what would warrant extending the coding scheme if we are already on the level of fictional languages today. Also note that many things like flags or emoji skin colors, etc. are now done via grapheme clusters, which conserves code points.
- coldpie 5y agoI have stapled your reply to the original printout for extra hilarity! :)
- lkuty 5y agolike "A Branchless UTF-8 Decoder" at https://nullprogram.com/blog/2017/10/06/ https://nullprogram.com/blog/2017/10/06/