4 ms·
> It would have been expensive, but all characters should have been fixed size 64bit values You're making the same mistake that numerous people made before you
by usrnm 5mo ago
> It would have been expensive, but all characters should have been fixed size 64bit values
You're making the same mistake that numerous people made before you: thinking that it's as simple as using arrays of large enough numbers. First they thought that two bytes per symbol would be enough, then four. Spoiler alert: it wasn't. And eight won't work either.
- bombcar 5mo agoUnicodeV6 - 128 bits per character!
- 201984 5mo agoWhy wouldn't 8 be enough? Surely 18,446,744,070,000,001,024 characters is enough for every writing system in the world.
- usrnm 5mo agoBecause that's not how Unicode works. It's not simply a table mapping numbers to all possible symbols. Unicode is full of special codepoints that have no meaning on their own, they serve as modifiers to other symbols and a single visible symbol can be formed by an arbitrary (in theory) long combimation of such codepoints. It doesn't matter how you encode it, it simply doesn't work as "codepoint -> symbol" and indexing in a unicode string is never O(1) and cannot be made O(1). Could we use a simple table approach? Maybe. But it wouldn't be Unicode
- jandrese 5mo agoI actually wonder if the combinatoral explosion of attempting to enumerate every possible character combination would exceed 2^64 bits. My intuition is that it might, and also such a system would be unworkably unwieldy. The size of the spec document would also suffer from the combinatoral explosion. Imagine a system that tries to encode a unique entry for every possible Zalgo character. Also, literally nobody wants to use 64 bit values to encode ASCII values. Even in our world of insanely large storage that would be breathtakingly wasteful.
- BobbyTables2 5mo agoAgreed, but it will take many generations for people to see characters in textual strings mainly as “code” instead of “data”.