8 ms·
"2.1 Sufficiency of 16 bits" Someone should go back in time and tell them...
by krispyfi 3y ago
"2.1 Sufficiency of 16 bits"
Someone should go back in time and tell them...
- netcruiser 3y ago"it is certainly possible, albeit uninteresting, to come up with an unreasonable definition of "character" such there are more than 65536 of them". They probably didn't foresee the "unreasonable" and "uninteresting" use of emojis.
- krispyfi 3y agoI was thinking more of Han Unification, which we now realize was a bad idea, and to this day makes it difficult if not impossible to correctly render a block of text that contains characters in more than one CJK language.
- thristian 3y agoTheir justification is: > Unicode aims in the first instance at the characters published in modern text (e.g. in the union of all newspapers and magazines printed in the world in 1988), whose number is far below 2¹⁴ = 16,384. ...and yeah, that's probably true. It wasn't until several years later that they joined forces with the ISO10646 project, which aimed to encode every character ever used by any human writing system, and 16 bits turned out not to be enough.
- jandrese 3y agoAdding all of the dead languages bloated the system very quickly. Shame UCS2 turned out to be such a mistake. Beyond being not just big enough, trading off space for compute power turned out to be the wrong trade, CPUs increased much more rapidly than memory sizes and data especially transmission rates.
- WalterBright 3y agoUCS2 was not a mistake. Unicode forgot its mission.
- otabdeveloper4 3y agoThe idea of fixed-length characters was the mistake. Unicode is fine.
- numpad0 3y ago24- or 32-bit fixed length would be probably fine… The issue is doing a unified codepoint right run into problems of accepting that zh-CN/zh-HK/zh-TW and ja-JP/kr-KR each uses slightly different definitions for every letters in common use, which is not just memory inefficient but also have to do with Taiwanese identity problem. Much easier to keep it to just zh-CN and zh-TW(recently renamed zh-Hans and zh-Hant) with zh-CN completely double defined with ja-JP, and kr-KR dealt as footnotes on ja-JP, which is what they did.
- lifthrasiir 3y ago> The issue is doing a unified codepoint right run into problems of accepting that zh-CN/zh-HK/zh-TW and ja-JP/kr-KR each uses slightly different definitions for every letters in common use, which is not just memory inefficient but also have to do with Taiwanese identity problem. I'm not sure what exactly you wanted to say, but all of six or seven variants of Han characters are substantially different from each other. Glyphwise there are three major clusters due to the different simplification history (of the lack thereof): PRC, Japan and everyone else. (Yes, Korean Hanja is much similar to traditional Chinese.) Even when Taiwan didn't exist at all there would have been three major clusters anyway. Also pedantry corner: ko-KR, not kr-KR. And zh-TW wasn't renamed to zh-Hant; it is valid to say zh-CN-Hant if PRC somehow wants to use traditional Chinese characters.
- numpad0 3y agoMy apologies for ignorance to Korean speakers. What I was trying to say is, to fix Unicode, we must accept that whatever many of those variants don't interchange, and therefore all require its own planes, which I believe had never been the Consortium's stance on the matter. And I think that should solve about half of problems with Unicode by volume.
- WalterBright 3y agoIt's a pity that Unicode forgot its purpose and added: 1. formatting 2. fonts 3. semantic meaning 4. multiple encodings for the same code point 5. you vote for my invented character and I'll vote for your invented character 6. multiple code points for the same glyph 7. an incomprehensible document describing all this nonsense 8. an impossibility to implement it all
- lifthrasiir 3y agoThe only fault of Unicode is 4 (and we have eventually solved it anyway). Everything else is a natural complication from human writing systems.
- WalterBright 3y agoFonts, for example. That is a style, and should be set by the rendering instructions (like CSS) not the code point. Semantics: there's a code point for the letter 'a' and another code point for the mathematical symbol 'a', although the glyphs are identical. Unicode points should not have semantic meaning - semantic meaning comes from the context in which the points are used. You know, like the text in a printed book. And so on, for the other bullet points.
- lifthrasiir 3y agoI have already explained why mathematical alphanumeric symbols exist in the first place [1]. They are an extremely small portion of Unicode; you need way more examples. By the way, if we indeed decide to encode characters according to glyphs, we would have to distinguish a double-storey `a` from a single-storey `a` (among others). They have clearly different glyphs, so why shouldn't we? Think about that. [1] https://news.ycombinator.com/item?id=28294032 https://news.ycombinator.com/item?id=28294032
- WalterBright 3y ago> I have already explained But somehow mathematical books have done just fine without any need for such encodings. Meaning in written language is inferred from the context, not the symbol. This is because meanings are infinite in variety, and symbols are finite. > Think about that I did. That's what fonts are for. Fonts come in endless variety, and so cannot be encoded into Unicode. They are an orthogonal property of rendered text. And so is formatting - it's an orthogonal property.
- wolverine876 3y agoIt's not true if you count current languages of East Asia, such as Chinese. Or does the document address that omission somehow? EDIT: Skimming it, I don't see how the issue is addressed, except in a diagram that says, * Modern-Use Ideographs: 54x256 = 13,824 * Added Ideographs: 64x256 = 16,384 Also it talks about deduplicating identical characters in Chinese, the Japanese alphabets, and Korean (Hangul). More interesting is the full explanation of modern use: Distinction of "moderm-use" characters: Unicode gives higher priority to ensuring utility for the future than to preserving past antiquities. Unicode aims in the first instance at the characters published in modern text (e.g. in the union of all newspapers and magazines printed in the world in 1988), whose number is undoubtedly far below 2^14 = 16,384. Beyond those modern-use characters, all others may be defined to be obsolete or rare; these are better candidates for private-use registration than for congesting the public list of generally-useful Unicodes. It's very interesting, and revealing about the author, to define use as newspaper and magazine publishing.
- lifthrasiir 3y agoUnicode 88 (wrongly) assumed that Han characters in current use are small enough (see p. 3 for specifics). This turned out to be massively incorrect because they have an extremely long tail. Same goes for Hangul. EDIT: > Beyond those modern-use characters, all others may be defined to be obsolete or rare; these are better candidates for private-use registration than for congesting the public list of generally-useful Unicodes. I do believe that the initial estimation was not very off. IICore 2.2 [1] contains 9,810 Han characters across all relevant countries and even assuming that every character should be encoded in two ways (to account for simplified-traditional splits) and IICore had some significant omissions we could have just encoded ~40,000 characters. If everything went well according to their plan. The reality however was that PUA for rare characters was an old concept---pretty much every East Asian character set has one---and proven so ineffective. PUA is by definition not interchangable, but rare characters are still in use and they have to be interchanged anyway. So people avoided them like a plague and argued for bigger character sets instad, East Asian character sets had to be ballooned up, and at the point of Unicode they became so large that a naive vision of Unicode 88 no longer worked. [1] https://en.wikipedia.org/wiki/International_Ideographs_Core https://en.wikipedia.org/wiki/International_Ideographs_Core