4 ms·
> I still can’t work out why it wasn’t obvious from the start that UCS-2 would never be enough) Surely certain people did know, but those people weren't in a p
by rectang 5mo ago
> I still can’t work out why it wasn’t obvious from the start that UCS-2 would never be enough)
Surely certain people did know, but those people weren't in a position to do anything about it.
Specifically, there were surely people who knew that because historical Chinese place names, Japanese nicknames, and so on, were not included in the original "Unicode" (it wasn't called UCS-2 yet) it was insufficient for complete expression of Asian languages.
There were also many people who objected to Han unification, which is a different problem.
But all of these objections were discarded because of the overwhelming mandate for a fixed-width encoding. The original "Unicode" was conceived as a "16-bit" initiative. Its 16-bit-ness was an essential aspect of the design and the Unicode Consortium did what they had to do to fit all scripts and characters "in modern use" into 16 bits.
From the Wikipedia article on Han Unification[1]:
> Some of the controversy stems from the fact that the very decision of performing Han unification was made by the initial Unicode Consortium, which at the time was a consortium of North American companies and organizations (most of them in California), but included no East Asian government representatives. The initial design goal was to create a 16-bit standard, and Han unification was therefore a critical step for avoiding tens of thousands of character duplications.
[1] https://en.wikipedia.org/wiki/Han_unification https://en.wikipedia.org/wiki/Han_unification
- jcranmer 5mo agoHan unification predates Unicode by about a decade; most of the early work in Unicode largely consists of copy-pasting the Japanese and Chinese governments' standards for unified CJK ideographs. Indeed, read some of the early histories of Han unification (e.g., https://www.unicode.org/versions/Unicode16.0.0/core-spec/appendix-e/ https://www.unicode.org/versions/Unicode16.0.0/core-spec/app...), and you'll notice that there's a lot of liasoning with East Asian technology groups in East Asian cities going on. I don't think any East Asian government representatives would have actually objected to Han unification! It's also worth noting that the original goal of Unicode wasn't to be able to faithfully represent all text, but rather to faithfully represent existing character sets. Only later do you get the impetus to actually include everything, as people become a lot less tolerant of "computer can't actually represent <X>" scenarios. Note too that a lot of the Han unification criticisms basically fall into the same bucket as, say, Medievalists, who want to preserve certain details of their source texts more faithfully than was the norm for computer systems in the 1980s.
- chrismorgan 5mo agoThere was never an adequate safety margin for anything but immediate (less than five year horizon) use—even at Unicode 1.1 it was more than half full, and they knew they weren’t done. And yet all kinds of major companies put all their eggs in that basket, and then doubled down with the monstrosity that is UTF-16, rather than backing out and going with UTF-8 instead, even though I strongly suspect it would have been easier for everyone involved in most cases, compared to the whole wchar shemozzle. Instead it took Windows twenty-five years to bridge the gap with a UTF-8 codepage (65001) that actually worked.