4 ms·
Maybe all Asian scripts was not planned to be included back then? Seems strange they would miscount so grossly otherwise.
by nikbackm 8y ago
Maybe all Asian scripts was not planned to be included back then?
Seems strange they would miscount so grossly otherwise.
- Dylan16807 8y agoIt's all down to CJK. Originally they allocated 21k codepoints to CJK, and if that was accurate then 16 bits would pretty much fit things. But we currently have 88k CJK characters assigned out of possibly more than 100k total. I can't easily find anything about how this went wrong and they got such a small number.
- vorg 8y ago> we currently have 88k CJK characters assigned If you also count the Unihan variations registered in Unicode's ideographic variation database by various Japanese outfits, encoded using the VS16 to VS255 characters after the codepoint they modify, there's another 8k or 9k unique characters assigned.