4 ms·
Considering how huge the Unicode standard is, it's more surprising that there are so little (known) errors. CJK characters alone account for thousands upon thou
by cooper12 6y ago
Considering how huge the Unicode standard is, it's more surprising that there are so little (known) errors. CJK characters alone account for thousands upon thousands, and these can sometimes vary by just a single stroke. I suspect the majority of redundancies or suboptimal choices were the result of subsuming so many existing standards though. There might also be plain errors in research but are probably in the more obscure blocks.
See also, ghost kanji: https://www.japantimes.co.jp/life/2018/10/29/language/ghost-kanji-lurk-japanese-lexicon/ https://www.japantimes.co.jp/life/2018/10/29/language/ghost-...
- duskwuff 6y ago> CJK characters alone account for thousands upon thousands The vast majority of CJK characters have systematic names like "CJK UNIFIED IDEOGRAPH-72AC" which aren't really subject to errors in the same way as the more verbose names used for other scripts.
- claudiawerner 6y agoI have to wonder what the systematic names refer to; is it a Chinese/Japanese dictionary ordering in which they were assigned, stroke count (but then, in what order are characters of the same stroke count organized), or some other or arbitrary ordering? As in, why is the 72AC character at 72AC, not at 72AD?
- duskwuff 6y agoThat's a good question. Unfortunately, a lot of old Unicode process documentation isn't available online -- you can see a couple of relevant-sounding documents at [1], but very little of it is available until 1999 or so. There are some pretty clear patterns to the character ordering, though -- if you look closely, you can see big runs of characters which share radicals. For example, characters 5000 through 500F are: 倀 倁 倂 倃 倄 倅 倆 倇 倈 倉 倊 個 倌 倍 倎 倏 all of which have the 亻radical on the left side. It's clearly not arbitrary. [1]: https://www.unicode.org/L2/L1990/Register-1990.html https://www.unicode.org/L2/L1990/Register-1990.html
- cynix 6y ago> all of which have the 亻radical on the left side. Except for 倉 it seems.
- deleted 6y ago[deleted]
- jschwartzi 6y agoIt looks like it's merged under the roof in that character.
- thaumasiotes 6y agoNo, a1369209993 is correct - 亻is the combining form of 人 (the "单人旁"); 倉 uses the more basic 人 form.
- a1369209993 6y agoSupposedly the two strokes at the top that look like a roof are somehow supposed to be the same character as "亻". You could argue that it's not any stupider than claiming that a upside-down vee with a crossbar is the character as a dee with the ascender cropped off, but it also isn't any less stupid, so... <shrug>.
- deleted 6y ago[deleted]
- lifthrasiir 6y ago> There are some pretty clear patterns to the character ordering, though -- if you look closely, you can see big runs of characters which share radicals. That's because those runs were allocated at once (e.g. CJK Unified Ideographs Extension G spanning from U+30000 to U+3134A) and they were systematically ordered by radicals and stroke counts. Note that radicals and stroke counts are fairly arbitrary and can differ among character sources (the Unihan database has a ton of them). While fairly predictable, this ordering is ultimately arbitrary.
- lifthrasiir 6y agoUnicode character names primarily exist for identifying characters. That's why we have the stability for those names after all: unless it is very much misleading (like swapped Lao letters) we would like to stick to them as long as possible. For many scripts we can come up with some set of names that are suitable for identification. Han characters, among others, aren't. One can say that Han "characters" are identified by its shape, but a line between two differently perceived characters is extremely unclear. Some may recognize two characters as same, some not, some would even try to add a stroke or so to differentiate the character. Han characters are thus identified by providing multiple properties for them (the Unihan database), and the character names are just placeholders.
- Muromec 6y ago>There might also be plain errors in research but are probably in the more obscure blocks. The one with Lao letters "LO LING" and "LO LOOT" is actually very blunt error -- names for them are swapped. It's pretty safe to assume that whoever was making the standard does not know this writing system.
- TheRealPomax 6y agoIt really isn't: it's safe to assume they did, because full language mappings are not the job of a single person, and don't get accepted willy-nilly (unlike emoji). However, clerical errors are SUPER EASY, so it's safe to assume someone accidentally pasted something in the wrong database-presented-as-a-spreadsheet close enough to the time of official revision publication, and when that isn't caught, now it's official.
- jschwartzi 6y agoHaving done a Japanese localization by communicating entirely in spreadsheets, this is extremely likely. In fact we had to fix the translations a couple of times because we'd send it off for feedback and get a different translation for the same English.