4 ms·
Thanks for that background. I was thinking that the (for lack of a better word) politics of CJKV and Han unification likely had a lot to do with it. To my eye
by SloopJon 7y ago
Thanks for that background. I was thinking that the (for lack of a better word) politics of CJKV and Han unification likely had a lot to do with it.
To my eyes, the letter "A" in English (Latin), Greek, and Russian (Cyrillic) looks identical. Why does it need three separate code points? Maybe it gets tricky when the relationship between uppercase and lowercase diverges. In fact, that does appear to be one of the arguments in this technical note:
https://www.unicode.org/notes/tn26/ https://www.unicode.org/notes/tn26/
- amaccuish 7y agoOne thing I can think of is that when I write in Russian, and I put text into "italics", it goes into the cursive script. So I guess the letter "a" is different at least from the roman letter.
- deleted 7y ago[deleted]
- legulere 7y agoYou could also see black letter and antique as two different scripts. In fact foreign words were written in antiqua. In Unicode however they’re put together (except for mathematics).
- gumby 7y agoI remember the politics of this well. It was all about legacy encodings (similar to the UCS-2 issue in the article). There were many encodings (“code pages” in MS land) for each language or character set, even Latin-derived ones. Just getting a single one for Cyrillic or Greek was a hurdle, with the most common in use as the “default”. True character unification would have been great, but was too much of a hurdle because you couldn’t just use an 8-bit character space plus a single offset. This was in the days of 16-bit machines, remember. This legacy support is why there are precomposed characters. It’s not quite the same for Han unification (people using those alphabets were already used to complex gyration was ao e). In that case the linguists, historians and other scientists were in agreement but they neglected politics. In Latin unification nobody cares that J sounds different in every language, but because Hanzi embody semantics, Japan may have a loancharacter because it once sounded the same as a word in use at the time OR because it had the same meaning as a word in use at the time so the sense of “unified” can be contentious, even though they are simply code points.
- mrighele 7y agoI don't think uppercase/lowercase is an issue and in fact it is something that already happens. For example in most languages lowercase "i" has the letter "I" as uppercase version, but not in Turkish, where it gets converted to "İ" (with the dot). In Turkish "I" is the uppercase version of "ı" (without the dot). In other words the same letter can be "uppercased" to a different letter depending on the current language
- simias 7y agoI think the most salient point in your link is: >Even more significantly, from the point of view of the problem of character encoding for digital textual representation in information technology, the preexisting identification of Latin, Greek, and Cyrillic as distinct scripts was carried over into character encoding, from the very earliest instances of such encodings. Once ASCII and EBCDIC were expanded to start incorporating Greek or Cyrillic letters, all significant instances of such encodings included a basic Latin (ASCII or otherwise) set and a full set of letters for Greek or a full set of letters for Cyrillic. Precedent for the purposes of character encoding was clearly established by those early 8-bit charsets. That's true, although it reminded me of a cool charset: the Russian KOI8-R encoding[1]. This encoding was created so that if you stripped the high bit (presumably on a system unable to handle russian properly) you ended up with semi-legible latinized Russian. Quoting wikipedia: >For instance, "Русский Текст" in KOI8-R becomes rUSSKIJ tEKST ("Russian Text") if the 8th bit is stripped; attempting to interpret the ASCII string rUSSKIJ tEKST as KOI7 yields "РУССКИЙ ТЕКСТ". KOI8 was based on Russian Morse code, which was created from Latin Morse code based on sound similarities, and which has the same connection to the Latin Morse codes for A-Z as KOI8 has with ASCII. [1] https://en.wikipedia.org/wiki/KOI8-R https://en.wikipedia.org/wiki/KOI8-R
- yencabulator 7y agoWow, they made the 8th bit essentially be the case bit, and losing+recreating it mostly just made everything uppercase.
- bonzini 7y agoNo, the case bit is still bit 5. Hiwei, unlike ASCII, KOI7/KOI8's Cyrillic alphabet sets bit 5 to 1 for uppercase characters. This way, even though Cyrillic text remains legible, you also have a clue that the encoding is KOI.
- tempguy9999 7y agoBloody hell, that is awesome and I don't often say that.
- wruza 7y agoThat would mix words from different alphabets when sorted in a lexicographical order (and/or break this order, respectively). You could get used to it, but traditionally it is a big nope. Let fonts erase the difference, but we want to know if a letter is greek, latin, cyrillic or something else. Sorting complex charsets may require some sort of collation, of course, but in simple cases the status quo is highly usable without one.