3 ms·
I'd really rather go back to the old way when there were lots of competing national encodings for each language, and the actual users of that language could vot
by quant18 17y ago
I'd really rather go back to the old way when there were lots of competing national encodings for each language, and the actual users of that language could vote with their web pages/documents for which one they preferred. Instead we have this one overarching encoding whose subparts were fixed for all time by fiat from committees before being put into practical use, and as you might expect, some of those committees really screwed things up.
For example, some fine upstanding gentlemen decided that in Unicode (and GB-18030), Mongolian ᠣ and ᠤ, which are printed/handwritten exactly the same, shall be two "different letters" U+1823 and U+1824, but the different forms of ᠳ are the "same letter" U+1833. (And of course, there's ᡩ U+1869 which looks like what you want for some forms of U+1833, but you're not supposed to use it because it's "only for Xibe").
The closest analogy I can give in English is an encoding which forced you to use different "a" codepoints for the characters in apple vs. fake because of their different pronunciations, while making a single codepoint for "k" and "ck" and "c" (but only sometimes) because they sound the same. If you ever saw an encoding like that, you'd no doubt say to yourself: "WTF? I'm not using this, I'll stick with ASCII/EBCDIC/Morse code, thank you very much".
- viraptor 17y agoPlease no... It ended up with a simple language like Polish having latin2, cp1852 (or something like that), mazovia, mazovia2, and probably some more homebrew encodings. I can't even imagine what would happen for completely different scripts like Indian. There are problems with unicode - ok, let's resolve them then. I still want to be able to address my email to the real name of person named in language A, living under address of country B, signing the email properly in language C. (where all parts use language-specific characters) Unicode is the first standard which allows me to do that in most cases, so I guess it's a step in the right direction.
- quant18 17y agoI can't even imagine what would happen for completely different scripts like Indian. As far as I know, Ge'ez (for Ethiopian languages) sets the record with 70+ encodings [1] which took all sorts of different approaches. The Unicode design process worked out quite happily for Latin and Cyrillic alphabet users because 1. There was widespread agreement about what is the smallest indivisible unit of the script (thanks to long history of literacy education, decades of typewriter usage, etc.). No one suggested brilliant schemes like encoding "O" as "C" plus a right-concave combining mark ")" or I as "T" plus an underline, for example. 2. Among the hundreds of millions of users of those scripts, there were enough countries which had a reasonable history not just of typewriter usage but also of computer usage, enough time for them to develop various competing encodings whose mistakes Unicode could learn from Inner Mongolian script pretty much presented the worst-case scenario compared to the above criteria: 1. The actual users of the script were a small and poor population with high illiteracy rates and not many computer users; and unlike e.g. Cambodians or Ethiopians, they had no big diaspora population of refugees living in the US or other high-tech countries either (hence no one fluent in English to advocate for them and point out problems in the proposed encodings). 2. As a result of #1, disproportionate amount of discussion surrounding the encoding was generated by scholars whose main aim was digitising quirky classical texts, not everyday people who wanted to write everyday things without the computer making them think of extraneous details they don't think of when they're writing by hand. 3. These scholars can't even agree what is the basic unit of the script (in Russian grad schools, they teach it as an alphabet; in Japanese grad schools, they teach it as a syllabary) [1] http://www.punchdown.org/rvb/papers/EriPaper3C.html http://www.punchdown.org/rvb/papers/EriPaper3C.html