9 ms·
These aren't mistakes of C, these are mistakes made hundreds and hundreds of years ago in the design of various language writing systems. English gets a lot ri
by smcameron 3y ago
These aren't mistakes of C, these are mistakes made hundreds and hundreds of years ago in the design of various language writing systems. English gets a lot right with its writing system, mostly in its choice of a small alphabet, a small set of punctuation marks, and an almost complete absence of diacritical marks.
- mikepurvis 3y agoMaybe, but we also have to understand that by an accident of history, computers and the internet were originated by English-speaking people, so it makes sense that the defaults (ASCII and so on) were geared toward the specific needs of that language— it makes sense that everything not-English is going to have some degree of feeling "bolted on" to an underlying framework built for English. If French or Greek or Russian or Japanese had happened to be the lingua franca of the Internet, the features those writing systems require would be first class, and the implementation of English would be the one that felt like a bit more of a kludge. Imagine if computers were designed around unicase Arabic languages [1], and then years later there was a "Han Unificiation"-type summit where it was decided that capitalization was to be handled with extra metadata, since an upper- lower-case latin letter are really just the same thing and why bother wasting double the codepoints on that. (I don't have a link handy, but I believe I read somewhere that you can see some of this effect in NES-era RPGs from Japan— that some of the narrative had to be simplified down for the English translations because the text boxes were originally built for a small number of complex characters rather than the larger number of simpler ones that would be required to capture the same meaning.) [1]: https://en.wikipedia.org/wiki/Unicase https://en.wikipedia.org/wiki/Unicase
- dahfizz 3y agoThe English language evolved alongside the technology. English used to be cursive-only, like something like Arabic. But the language evolved to fit into a modern world, and I think other written languages would benefit from a simplified "print" version as well.
- dontlaugh 3y agoAlmost every language already has a print version, newspapers and books aren't a new thing. Computers merely got it wrong several times, mostly for silly reasons.
- serentty 3y agoThe narrative that Han unification is this thing imposed by ignorant westerners on East Asian computer users is simply not true. The criteria for which characters were unified was mostly based on the criteria used in legacy East Asian encodings, which already had to deal with the question of what counts as the same character and what does not. Unicode has round trip compatibility with the old encodings, too, without the use of any of that extra metadata, which is only used for incredibly minor character variations which are pretty much never semantically meaningful the way capitalization is. A human copyist might change one for another in copying a text by hand, just as when copying English, you do not consider it semantically meaningful whether lowercase “a” is drawn as a circle and a line, or a circle and a hook.
- mikepurvis 3y agoIs there a longer-form piece that would be helpful for me to read giving more of this perspective and showing receipts on it? As someone with little knowledge outside the Wikipedia article, it seems to me most damning that Shift JIS is still so popular in Japan.
- kps 3y agoUnicode 1.0 Volume 2 Chapter 2 — https://www.unicode.org/versions/Unicode1.0.0/V2ch02.pdf https://www.unicode.org/versions/Unicode1.0.0/V2ch02.pdf
- serentty 3y agoI wish I could think of a good long form article about this. The best I can immediately think of is a pretty informative FAQ about Japanese encoding. https://www.sljfaq.org/afaq/encodings.html https://www.sljfaq.org/afaq/encodings.html What I will point out though, that you can easily verify with a simple search, is that Shift-JIS can be represented losslessly in Unicode, and that this has been true for as long as Unicode CJK has existed. The continued popularity of Shift-JIS is worth noting, but it is also important to note that its continued use is not a stable thing, and has been declining for two decades. The most popular websites in Japan no longer use it, and among smaller websites, the percentage that use it gets smaller each year. Secondly, there is absolutely no benefit to the user in terms of that can be encoded and distinguished, because it will get converted into Unicode at some stage in the pipeline anyway. Any modern text renderer used in a GUI toolkit will use Unicode internally. “Weird byte sequences to indicate Unicode characters” is really all that legacy encodings are from the perspective of new software. They are essentially just incompatible ways of specifying a subset of Unicode characters. As for why it has held on so long, there are a few reasons. The Japanese tech world can be quite conservative and skeptical about new things. That is one factor. But I think another is that Japan was really a computing pioneer in the 1980s, and local standards ruled supreme. Compatibility was not a big concern, and even the mighty IBM PC and its clones barely made an impact there for a long time, as it was completely eclipsed by Japanese alternatives. Now, everyone is forced by our increasingly interconnected world to work on international standards, and I can’t help but feel that there is some resentment at not being about to just “do their own thing” anymore. Every time a new encoding extension is proposed, they have to present it to an organization that includes China, Korea, Taiwan, and Vietnam, who will scrutinize it. A few years ago JIS (the Japanese standards organization) actually proposed that each East Asian country should just get their own blocks in Unicode and they should be able to encode whatever they want with no input from others. Of course, none of the other East Asian countries took their proposal seriously. I wish I could find the proposal, because I hate saying stuff like this without sources, but I can tell you that it is buried somewhere among all of the proposal documents that you can find on the website of the Ideographic Research Group, which is the organization that East Asian countries participate in which is responsible for CJK encoding in Unicode. You might find it here with enough scrolling. I have to get on the subway though, so I have to end this comment here. https://appsrv.cse.cuhk.edu.hk/~irg/ https://appsrv.cse.cuhk.edu.hk/~irg/
- cryptonector 3y agoASCII was always a multi-byte codeset. The characters were designed so that one could use overstrike to write things like á (á) by writing a BS ' or ' BS a. This worked for most lower-case characters that have diacritical marks in Latin scripts. It didn't work for upper-case characters unless the terminal or printer understood this and had special fonts, but then, that's why Spanish historically doesn't require accenting upper-case letters. This all goes back to typewriters, where this is how one typically wrote accented characters. And most of this got lost. Though we still use overstrike for underline and bold in the terminal.
- nemetroid 3y agoLanguages with diacritical marks have added them because they are beneficial in disambiguating different sounds. English's lack of diacritical marks is why romanized Korean is so difficult to read, with digraphs used for monophthongs.
- serentty 3y agoIt doesn’t matter. If C is supposed to be a language for solving real-world problems, its strings have to be real-world strings. Wishing writing were more elegant gets you nowhere.
- biorach 3y ago> English gets a lot right with its writing system More like English uses a relatively simple writing system which has the drawback of being unable to properly represent all the sounds of the language. Like a lot of things in life it's a question of tradeoffs, not being right or wrong