4 ms·
UTF-8 must be the default and only encoding. Why does anything else still exist?
by plugnburn 11y ago
UTF-8 must be the default and only encoding. Why does anything else still exist?
- jmnicolas 11y agoYes but UTF-8 with or without byte order mark ? ;-)
- plugnburn 11y agoWithout. BOM (when used for UTF-8) is an obsolete crap invented by necrosoft in order to make their software incompatible with normal.
- TazeTSchnitzel 11y agoIt's not a Microsoft invention, and MS's use of it is really quite sensible. They had a problem of distinguishing UTF-16, UTF-8 and non-Unicode (possibly a single-byte "extended ASCII" type encoding, possibly some multi-byte monstrosity) text files. Since UTF-8 and ASCII-compatible encodings look similar when there aren't many >U+007F characters in use, and identical if none are in use, they could get confused. Prepending a Byte Order Mark solves this problem, in that it makes a file unambiguously UTF-8 (or UTF-16, for that matter).
- tempodox 11y agoHow do you have a BOM in the shell?
- plugnburn 11y agoSome masochist M$-fan could invent even this just in order to justify the difference from civilized world.
- Tiksi 11y agoANSI must be the default and only encoding. Why does anything else still exist?
- plugnburn 11y agoBecause, you see, not everyone in the world uses Latin characters. UTF-8 must become a new standard instead of that whole obsolete encoding zoo.
- Tiksi 11y agoAnd in 20-30 years we'll likely be saying the same about UTF-8. I figured that "ANSI" would give away that I wasn't being serious since it's not actually an encoding.
- plugnburn 11y ago> And in 20-30 years we'll likely be saying the same about UTF-8. Well... If we will, why not? But the thing is that in 20-30 years we won't be able to invent any new writing systems that UTF-8 won't cover. Single-byte encodings were doomed because of their single-byteness. The same awaits two-byte encodings like UCS-2 (aka UTF-16BE) - we already have extended code points for something that glamour hipsters call "emoji". Variable-byte encoding will never become obsolete.
- TazeTSchnitzel 11y agoUnicode is currently limited to 21 bits for compatibility with UTF-16. Eventually we might manage to exhaust all available codepoint space, and with that we'd have to move to yet another encoding with a whole new kind of surrogate pairs. Though UTF-8 could originally handle 31 bits, that's no longer the case.
- plugnburn 11y agoSo I see 2 steps here: dropping UTF-16 altogether (well, already, because there are plenty of extended codepoints above 0xFFFF), and when approaching the 31-bit limit - inventing something like "zero-width codepoint joiner" to compose codes of arbitrary length. For example, in a hypothetical alien language, a hypothetical character "rjou" would have a code 0x2300740457 (all the previous codes are exhausted). We can't express this with a single code, so actually we split it into 2-byte parts and write "#" (0x0023), joiner, "t" (0x0074), joiner and "ї" (cyrillic letter yi, 0x0457). As we have a joiner between these codes, we know that we must interpret and display them not as a "#tї" sequence but as a single alien "rjou" character. I think you get the idea.
- Grue3 11y agoBecause you don't want to give the geniuses who came up with stuff like "Han Unification" a monopoly on encoding.