8 ms·
What Every Programmer Absolutely, Positively Needs To Know About Encodings (2011)
- torstenvl 5y agoThe world of text encodings is pretty insane, especially when you start getting into the realm of what seems like endless variations on multi-level encodings, like the bajillion different character set encodings for quwei/kuten CJK encodings. I'm only scratching the surface right now, and just wrote a CPG 932 → Unicode lookup utility. https://pastebin.com/4PYmEjQZ https://pastebin.com/4PYmEjQZ (For internal testing, forgive any sloppiness, but feel free to use the kuten table if you happen to have a niche project with one-way mapping. The tables themselves are facts and not creative expression, so should not be copyrightable, but I'm dedicating the project to the public domain anyway.)
- ben-schaaf 5y agoI recently implemented GB18030, which is a 1,2 or 4 byte encoding with a giant code page covering the entirety of unicode.
- dhosek 5y agoBack in the 90s I had a 4-inch binder stuffed with every encoding registered with ECMA. I'll take UTF-8 with all of its warts over the mess that was the pre-Unicode encoding hell.
- kmeisthax 5y agoUTF-8 actually isn't the warty bit of Unicode. It's actually fairly reasonable[0], with the only downside being that you can't efficiently index strings by integer. If you exclusively use UTF-8 everywhere in your application, you will remain wart free. UTF-16 is the warty bit. Why? Because it was designed to be a fixed-width encoding, under the assumption that 65,536 code points would cover all languages we needed. Well, spoiler alert: this wasn't enough, even with the mess that was UniHan. So they carved out a chunk of those code points from the 16-bit part of Unicode, called them "surrogate code points", and used them to extend UTF-16 into a variable-width encoding. Thus, they could have 16 planes of 65,536 characters each. They weren't valid characters beforehand, so obviously we could just repurpose them as flags for "astral" characters outside the Basic Multilingual Plane. Simple, right? Except that plenty of code to work with UTF-16 strings already existed, and assumed UTF-16 was a fixed-width encoding. If you chucked an astral character at it, it'd think it's two rather than one. It'll also happily split or rearrange astral characters into invalid sequences of codepoints. Furthermore, if you ask these systems for UTF-8, it'll happily encode the surrogates, leading to two nested layers of variable encodings. Oh, and the two biggest users of UTF-16 were Windows NT and JavaScript. So we're absolutely stuck with fixed-length UTF-16 and the mangled WTF-8[1] encodings they will output until the end of time. Furthermore, if we ever do fill up the astral planes, we'll have to define superastral characters, which means having supersurrogate code points that will break modern UTF-16 again and make WTF-8 handling even more complicated. [0] Unless you use MySQL, where someone decided to break UTF-8 early on and now there's "utf8mb4" to un-break it. [1] UTF-8 which contains encoded UTF-16 surrogates, as defined in https://simonsapin.github.io/wtf-8/ https://simonsapin.github.io/wtf-8/
- bruce511 5y agoYou are completely right. Utf-16 is either 2 or 4 bytes long per code point (not per character) but a lot of older programs (and older programmers) see it as a "2 byte per code point" string. There is a 2 byte per code point encoding, its called UCS2 so its important to make sure when something says utf-16 they don't mean UCS2. I am not familiar enough with UCS2 to know if it also supports multiple code points per char - but I expect it does - so the mental model of 2 bytes per char breaks down even there.
- kmeisthax 5y agoIt depends on which UCS-2 you're talking about. Originally, UCS-2 was variable-width, and then it was fixed-width, and then UTF-16 was proposed to make it variable-width again. The reason for this is actually because there used to be a different standard called UCCS, which was in competition with Unicode at one point before being withdrawn and replaced by it. But the actual historical drama is very relevant for why UTF-16 is such a warty mess. ISO proposed UCCS (then ISO 10646), where you had 31-bit code points. Each codepoint was organized by it's bytes into "groups" (high byte), "planes", "rows", and "cells" (low byte). Furthermore, you couldn't have a group, plane, row, or cell that was an ASCII or ISO control code (00-1F/80-9F), so everything was in Group 20 and Plane 20. The standard then defined 1, 2, and 4-byte encodings of these 31-bit codepoints. These were UCS-4, UCS-2, and UTF-1[0]. UCS-2 as defined in UCCS very much supports multiple code points per character; because UCS-4 was supposed to be the fixed-width encoding. The extension mechanism is very different from UTF-16 surrogates and is actually the same as defined in ISO 2022. That is, when you wanted to use a character from a different group or plane, you would actually encode special code-switching words that said you wanted to change all further characters in the string to that new group or plane. This is also why control codes were forbidden in UCCS code points, for the same reason why UTF-16 surrogates are forbidden in Unicode control points. The underlying problem with UCS-2 (and, for that matter, UTF-1) was that it was not self-synchronizing. In other words, if I take a UCS-2 string and delete one byte from it, all the characters onwards flip their byte order. If I delete a word that's part of a code-switch, then all characters onwards will be misinterpreted in the wrong group or plane. This is a terrible property for text encodings where erroneous deletions or mutations can occur. A couple of things happened that killed UCCS: - Software companies vehemently opposed 31-bit codepoints as wasteful and insisted on doing everything in 16-bits. They promoted a competing standard called Unicode which defined one 16-bit, fixed-width (and, thus, self-synchronizing) encoding just called "Unicode". - When designing Plan 9[1], the designers of that OS realized they could make UTF-1 self-synchronizing, and proposed UTF-8 as an alternative encoding for UCCS. This also had the advantage of not requiring integer division to decode, which made it faster on contemporary processors. ISO ultimately caved and withdrew the original version of ISO 10646. Unicode 1.1 was altered slightly to please ISO, but ISO 10646 was altered drastically to match Unicode. The resulting standard became ISO 10646-1:1993; and future versions of ISO 10646 are more or less identical to specific Unicode versions. At this point our nomenclature changes, because in Unicode 2.0 they define or reference several encodings[2][3]: - UCS-2: Fixed-width 16-bit words. At this point, UCS-2's code-switching mechanism had been withdrawn from ISO 10646, making it identical to just an array of Unicode 1.1 codepoints. - UCS-4: Still specified by ISO 10646 to support 31-bit codepoints, but with the control code restriction withdrawn. Since they were now using group 0 and plane 0 exclusively, the Unicode 2.0 spec says to treat this as codepoints with extra zero padding. At some point in the future this would become UTF-32 with a restriction to 16-bit + astral codepoints. - UCS Transformation Format, 7-bit form; or UTF-7: Backwards compatibility for MIME, because e-mail is evil. - UCS Transformation Format, 8-bit form; or UTF-8: Backwards compatibility for US-ASCII and UNIX filesystems. This is also referred to as "File System Safe UTF", "FSS-UTF", or "UTF-2", but caps characters at 4 bytes, prohibiting the use of 31-bit codepoints. - UTF-16: The historical mess I mentioned in the previous comment. Interestingly enough, several parts of the 2.0 spec differ as to whether or not "Unicode" means UCS-2 or UTF-16. In the UTF-8 appendix they specify 4-byte UTF-8 characters as having a "Unicode value" consisting of two UTF-16 surrogates. But in other parts of the spec, such as Table C-1, they list "UCS-2" and "Unicode" as equivalent. I have a feeling that surrogates were a last-minute compromise between ISO and Unicode, given that the UTF-16 spec warns against using existing private-use superastrals[4] already allocated in ISO 10646. [0] Wikipedia calls this UTF-1 but I don't have access to the old versions of ISO 10646 to check if that standard actually called it that or UCS-1. The Unicode spec calls UTF-8 "UCS Transformation Format-8" so it's possible that the "UTF" nomenclature actually came from ISO first. [1] Plan 9 is a really fascinating OS, because, while it wasn't very commercially successful, it wound up being highly influential. Things like /proc in Linux were ripped wholesale from Plan 9. It even had C extensions that the designers would later copy in Go (which they also made). [2] https://www.unicode.org/versions/Unicode2.0.0/appA.pdf https://www.unicode.org/versions/Unicode2.0.0/appA.pdf [3] https://www.unicode.org/versions/Unicode2.0.0/appC.pdf https://www.unicode.org/versions/Unicode2.0.0/appC.pdf [4] Term I coined in the previous comment for characters beyond U+10FFFF. Encoding them in UTF-16 would require defining a second level of astral surrogate pairs, themselves being encoded as pairs of surrogates.
- deleted 5y ago[deleted]
- agumonkey 5y agoA nice complement to nedbat's https://nedbatchelder.com/text/unipain/unipain.html#1 https://nedbatchelder.com/text/unipain/unipain.html#1
- kgm 5y agoI would also recommend https://www.joelonsoftware.com/2003/10/08/the-absolute-minimum-every-software-developer-absolutely-positively-must-know-about-unicode-and-character-sets-no-excuses/ https://www.joelonsoftware.com/2003/10/08/the-absolute-minim...
- cxcorp 5y agoThis is linked to in the second paragraph of the post
- oshiar53-0 5y agoFun fact: GB 18030 is a Unicode Transformation Format. Example: \N{THINKING FACE}\N{FACE WITH TEARS OF JOY}\N{FACE SCREAMING IN FEAR}\N{SMILING FACE WITH SMILING EYES AND THREE HEARTS}\N{PERSON DOING CARTWHEEL}\N{FACE WITH NO GOOD GESTURE}\N{ZERO WIDTH JOINER}\N{FEMALE SIGN}\N{VARIATION SELECTOR-16}\N{EYES}\N{ON WITH EXCLAMATION MARK WITH LEFT RIGHT ARROW ABOVE}\N{SQUARED COOL}\N{VARIATION SELECTOR-16} In UTF-8: 00000000: f09f a494 f09f 9882 f09f 98b1 f09f a5b0 ................ 00000010: f09f a4b8 f09f 9985 e280 8de2 9980 efb8 ................ 00000020: 8ff0 9f91 80f0 9f94 9bf0 9f86 92ef b88f ................ In GB 18030: 00000000: 9530 cd34 9439 fc38 9530 8335 9530 d636 .0.4.9.8.0.5.0.6 00000010: 9530 d130 9530 8535 8136 a439 a1e2 8431 .0.0.0.5.6.9...1 00000020: 8235 9439 cf38 9439 e537 9439 8b32 8431 .5.9.8.9.7.9.2.1 00000030: 8235 .5
- lifthrasiir 5y agoWhich is carefully designed to work around existing codes that only expect at-most-two-byte-long encoding, e.g. Windows's IsDBCSLeadByte(Ex). Normally a bad design for a new-ish encoding, but a reasonable one given that it's meant to be a superset of GBK---an already bad but widespread encoding.
- bombcar 5y agoWhat is a Unicode transformation format?
- e-dt 5y agoIt's an encoding that encodes all of Unicode. The "UTF" in UTF-8, etc. stands for Unicode Transformation Format.
- bruce511 5y agoA detail skimmed over in the article, but one which has significant importance is that; "one code point in unicode does not necessarily map to one character on the screen." A "character" can, and often does, get constructed from multiple code points. This doesn't help an already complicated sorting issue (who knew "sort these names alphabetically" could be an ambiguous statement).
- dhosek 5y agoThat said, the Unicode standard does a great job of tackling that question and providing a framework that can cover all the various language standards (and even with a fully normalized Unicode string you still need to do things like handle ch as a single character for sorting in Czech or Spanish).
- pezezin 5y agoIn Spanish, "ch" and "ll" stopped being considered a single character for sorting purposes in 1994. In have some dictionaries printed before that date, and I find them quite confusing...
- dhosek 5y agoAs someone who learned his Spanish before 1994 (and whose dictionary also predates that time), I was unaware of this until now.
- Toutouxc 5y agoAs a Czech I would love "ch" to go fuck itself for that exact reason.
- oldsecondhand 5y agoIn Hungarian "dzs", "zs" and "sz" are so called compund letters, however most software treats them as separate characters, because it would be a pain in the ass do sorting the correct way. (It would either require changing the input method, or software would need to be aware of etymology.) It's only sorted the right way in paper dictionaries.
- deleted 5y ago[deleted]
- ncmncm 5y agoPHP-centric, not mentioned in the title. Most of it is relevant to everybody, but it is jarring to run into stuff about PHP. Isn't that dead yet? Of greater moment is that the article keeps talking about "characters", which is an undefined term in Unicode. Unicode offers you code points, code units, graphemes, grapheme clusters, and ... other things, none of which maps to the grouping of dots you see on your screen (and probably cannot imagine how to type in). "Character" has outlived its sell-by date. Let it be retired and buried with dignity, but with a good thick slab of concrete on top. It also fails to mention "expanded form" and "canonical form", and other ways that two completely different sequences of bits mean, at some level, the same text. Different forms are convenient for different things; there is a shortest possible representation nice for sending and storing, and a maximally decomposed representation that might be best for editing if you like adding and removing diereses ("umlauts") and accents piecemeal. And it fails to mention WTF-8, a way to package up byte sequences that are not valid UTF-8, but may have valid UTF-8 characters that you want to display in case they offer the poor human a clue as to what was intended. WTF-8 sequences often arise in file systems and databases that don't enforce any particular encoding, but just store whatever bytes the benighted programs users run provide as, e.g., names for files. You wish you could display them in sorted order. There had better be a way to point at it, because there is no way to type it. But you have to store it, because that is the only way to tell the OS which file you wanted to rename or delete. Deletion is tempting, but we can't, always, can we?.
- ahartmetz 5y agoPHP is very much not dead. WordPress still seems to be popular. Facebook replaced its optimizing PHP VM with another one a few years ago. And most importantly, PHP is a low-friction, decent performance way to create simple dynamic websites because it was made for that.
- jraph 5y agoCome on, the PHP part is small and at the end, and most things said in this part apply in any language that works the same way, notably C and C++ which also don't care about what you put into their strings.
- 867-5309 5y ago
- _8j50 5y agoShould mention enduanness as well and ebcdic. There are big endian versions of UTF-*
- samatman 5y agoNot -8 there isn't. One of the many reasons to deprecate all other Unicode encodings whenever it's in one's power to do so.
- seles 5y agoAccording to the article, PHP can handle other encodings by just treating sequences of strings as byte sequences and not caring what the encoding is. There example: $string = "漢字"; But if you are using say UTF-8 and one of those Chinese characters has one of its bytes have a value of 34 (the ascii value of "), then wouldn't the string terminate prematurely? Edit: to answer my own question, quote from wikipedia: ASCII bytes do not occur when encoding non-ASCII code points into UTF-8
- chmod600 5y agoAlso, the compiler might be treating the input file as UTF-8, while the semantics of the language may treat string literals as the sequence of bytes when encoded as UTF-8.