4 ms·
This article was written before UTF-8 became the de-facto standard. According to Wikipedia, UTF-8 encodes each of the 1,112,064 valid code points. Much more t
by lcuff 10y ago
This article was written before UTF-8 became the de-facto standard. According to Wikipedia, UTF-8 encodes each of the 1,112,064 valid code points. Much more than Goundry's (the author's) 170,000. Goundry's only complaint against UTF-8 is that at the time, it was one of three possible encoding formats that might work. Since it has now been widely embraced, the complaint is no longer valid.
In short, Unicode will work just fine on the internet in 2016 as far as encoding all the characters goes. Problems having to do with how ordinal numbers are used, right-to-left languages, upper-case/lower-case anomalies, different glyphs being used for the same letter depending on the letter's position in the word (and many other realities of language and script differences) all need to be in the forefront of a developer's mind when trying to build a multi-lingual site.
- WayneBro 10y ago> In short, Unicode will work just fine... > Problems............all need to be in the forefront of a developer's mind when trying to build a multi-lingual site. It will work. Just fine though? It sounds like way too much work!
- Dylan16807 10y agoUnicode handily solves the problem of storing text. Manipulating text, though, is inherently nightmarish. No format can prevent that.
- deleted 10y ago[deleted]
- syncsynchalt 10y agoThis article is written years after the 16-bit problem was solved, yet seems to be completely unaware of the solution. (It's unclear if the mentions of adding "octet blocks" is a misunderstanding of the solution, or a dismissal.) It mentions Unicode 3.1, yet surrogate pairs were introduced in Unicode 2.0 five years before the doc was written. In actuality we don't concern ourselves overmuch with planes, "octet blocks", and so on these days. There's just a single numbered list of characters* that we call "Unicode", and a few fairly minimal algorithms for transforming a list of characters into a bytestream ("UTF-8", "UTF-16", etc). Still I love it as a historically interesting document. It's jarring to see "Oriental" used, I suspect the same author would never use that word in a professional context today. [*] I use "characters" here when "code points" would be more accurate, but the details of that comparison aren't meaningful to this argument. I also don't go into normalization and so on, as this article seems to be more about the feasibility of fitting >100K code points into actual-bits-on-the-wire.
- LukeShu 10y agoIt doesn't have anything to do with UTF-8. It has to do with Unicode‡ growing from a 16-bit address space, into a ~20.1-bit address space; 17 16-bit "planes"‡‡ (log₂(17×2¹⁶)≈20.1bits). This expansion happened sort-of gradually. Unicode 3.1 (2001) assigned the Unicode 3.0 character set to be "plane 0", and added 14 additional planes‡‡‡ (for a total address space of log₂(15×2¹⁶)≈19.9bits). I'm not sure exactly which version between 3.1 and 9.0 the additional two planes in. That is to say, Unicode 3.1 solved the address-space problem. So why does the article have a section "Why Unicode 3.1 Does Not Solve the Problem"? Well, there are two answers: the one that the article suggests, and the one giving the author the benefit of the doubt. The way the section is written, it seems that the author thinks the the address space of 16 bits + 16 bits is 2×2¹⁶ instead of 2^(2×16); because they are in separate "16 bit blocks". They seem to think that maybe 1 bit was added, but think that that just grew the encoding size by 16 bits. It honestly read to me like it was written by a linguist who only had a rudimentary grasp of programming; I was surprised when it said the author was a programmer at the bottom. Giving him the benefit of the doubt that the argument was only poorly expressed, not poorly thought: Unicode 3.1 expanded the address space by another 917,504 encodable characters; more than enough! However, it didn't actually define 900,000 more characters; it only defined 44,946 more characters (as the author noted). But it wasn't limited to 16 bits anymore; it had a full 19 (more actually; 19.9-ish!) to work with. The author even mentions that 18 bits would have been plenty. Well Unicode 3.1 got them! They just weren't allocated yet. That said, Unicode 9.0 (2016) still only has about 128,000 characters defined in it. A far cry from the author's claim of 170,000 characters needed to satisfy asian languages. ‡: To say otherwise would be to confuse Unicode with its encodings, a mistake that only leads to confusion. ‡‡: The term "plane" comes from ISO/IEC standards dealing with character sets. Unicode 3.0 corresponded to the ISO/IEC "Basic Multilingual Plane"; so each 16-bit group of characters got its own cutesy name as a "Plane" to match. ‡‡‡: Why 15 planes, then 17? It has to do with what was encodable with existing encodings. That isn't to say that Unicode was limited by the encodings; but that it was informed by them. The growth beyond plane 0 meant that UCS-2 had to be phased out for UTF-16 (its successor), as UCS-2 couldn't encode anything but plane 0. However, seeing that UTF-16 could only encode 17×2¹⁶ characters, it made sense to limit the number of planes to 17 if there isn't a pressing need for more; as doing so would require obsoleting UTF-16. And given that the current address space is only about 12% utilized, there's no reason to mandate phasing out UTF-16 yet.