3 ms·
I think Bob unfairly discounts the work that IBM did here. It's clear that there was synchronicity between the wants of the Plan 9 team and what IBM was doing.
by Todd 12y ago
I think Bob unfairly discounts the work that IBM did here. It's clear that there was synchronicity between the wants of the Plan 9 team and what IBM was doing. The need was clearly 'in the air' back then.
Upon reading this, it doesn't appear to me that Bob watched Ken design this system on the back of a placemat one fine day. Instead, they got a complete design with working code from a team at IBM. Being the seasoned bit twiddlers that they were, they realized that they could do it better. Their design is clearly an improvement, and their code is more elegant. In addition to getting the conceptual head start, they also appear to have received the full support of the IBM team in promoting and promulgating the RFC as a standard. To me, it's also story of humility vs. hubris.
As an aside, Windows NT shipped in 93, and wasn't able to benefit from this new encoding format. As a result, UTF-16 is baked-in and Windows developers have been living with it ever since (even if file formats are largely UTF-8 now, the API is still solidly UTF-16).
- TazeTSchnitzel 12y agoMS sadly won't add a UTF8 codepage. If they did, people could write UTF8 Windows apps. Alas. http://utf8everywhere.org http://utf8everywhere.org
- vorg 12y agoOne big obstacle to UTF8 uptake is in 2003 Unicode was clipped down to about a million codepoints, instead of the 2.1 billion originally specified by Pike and Thompson. Only 137,000 of those million are available for Private Use, and if a user recognizes the ConScript allocations such as the Kinya Syllables in F0000..F0E6F, those 137,000 are cut down quickly. There's only room left for a few dozen more Hangul-style syllabaries!
- TazeTSchnitzel 12y agoYeah, I do wonder how long those 21 bits will last. They clipped it down to ensure it'd remain UTF-16 compatible, but I worry if we might need to extend UTF-8 in future to support more codepoints, and add some horrible new set of surrogate pairs, encoding using surrogate pairs, to UTF-16.
- vorg 12y agoAnyone can extend UTF-16 to the UTF-8 repertoire of codepoints by using the top 2 private use planes as high and low 2nd-tier surrogates respectively. Hopefully, by the time UTF-8 (and UTF-32) need to be extended, UTF-16 will have fallen into general disuse and almost no-one will need to use those 2 levels of surrogates. Java/Android and .NET strings aren't being broken anytime soon, however, so it seems to me people should start talking about the details of such a 2nd-tier surrogate system.
- enneff 12y agoI'm not sure what point you're making. You say almost exactly what Rob says (who's "Bob" btw?), which is that Ken invented the UTF-8 encoding scheme to improve on IBM's proposal. I don't see hubris here; just a desire to give Ken the recognition he deserves for some clever work.
- Todd 12y ago(Thanks for the correction.) Perhaps I was reacting to the first paragraph of the post. He says the 'incorrect story' is that IBM designed it and Rob and Ken just implemented it. The correct version is that Ken designed it. I think that's a bit strong and doesn't give enough credit to the IBM team. Upon a more careful read it appears that they designed very similar approaches in parallel, with Ken's version being more elegant. I still don't think it gives enough credit to the IBM team. But you're point it well taken.