14 ms·
With all due respect, it seems to me that the documents contradict the claim in the initial email. Rob's initial claim is that Ken came up with UTF-8 entirely
by Camillo 7y ago
With all due respect, it seems to me that the documents contradict the claim in the initial email.
Rob's initial claim is that Ken came up with UTF-8 entirely from scratch, without even looking at IBM's proposal.
But Ken's oldest document from Sep 2, 1992 has his changes simply appended after the original FSS-UTF document from IBM (starting at "We define 7 byte types", as mentioned in the email), and notes the changes from original spec.
Then the final document that he sent out (Sep 8) is basically the FSS-UTF document with his changes applied (this is also mentioned in the email!).
There are two changes, basically:
1. Use 10 instead of 1 as the prefix for continuation bytes, so you can synchronize from an arbitrary location.
2. Once the bits are reassembled, use the value as-is instead of adding a bias constant, which simplifies the code at the cost of a tiny bit of packing efficiency.
Given the documentation provided, I would say that the fairest description of the development of UTF-8 is that IBM came up with the initial design, and Plan 9 made two improvements on it to produce the final UTF-8 design.
- leshow 7y agoOne thing I was confused about. The document says there are 7 byte types, but I thought UTF-8 was variable width up to only 4 bytes. Did I misunderstand something?
- hyperman1 7y agoBoth are correct: This original UTF-8 encoding can encode values up to 2^32. But because UTF-16 encoding limits possible values to 16 planes of 64K values, unicode has a hard limit of 2^20 codepoints. This means UTF-8 encoded values of more than 4 bytes can never represent a valid unicode codepoint even if they produce a valid 32 bit numerical value.