3 ms·
> Because Unicode is not an encoding. > Overall, Unicode is yet another encoding scheme. ?
by ____tom____ 1y ago
> Because Unicode is not an encoding.
> Overall, Unicode is yet another encoding scheme.
?
- Terr_ 1y agoYeah, author seems to have made a mistake there. > Unicode is a large table mapping characters to numbers and the different UTF encodings specify how these numbers are encoded as bits. Overall, Unicode is yet another encoding scheme. I would guess this represents a confusion between the narrow abstract definition of Unicode versus the way it is casually used as an umbrella term which includes stuff like Transformation Formats.
- jibal 1y agoThe author doesn't understand what a character is, despite the Unicode standard making it very clear that character != codepoint
- btilly 1y agoThat's just somewhat sloppy. Unicode is not an encoding of text to bits. It is an encoding of text to numbers. There are a variety of encodings of text to bits based on how those numbers are to be encoded into bits. Though technically Unicode isn't even quite that. For example "é" can be encoded as U+00E9 or as U+0065,U+0301. Going the other way, "水", U+6C34, is drawn differently in simplified Chinese, Japanese, and traditional Chinese. Unicode calls this, "language-sensitive glyph variation". Which means that the correspondence between text and Unicode is many to many both ways. And then the Unicode can show up in bits and bytes again in multiple ways.