3 ms·
Here's a potentially random question: Is UTF-8 essentially a fancy way of representing arbitrary precision BigInts where the character-to-integer mapping is spe
by puzzledobserver 6y ago
Here's a potentially random question: Is UTF-8 essentially a fancy way of representing arbitrary precision BigInts where the character-to-integer mapping is specified by Unicode, and which coincides with ASCII encodings on a certain range of numbers?
- 1ris 6y agoUTF-8 is not arbitrary precision. There is a artificial limit at 2^20. And Unicode does not map theses to characters. Unicode maps these to so called "code points". These "code points" then are mapped to charaters ("glypheme clusters" in unicode speak) in a very complicated way. A sequence of code points can be different kinds of normal forms, for example.
- masklinn 6y ago> UTF-8 is not arbitrary precision. There is a artificial limit at 2^20. Since that limit is artificial it can safely be ignored. UTF-8 was designed as a 32b encoding. I don't think it is arbitrary-precision either though, since it was also designed to be self-synchronising. > glypheme Grapheme. A grapheme is more or less a user-perceived character, a "grapheme cluster" is a group (cluster) of codepoints roughly corresponding to a grapheme. > A sequence of code points can be different kinds of normal forms, for example. That has limited relevance, since not all clusters exist in precomposed forms they have to be handled regardless of normalisation, there's just some redundancies.
- masklinn 6y agoUTF-8 is a "fancy" way of representing 32 bit integers in a self-synchronising byte format. I can see an extension to 7 bytes (by making the leading byte 11111110 which I think would still be unambiguous) going up to 37 bits, but at 8 bytes you'd start getting collisions in the byte patterns and would lose the self-synchronising properties.
- kstenerud 6y agoYou could in theory extend the pattern up to 42 bits: 0xxxxxxx 110xxxxx 10xxxxxx 1110xxxx 10xxxxxx 10xxxxxx 11110xxx 10xxxxxx 10xxxxxx 10xxxxxx 111110xx 10xxxxxx 10xxxxxx 10xxxxxx 10xxxxxx 1111110x 10xxxxxx 10xxxxxx 10xxxxxx 10xxxxxx 10xxxxxx 11111110 10xxxxxx 10xxxxxx 10xxxxxx 10xxxxxx 10xxxxxx 10xxxxxx 11111111 10xxxxxx 10xxxxxx 10xxxxxx 10xxxxxx 10xxxxxx 10xxxxxx 10xxxxxx
- masklinn 6y agoTrue, not sure why I figured FF was not an acceptable LBP. Possibly because I assumed you'd need a special continuation byte and all of them would be unavailable, but the existing continuation bytes work fine so that's just not correct. FF could even lead to special patterns of continuation bytes allowing for smuggling more data in there I guess.
- lifthrasiir 6y ago> 32 bit integers The original UTF-8 was at most 31 bits long (one bit in the leading byte, 6*5 = 30 bits in six subsequent bytes). > you'd start getting collisions in the byte patterns and would lose the self-synchronising properties I'm not sure why. The main requirement for self-synchronization in this case is that the leading byte doesn't appear anywhere else, which would be the case for the leading byte FF. It would require substantial modifications to go beyond 7 bytes (say, it would have alternative means to denote the length of following data bytes) but it is surely possible.
- deleted 6y ago[deleted]
- lifthrasiir 6y agoThere is a proposal for infinitely long UTF-8 extensions: http://ucsx.org/%E2%88%9E8 http://ucsx.org/%E2%88%9E8
- goto11 6y agoIt only goes to 31 bits, otherwise you are correct. Although the term "character" is ambiguous so Unicode use more specific terms. An encoded integer represent a "code point" which usually corresponds to a character, but in some cases characters are represented by multiple code points. For example "â" might be represented by the code point for "^" following the code point for "a". This in turn opens a whole can of worms since the same character can be represented as code points in multiple ways.
- Someone 6y agoThe encoding scheme goes to 31 bits, but UTF-8 nowadays goes to ≈20 bits. https://en.wikipedia.org/wiki/UTF-8#Invalid_code_points https://en.wikipedia.org/wiki/UTF-8#Invalid_code_points: “Since RFC 3629 (November 2003), the high and low surrogate halves used by UTF-16 (U+D800 through U+DFFF) and code points not encodable by UTF-16 (those after U+10FFFF) are not legal Unicode values, and their UTF-8 encoding must be treated as an invalid byte sequence.”
- upofadown 6y agoMinor quibble, UTF-8 is used to represent Unicode code points, not integers. People like to map the bit patterns to integers in various bases for stuff like ordering but it is inherently a different sort of thing.