4 ms·
Hence I said "UTF-8 as currently defined". You could produce something that uses the same basic scheme as UTF-8 (using the leading byte to indicate the total n
by ubernostrum 8y ago
Hence I said "UTF-8 as currently defined".
You could produce something that uses the same basic scheme as UTF-8 (using the leading byte to indicate the total number of bytes used for the code point), but it would not be UTF-8 as we know it (which caps at four bytes per code point), and different encoders/decoders would need to be developed.