3 ms·
> Unicode characters are limited to the range U+0000..U+10FFFF, so they fit in 21 bits and require no more than four bytes (in the UTF-8 encoding form) or two 1
by vorg 8y ago
> Unicode characters are limited to the range U+0000..U+10FFFF, so they fit in 21 bits and require no more than four bytes (in the UTF-8 encoding form) or two 16-bit code units (in UTF-16)
The private use planes U+F8000..U+FFFFF and U+100000..U+10FFFF can be used as the high and low surrogate ranges in a surrogate-pair encoding scheme similar to that of UTF-16. That would give a range of U+00000000..U+7FFFFFFF, which is in 2^31 space. If the Unicode Consortium removed their arbitrary upper limit of U+10FFFF for codepoints, the UTF-8 and UTF-32 schemes as presently defined would naturally fit into this space without needing to use any surrogation. Only UTF-16 would require any surrogation -- its present scheme and those private use planes for a 2nd tier -- but hopefully the use of UTF-16 would be well on its way out by then.