10 ms·
The biggest codepoint in Unicode fits into 4 bytes of UTF-8. UTF-8 would allow up to 6 bytes, but those codepoints are not in use currently. If they ever become
by warpspin 5y ago
The biggest codepoint in Unicode fits into 4 bytes of UTF-8. UTF-8 would allow up to 6 bytes, but those codepoints are not in use currently. If they ever become in use, yes, you'd probably need a new character set again. But then a lot more things will break, as higher codepoints would be incompatible with UTF-16 also.
- ghusbands 5y agoUTF-8 only allows 4 bytes, since 2003: https://datatracker.ietf.org/doc/html/rfc3629 https://datatracker.ietf.org/doc/html/rfc3629