3 ms·
Basically, you don't want to use UTF-8 to encode UTF-16 surrogate code points The awful truth is that there is such a beast. UTF-8 wrapper with UTF-16 surrogat
by CountSessine 5y ago
Basically, you don't want to use UTF-8 to encode UTF-16 surrogate code points
The awful truth is that there is such a beast. UTF-8 wrapper with UTF-16 surrogate pairs.
https://en.wikipedia.org/wiki/CESU-8 https://en.wikipedia.org/wiki/CESU-8
- nayuki 5y agoIs CESU-8 a synonym of WTF-8? https://en.wikipedia.org/wiki/UTF-8#WTF-8 https://en.wikipedia.org/wiki/UTF-8#WTF-8 ; https://simonsapin.github.io/wtf-8/ https://simonsapin.github.io/wtf-8/
- LegionMammal978 5y agoNo. Let UTF-8* denote UTF-8, except that surrogate code points are allowed and have no special meaning. To encode a string in CESU-8, each supplementary character is converted to a surrogate pair, leaving existing surrogates alone, then each codepoint is encoded as UTF-8*. To encode a string in WTF-8, each surrogate pair is converted to a supplementary character, leaving unpaired surrogates alone, then each codepoint is encoded as UTF-8*. So really, they have the opposite effect; CESU-8 always uses surrogate characters, whereas WTF-8 removes them if possible. Both are similar in that they directly UTF-8*-encode unpaired surrogates.
- account42 5y agoThe obvious advantage being that WTF-16 -> WTF-8 conversion maps all valid UTF-16 to the corresponding valid UTF-8 and only unpaired surrogates will produce invalid UTF-8 - but only invalid because UTF-8 explicitly disallows encoding those surrogates and not because the actual encoding differs, so you can almost always treat WTF-8 as UTF-8.