3 ms·
I don't see what's inherently wrong with UTF-16 surrogates. If I am not wrong, a given UTF-16 codeunit is unambigously either a complete code point, a first sur
by cylemons 3y ago
I don't see what's inherently wrong with UTF-16 surrogates. If I am not wrong, a given UTF-16 codeunit is unambigously either a complete code point, a first surrogate, or a second surrogate.
Why should we expect invalid utf-16 strings to be representable in utf-8 or 32? I don't see anyone trying to represent invalid utf-8 in utf-16 or 32.
- kps 3y ago> Why should we expect invalid utf-16 strings to be representable in utf-8 or 32? We shouldn't care. UTF-16 should just be an encoding and its internal details shouldn't leak into Unicode code points. There's just no good reason to exclude code points U+D800–U+DFFF merely because 0xD800–0xDFFF happen to be used specially in UTF-16 encoding, just like U+0080–U+00FF aren't excluded merely because (most of) 0x80–0xFF are used in UTF-8 encoding.
- cylemons 3y agoIs having a hole from U+D800 to U+DFFF such a big deal? The parent comment was specifically talking about surrogate pairs. That to me looks more like buggy implementation issue rather than standards issue.
- kps 3y agoThe main issue is that it adds validation code (if one is sticking to the standard) for things that don't care about UTF-16 at all. It does occupy 1/32 of the BMP, displaying a couple thousand potential actual characters (making them take an extra byte in UTF-8, and an extra two in UTF-16).
- chrismorgan 3y agoAs a hole, it would only be annoying and a performance penalty for validation. But by its very design, it will leak, and it does in such ways that it became the worst thing to ever happen to Unicode. I don’t know of a single language or library that uses UTF-16 for strings that validates strings: every last one actually uses sequences of UTF-16 code units, potentially ill-formed, and has APIs that guarantee this will leak to other systems. This has caused a lot of trouble for environments that then try to work with the vastly more sensible UTF-8 (the only credible alternative for interchange). Servo, for example, wanted to work in UTF-8, for massive memory savings and performance improvements, but the web has built on and depends on UTF-16 code unit semantics so much that they had to invent WTF-8, which is basically “UTF-8 but with that hole filled in” (well, actually it’s more complicated: half filled in, permitting only unpaired surrogates, so that you still have only one representation). So: the problem is that the Unicode standard was compromised for the sake of a buggy encoding (they should instead have written UCS-2 off as a failed experiment), and every implementation that uses that buggy encoding is itself buggy, and that bugginess has made it into many other standards (e.g. ECMAScript).
- cylemons 3y agoLet say A is an ill formed utf-16 string with unmatched surrogates. The problem comes when trying to convert A to utf-8. Is this the leak you are talking about?
- chrismorgan 3y agoThat’s one of the two situations I speak of: when it happens in practice. The other is… well, much the same really, but when it makes it into specs that others have to care about. The web platform demonstrates this clearly: just about everything is defined with strings being sequences of UTF-16 code units (though increasingly new stuff uses UTF-8), so then other things wanting to integrate have to decide how to handle that, if their view of strings is different: whether to be lossy (decode/encode using REPLACEMENT CHARACTER substitution on error), or inconvenient (use a different, non-native string type). Rust has certainly been afflicted by this in a number of cases and ways, generally favouring correctness.