3 ms·
Approximately nothing that uses UTF-16 validates it (I can’t think of a single thing that does), so surrogates will work fine. At least until a UTF-8 system tou
by chrismorgan 23d ago
Approximately nothing that uses UTF-16 validates it (I can’t think of a single thing that does), so surrogates will work fine. At least until a UTF-8 system touches it, because they normally do validate (Go is an uncommon exception in not validating).
- chungy 23d agoThe thing is that Windows is actually UCS-2, a consequence of the OS predating UTF-16 by a few years, and coming into existence when Unicode originally thought 16 bits was going to be enough. For backwards compatibility with old file systems, there's no way the OS can start enforcing surrogate codepoints as forbidden from names. You can just so happen to pretend it's UTF-16 until it's not.
- chrismorgan 23d agoNot much is actually UCS-2; almost all things like that are rather unvalidated UTF-16, also known as sequences of UTF-16 code units.
- account42 22d agoThat's correct. Windows does decode surrogate pairs whenever filenames are displayed so its no longer UCS-2.
- Dwedit 22d agoI looked it up. The one place where UTF-16 validation takes place is when using WideCharToMultiByte to convert to UTF-8 text. Before Vista, that was not validated.
- account42 22d agoWhat do you mean by not being validated. UTF-16 to UTF-8 conversion necessitates special treatment for surrogates to correctly convert matched pairs of them to their proper UTF-8 encoding of the code point they represent. The question is what you do when you encounter unmatched pairs: - Abort with an error (not useful) - Replace the unmatched surrogate with a replacement character (afaik that's what WideCharToMultiByte has always done_ - Treat unmatched surrogates like any other non-surrogate code unit and encode the code point they represent (which are reserved for surrogates) as UTF-8 like you would any other code unit. This gets you the WTF-8 encoding which is what you want if you need to lossless represent Windows almost-UTF-16 strings de-facto-but-no-de-jure-UTF-8.