2 ms·
Meanwhile, Windows uses UTF-16 everywhere internally, so "a" and "あ" are both 2 bytes large. You'd still have to exclude surrogate pairs and the Unicode contro
by Dwedit 19d ago
Meanwhile, Windows uses UTF-16 everywhere internally, so "a" and "あ" are both 2 bytes large. You'd still have to exclude surrogate pairs and the Unicode control characters from a filename.
- chrismorgan 19d agoApproximately nothing that uses UTF-16 validates it (I can’t think of a single thing that does), so surrogates will work fine. At least until a UTF-8 system touches it, because they normally do validate (Go is an uncommon exception in not validating).
- chungy 19d agoThe thing is that Windows is actually UCS-2, a consequence of the OS predating UTF-16 by a few years, and coming into existence when Unicode originally thought 16 bits was going to be enough. For backwards compatibility with old file systems, there's no way the OS can start enforcing surrogate codepoints as forbidden from names. You can just so happen to pretend it's UTF-16 until it's not.
- chrismorgan 19d agoNot much is actually UCS-2; almost all things like that are rather unvalidated UTF-16, also known as sequences of UTF-16 code units.
- account42 19d agoThat's correct. Windows does decode surrogate pairs whenever filenames are displayed so its no longer UCS-2.
- Dwedit 19d agoI looked it up. The one place where UTF-16 validation takes place is when using WideCharToMultiByte to convert to UTF-8 text. Before Vista, that was not validated.
- account42 19d agoWhat do you mean by not being validated. UTF-16 to UTF-8 conversion necessitates special treatment for surrogates to correctly convert matched pairs of them to their proper UTF-8 encoding of the code point they represent. The question is what you do when you encounter unmatched pairs: - Abort with an error (not useful) - Replace the unmatched surrogate with a replacement character (afaik that's what WideCharToMultiByte has always done_ - Treat unmatched surrogates like any other non-surrogate code unit and encode the code point they represent (which are reserved for surrogates) as UTF-8 like you would any other code unit. This gets you the WTF-8 encoding which is what you want if you need to lossless represent Windows almost-UTF-16 strings de-facto-but-no-de-jure-UTF-8.