3 ms·
I think NTFS is UCS-2, that is why WTF-8 was invented https://simonsapin.github.io/wtf-8/ https://simonsapin.github.io/wtf-8/
by wolf550e 3y ago
I think NTFS is UCS-2, that is why WTF-8 was invented
https://simonsapin.github.io/wtf-8/ https://simonsapin.github.io/wtf-8/
- layer8 3y agoSurrogate pairs are interpreted by client software, so it’s UTF-16 in that sense. The file system just doesn’t ensure that there won’t be unpaired surrogates, or other noncharacters. This is similar to strings in .NET, Java, and JavaScript.
- heinrich5991 3y agoWhich unfortunately means you can't rely on it being UTF-16.
- naniwaduni 3y agoNor should you. Even a well-formed sequence of utf-16 codepoints can be utter nonsense; there's approximately no level of abstraction between "sequence of fixed-width code units" and "run it through a full-blown a font rendering stack" where it makes sense to assume your input is "well-formed".