4 ms·
The OS treats strings as a blob, yes, but typically specifies that they're a blob of nul-terminated data. Unfortunately some text encodings (UTF-16 among them)
by syncsynchalt 1y ago
The OS treats strings as a blob, yes, but typically specifies that they're a blob of nul-terminated data.
Unfortunately some text encodings (UTF-16 among them) use nuls for codepoints other than U+00. In fact UTF-16 will use nuls for every character before U+100, in other words all of ASCII and Latin-1. Therefore you can't just support _all_ text encodings for filenames on these OSes, unless the OS provides a second syscall for it (this is what Windows did since they wanted to use UTF-16LE across the board).
I've only mentioned syscalls in this, in truth it extends all through the C stdlib which everything ends up using in some way as well.
- codedokode 1y agoYou should not be passing file names in different encodings because other apps won't be able to display them properly. There should be one standard encoding for file names. It would also help with things like looking up a name ignoring case and extra spaces.
- syncsynchalt 1y agoI mean, I agree there _should_ be one standard encoding, but the Unix API (to pick the example I'm closest to) predates these nuances. All it says is that filenames are a string [of bytes] and can't contain the bytes '/' or '\0'. It is good for an implementation to enforce this at some level, sure. MacOS has proved features like case insensitivity and unicode normalization can be integrated with Unix filename APIs.
- 1718627440 1y agoYou're right I missed that. Sounds like blob size should be communicated out of band.