3 ms·
Why is that a problem? That's how unix does it as well. A file path is just a sequence of bytes. Don't try to validate it in any sort of encoding, it's just by
by dahfizz 3y ago
Why is that a problem? That's how unix does it as well.
A file path is just a sequence of bytes. Don't try to validate it in any sort of encoding, it's just bytes that the OS uses to identify a file.
- chungy 3y agoThe problem is that if the operating system allows sequences that are not valid UTF-16, then you cannot just state "it uses UTF-16": anything that is expecting UTF-16 and only UTF-16 is going to break on valid file names. "That's how unix does it as well" is the exact same thing. Most people these days expect file names in UTF-8, but not all file systems restrict the valid character sequences to be UTF-8 compliant. If you treat everything as if it must be UTF-8, your application may just break when it encounters a name that isn't UTF-8. (This detail gets messy because there are configurations of, eg, ext4 and ZFS that limit valid file names to be valid UTF-8.)
- OkayPhysicist 3y agoDoes this actually matter? Who is out there, in the wild, creating file paths that use incredibly cursed random sequences of bytes? Maybe it's just me, but I don't really see the need to accommodate people who do stupid things to see what breaks.
- chrismorgan 3y agoOn Windows, it doesn’t matter much because, although the file system doesn’t validate Unicode, the usual system calls that work with file names do, so it’s very rare to encounter non-Unicode file names. On Linux, it’s not as rare as you might expect to end up with a few non-UTF–8 file names, though they’ll normally be on things you won’t often touch directly, so things like desktop software can reasonably only support UTF-8 paths.
- OkayPhysicist 3y agoLike what? I'm getting strong XKCD "Workflow"[0] vibes out of the idea that someone would have unintelligible paths on their system [0] https://xkcd.com/1172/ https://xkcd.com/1172/
- chungy 3y agoI'll just say: consider yourself very lucky to not having encountered it. (Especially if you've spent any time downloading random zip archives from the internet...) I have quite a number of times. There's no one "thing"; it's usually some legacy app that write these names. And when you get a program that refuses to open any non-UTF16 compliant file name, you'll have a time of grief.
- derobert 3y agoSome unarchivers (especially ones which aren't Unix natives, like say rar) seem to love to do it. They're just buggy, of course. Beyond that, before the mid-2000s, it was common to use non-UTF-8 locales on Linux. So I'm sure I still have ISO-8859-1/—15 encoded file names somewhere, especially in archived data. They're not always trivial to rename either, because there might be references to them by name. (Or, in odd cases, you can't convert the name to UTF-8 because you hit a filename length limit, since UTF-8 is more bytes). I believe wanting to access data from 20 years ago is a perfectly reasonable use case. It's not so bad if a program can't display the file name right, as long as it doesn't crash with an exception or refuse to open the file. Unix file names have been defined as arbitrary sequences of octects except / and NUL for 30+ years.
- dataflow 3y ago> Don't try to validate it in any sort of encoding That works fine until it becomes literally impossible when your language/library/framework forces you into picking an encoding. Even Microsoft's own APIs that add "UTF-8 support" choke when they encounter file names with invalid UTF-16 sequences.