11 ms·
This was discussed to death on the kernel mailing lists, you should go read them. The principal question is whether the tool is more important, or the end user
by james412 6y ago
This was discussed to death on the kernel mailing lists, you should go read them.
The principal question is whether the tool is more important, or the end user. Why did I pay for this machine if it weren't intended to facilitate me? That's the bottom line with most of these kinds of technical "correctness" arguments
And as for whether userspace should catch up, thanks to OS X for the most part that already happened a long time ago for a ton of open source packages
- jstimpfle 6y agoCould you provide pointers to these discussions? As a technical person, case sensitive filesystems have been a loss both from a programmer's and also from an end user's perspective.
- ekr 6y agoHere's Torvalds' view on the matter: https://lwn.net/ml/linux-fsdevel/CAHk-=wg2JvjXfdZ8K5Tv3vm6+bKRedotF5cr5AwVZVBypVfdAQ@mail.gmail.com/ https://lwn.net/ml/linux-fsdevel/CAHk-=wg2JvjXfdZ8K5Tv3vm6+b... I also side with this view, namely that this is something that would be better placed in the userspace rather than the kernel, which really doesn't need more complexity for things that add so little value (negative value to some).
- agwa 6y agoHe must have changed his view because he ultimately allowed the feature. Does anyone know what changed his mind?
- dm319 6y agoWhile Linus kicks off about things, he doesn't tend to outright refuse things. I don't think he sees himself as the gatekeeper of the kernel and that is evident in the way the kernel developed right from the beginning. That has attracted criticism from the likes of Ken Thompson who thought that too much crappy code was allowed into linux. |I've looked at the source and there are pieces that are good and pieces that are not. A whole bunch of random people have contributed to this source, and the quality varies drastically
- asveikau 6y ago> Why did I pay for this machine if it weren't intended to facilitate me? I happen to agree with the idea that the filename should be a dumb blob of bytes and the kernel should not do case folding, as it is the wrong layer for that, eg. the user can change their language but it won't update what has been written to the disk in thousands or millions of places where you could suddenly have a filename collision somewhere based on those rules changing. But, I do hope you get that refund for your Linux.
- kochthesecond 6y ago> dumb blob of bytes Well, now your filename is invalid utf8. How should programs display it or even address such a file?
- jcelerier 6y ago> How should programs display it what's wrong with foo����.txt > or even address such a file? ... by using the array of bytes ?
- ygra 6y agoIt's ambiguous, for example.
- jcelerier 6y agoso are a file named Hello.txt and another one named Нello.txt
- colejohnson66 6y agoThe fact that if one has two files, say “test{invalid bytes}.txt” and test{other invalid bytes}.txt”, both have replacement characters inserted at the same spot and would decode to the same codepoints.
- asveikau 6y agoHow does the UI framework act when you set a label to such payload? How does your web browser act when it sees it in HTML? I have found working on apps that see a lot of usage in varied markets that as much as we wish to see the best and ideal conditions, malformed utf-8 surfaces in the real world pretty often.
- azalemeth 6y agoArguably the biggest difference that I notice when using linux as a desktop environment as an end-user is that it trusts "you", the [sometimes root] user, to a far greater extent than other operating systems. It is for this reason that I enjoy it. It also means that if you want to do something highly annoying and unusual, you can – and arguably case-insensitive filesystems are a subset of that.
- snazz 6y agoI agree that case-insensitive filesystems should be an esoteric feature, but given that they’re the default on Windows and macOS, it should definitely be a well-supported option on Linux for the sake of compatibility.
- magicalhippo 6y agoI've set the Samba shares on my NAS to be case-sensitive, as making them case-insensitive slows down directory access by orders of magnitude. I've been running this for years accessing them both from Linux desktops and Windows desktops, and only once have I had an issue that required me to manually rename something on the NAS. This makes sense as most applications don't care about the filename, and will just use what you supply, or generate one and use that string all over.
- jra_samba 6y agoYep, that's true. It's the cache misses that kill performance. If the client asks for file "Foo", and the (l)stat fails to find it, then we have to scan the whole directory looking for any case-differing versions of "foo" "FOO" "fOo" etc. Very costly, but the only way to give case-insensitivity.
- jimsmart 6y ago> Very costly, but the only way to give case-insensitivity. That is very costly — but it's certainly not the only way to provide case-insensitivity, nor is it the recommended way, and I'd be surprised if any implementation of case-insensitivity actually did what you say. Normally one would lowercase (or uppercase) both strings, and then do the comparison. The complexity here usually comes when case-folding various tricky locales.
- akira2501 6y ago> That's the bottom line with most of these kinds of technical "correctness" arguments The problem is emergent behavior. We can create any number of features, but we have a really hard time testing all the available configurations that result. Engineers rely on simplicity as a way of warding off this particular problem, because the other bottom line is people don't want to pay for a system that loses or invalidates the work they've put into it.
- ploxiln 6y agoThis feature doesn't facilitate you at all. It's a historical mistake that macOS and Windows have preserved, and refined a bit over the years. The only real purpose of this is for easier/faster compatibility with software developed and tested on macOS and Windows, which has accidental case inconsistency in file name references in the code, which happens to work fine on macOS and Windows. TFA could have made that argument, but it didn't, it made incorrect arguments instead. For example, a user might type in a lower-case name for a file one time, and a capitalized name for a file another time, and intend to access the same file. But what user is typing in a whole filename the second time? They're picking it from a list, or if they're very advanced, using completion in a terminal. TFA also mentions non-English languages, in the context of unicode normalization ... but non-western-european languages won't be handled correctly by any universal case-folding algorithm anyway, with Turkish being the most common example. The whole strategy just doesn't work out well. Many low-level filesystem developers have known this for over 20 years. It doesn't work better for anyone, but non-technical people just aren't aware of why or how it increases complexity and costs, and reduces performance and reliability.
- efdee 6y agoSurely the mistake was case sensitive filenames. I can't think of a single end-user use case where this behavior is desirable.
- msla 6y agoEnsuring filenames don't get destroyed by an OS that refuses to understand a given language. Case is a complicated mess once you leave ASCII, and that's partially because ASCII is lying to you about how English case works: Yes, the English language has title case, and ASCII conflating it with upper-case does not negate that. Move on to most anywhere else and the notion that it's fast, reliable, and safe to convert case gets lost in the realities of human writing systems.
- arghwhat 6y agoIn a world of ASCII, maybe. But fixing one problem for a small group of people is a giant can of worms for the rest of the world. Normalization of compound characters, exotic character sets, emoji, different classes of upper/lower-case letters, normalization of compound what-not. And even then it still doesn't fix the issue outside of the world of ASCII. A filename written in hiragana and katakana is logically the same to the end-user, but they are still distinct. Simplified and traditional Chinese, Hangul and romaja, pinyin, devanagari, thai, and the list goes on. Case-insensitive filenames fixes nothing, but breaks everything. There is only one sensible thing to do with arbitrary user-input, and that is to leave it be.
- igetspam 6y agoWhich has caused countless problems with interpreted languages and cross platform functionality. Things lime ruby in OSX will gladly less you mangle your include strings on OSX, which causes a "works fine in Dev" problem. I'm not a fan of case insensitive filesystems because I have to manage services.
- alerighi 6y agoI disagree. The error is to consider paths as a high level information, that the user has to know about, rather of what they really are, a low level information, that potentially the user never sees (for example consider mobile operating systems like Android/iOS). In practice the case insensitive thing if we want to call it should be implemented more high level, in the file manager, rather than in the filesystem itself. That is even what newer versions of Windows/NTFS do! Recent versions of NTFS are in fact case sensitive, and if you mount a NTFS volume with Linux in fact you can create two file with the name differing only by the case: the whole case-insensitive thing is handled at an higher level in the Windows APIs.
- dietr1ch 6y agoIt's not enough to have similar matching at the fs level, applications need to have the same functionality around, otherwise anything that indexes the fs will have false negatives before reading the files. Now, if this needs to be taken care of at the application level, then why have this misfeature? It'd be better to have a good library for matching this that could be aware of the language and locale (or maybe multiple lang,locale pairs) instead of throwing this feature into the fs and call it a day. Also, if some applications benefit from having insensitive matching, like things built to run on Windows/Mac, then having a wrapper that fixed the fs access with this matching library would be enough, no need to force other applications to use insensitive matching because a single one needs it.
- ComputerGuru 6y ago> And as for whether userspace should catch up, thanks to OS X for the most part that already happened a long time ago for a ton of open source packages That’s a bold claim. Open source developer here (on the fish team): given a user-entered path, how do you make it canonical? Specifically, without recursively opening directories and listing contents, how can you convert /foo/baR to the correct case as it appears in the filesystem index? Without following any symlinks, should any of the path components be a symlink?