5 ms·
HFS+ is probably the worst filesystem in common use right now; even FAT has the benefit of simplicity. Most of its issues, however, are with its horrific imple
by gchpaco 12y ago
HFS+ is probably the worst filesystem in common use right now; even
FAT has the benefit of simplicity. Most of its issues, however, are
with its horrific implementation; the Unicode naming is kind of bad
but Linus manages to be wrong about several things.
Regarding case sensitivity: it is generally accepted among the user
interface crowd that (Western) users don't really understand that 'C'
and 'c' are different things; they're "both" 'c'. Case-preserving is
thus the accepted practice. However case manipulation is not an
operation that can be done absent a locale; my go to example here is
that 'i' upcases to 'I' unless you're a Turk in which case it upcases
to 'İ'. Similar although not quite as bad is the fact that 'ß'
upcases to 'ẞ' U+1E9E in some exotic circumstances; see
http://en.wikipedia.org/wiki/Capital_ẞ http://en.wikipedia.org/wiki/Capital_ẞ for details. Similar
limitations apply to sorting, which users also expect.
Regarding Unicode: NFD is a normalization format; it converts 'é'
U+00E9 and 'é' U+0065 U+0301--which are semantically identical--into
the same coding. As it happens NFD picks U+0065 U+0301 for that
string; NFC picks U+00E9. Any time there is ambiguity, NF[CD] will
retain the original ordering. Calling it "destroying user data" is
meaningless histrionics. Most of the time we tend to use NFC. I am
told that NFD has certain advantages for sorting, where one might want
to match the French word 'être' with the search string 'etre'; in NFC
this requires large equivalence tables but in NFD the root character
is the same in both cases. Linus's claim 'Even the people who think
normalization is a good thing admit that NFD is a bad format, and
certainly not for data exchange.' has a big [citation needed] tag
attached.
As it happens, my personal belief is the following: Given that users
expect case sensitivity and locale specific ordering, which complicate
filesystem design tremendously. Given that users mostly interact with
the system through GUI dialogs, which already hide system files (files
with the hidden bit in HFS+, or files starting with '.'). Therefore,
extract the case sensitivity to a layer, used by the GUI, which can
understand the user's locale and so fold case properly. This layer
should be available to command line applications so that they can use
the same rules if they so choose. The underlying filesystem will then
be case insensitive, but is still used to encode Unicode data; the
right thing to do here is to normalize. Either NFC or NFD is fine,
really.
For pedants: the related NFKC/NFKD forms add a canonicalization step
and are absolutely not semantically safe in any way, for all that
they're useful for sorting.
- yuhong 12y agoYea, but it would be better to convert filenames from disk at comparison time.
- gchpaco 12y agoThe main reason all this tends to get pushed down to the FS layer is because of the question "what happens if the user accidentally creates two files with the 'same name' but different coding sequences". When it's the difference between README and readme I'm inclined to be like "well, that was dumb, move on with life" but it's different somehow for visually indistinguishable situations.
- Someone 12y agoAnd then, you receive a zip file from Linus that has file.c, File.c and FILE.c files on it. You extract it, and then? They either end up on disk, breaking your case-insensitive UI layer (yes, you can see those files, but can you copy them elsewhere?), or they don't, breaking the makefile that's also in the archive. Here be dragons. Locale-specific ordering of course _must_ be done outside the disk because disks may move between systems with different locales, locales can be changed at will, and multiple users could read the same directory with different active locales (well, must: one could store a locale for sorting per directory and force that on he user, but that is madness) Also, reading http://dubeiko.com/development/FileSystems/HFSPLUS/tn1150.html http://dubeiko.com/development/FileSystems/HFSPLUS/tn1150.ht... (which I can't find anymore on Apple.com), HFS+ doesn't quite use full NFD because it sometimes destroys information that Apple deemed worth keeping.
- tedunangst 12y agoCreating three files named file.c, File.c, and FILE.c seems like being user hostile for the sake of being user hostile. Even for people using a case sensitive filesystem, it's a dick move. Imagine talking about your project over the phone. "Yeah, the problem is in file.c. No, not File.c, file.c. Idiot! I clearly said file.c, not FILE.c!"
- 12y ago