6 ms·
This issue is what makes me paranoid when putting a lot of files into a directory. Having directories whose sizes cannot be shrunk makes me really uncomfortable
by barosl 5y ago
This issue is what makes me paranoid when putting a lot of files into a directory. Having directories whose sizes cannot be shrunk makes me really uncomfortable, so I try to avoid such a situation as much as possible. What's unfortunate is that you cannot predict at what point a directory will grow above its initial size, because the size of a directory is affected by not only the number of files, but also the varying lengths of their filenames. It's a complete mess.
I wondered if NTFS on Windows is also not capable of shrinking already-grown directories but could not find information about it. As Windows doesn't report the size of a directory itself, it is hard to test the situation.
- nayuki 5y agoIt's hard to shrink the Master File Table (MFT) on Windows/NTFS. Each file or directory consumes 1 KiB in the MFT. So this is similar to the ext* directory-shrinking problem, but at the level of the entire volume instead.
- lisper 5y agoThe real problem here is that the unix file system is kinda sorta like a database but not really. So people try to use it like a database and it kinda sorta works, but not really. Someone ought to write a clean-sheet OS with an embedded copy of SQLite built in to the kernel. That would kick some serious tushy.
- xxpor 5y agoThat didn't go very well for MS back in the mid 2000s
- lisper 5y agoHuh? What are you referring to?
- croes 5y agoProbably WinFS https://en.m.wikipedia.org/wiki/WinFS https://en.m.wikipedia.org/wiki/WinFS
- lisper 5y agoOK, yeah, that's not quite the same as what I'm suggesting. WinFS was intended to be used at the application level. I'm talking about using SQLite (or something like that) to store filesystem metadata, more like the resource fork in the original MacOS, except that the resource fork was per-file and what I'm suggesting here would use the embedded DB to store directories (in addition to per-file metadata). The schemas would be part of the OS design. Applications would not be able to modify them or add new ones.
- lmm 5y agoYou can deploy your application as a unikernel and if you don't need a filesystem then you don't have to include one. I really think that's the future.
- smallstepforman 5y agoBeOS has the Be File System, a light database on top of a filesystem, and you can query to your hearts content. https://web.archive.org/web/20170213221835/http://www.nobius.org/~dbg/practical-file-system-design.pdf https://web.archive.org/web/20170213221835/http://www.nobius...
- Spooky23 5y agoThere are other considerations with Windows as well - if you’re using SMB, large file counts in a directory will create performance issues with SMB shares.
- magicalhippo 5y agoMaybe you're referring to something else, but what I noticed was due to case sensitivity. Turning off case sensitivity lead to orders of magnitude difference in directory performance, and since most applications just use what they get from the system for filenames, there's very few problems in practice. My main machines are all Windows and I've been running my NAS with case sensitivity off for almost a decade now, and only a few times did I have to manually rename some file through the NAS (two files with same name but different case). I use my NAS actively for a lot of things, including sharing files across my machines.
- southerntofu 5y ago> Turning off case sensitivity lead to orders of magnitude difference in directory performance What filesystem are you using? I assume case-insensitivity means your filesystem does not support UTF-8 filenames. Is that the case?
- zaarn 5y agoNo, case insensitive just means that the filesystem considers uppercase letters and lowercase letters to be the same. You have to be Unicode (or character set) aware for that. You can set ZFS and in newer Linux Kernels for a filesystem to be case insensitive, and neither really cares about UTF8 to begin with as long as the filename contains no NUL characters. Windows only requires the filename to be somewhat valid UCS-2 (ie, UTF-16 with the safeties off) on NTFS, FAT does the same for ASCII (though nothing stops a kernel from putting UTF8 in a FAT filename.
- southerntofu 5y ago> No, case insensitive just means that the filesystem considers uppercase letters and lowercase letters to be the same. You have to be Unicode (or character set) aware for that. I assumed in your case that meant ASCII encoding, but i still don't understand how turning off case sensitivity would speed up things. Was that a typo, or am i missing something? > nothing stops a kernel from putting UTF8 in a FAT filename Except interoperability with other systems who may access this filesystem of course. Thanks for this explanation.