4 ms·
> Of what use is a filesystem with many empty files? You could use it to benchmark things that do things to all the files in a directory tree. For example, what
by ChrisSD 3y ago
> Of what use is a filesystem with many empty files? You could use it to benchmark things that do things to all the files in a directory tree. For example, what is the fastest way to list all those files? Delete them? Back them up? Create an archive file with all of them?
This can be useful up to a point. But be careful of over optimizing for the pathological case at the expense of typical usage. Most of the listed operations are usually done on a small number of files. And even when they seem a lot to the user it's still a teeny tiny fraction of one billion.
And that's before we even mention directories.
- dspillett 3y agoAlso, be wary of making general conclusions about the performance of a large operation over zero-length files. There could be a significant gap between that and dealing with 1-byte files because of block allocations for that data. Some filesystems support inlining whereby small files do not require extra allocation⁰, that allocation could be very local (in the file's only inode) or not (elsewhere in the MTF in NTFS, so more local than “anywhere on disk”), so there can be a performance profile change when you hit the inlining limits instead of or as well as one when jumping from nothing to a single byte. There are circumstances where a great many very small files are present: a mail archive (though with the amount of headers in a modern SMTP delivered message for host transit tracking and spam/other identification process notes, these files are not as small as they once were), similarly a usenet archive, a web cache, source repos, …. Also I've seen some tools use small files as per-user flags or session notes, where that don't want the extra dependency of a DB layer, to parse/update a more complex structure for on-disk updates, or the cost of serialising an in-memory structure in a large regular write. This results in very small or even zero-length files, and could balloon in numbers with a great many users. But it is getting more common for such things to be in a sqlite DB or similar instead. -- [0] Ext4 can inline minute files, 60 bytes or less, in the inode if the relevant option is enabled¹ on filesystem creation (IIRC it can't be turned on after the fact), NTFS can inline small files (I think about 600 bytes) in the MTF structure. [1] It isn't enabled by default because of a potential rare issue that results in the chance of corruption during recovery after an unclean unmount (due to power drop for instance). IIRC the conditions are: if you create a small file that is inlined, then append to it enough that it can no longer be inlined so a block is allocated, and the unclean unmount happens immediately after, the new data may be lost.
- patwolf 3y agoIt does seem certainly seem atypical, and I'd be curious if anyone has ever encountered something like this in the real world. It seems unlikely because 1. It takes so long to create that many files in the first place. In this case 26 hours of doing nothing but file creation. 2. If the files contained any data at all, you would need a lot of disk space. If each file was 64k, you'd need 640TB of disk. It seems more likely you'd run out of disk space before you reached a billion files. 3. Someone developing such a system where that many files could be created would probably realize that it's a poor design and figure out a way to mitigate it before it ever happened. Nevertheless, it's fun to think about.
- lanthade 3y agoBack in the late 2000’s I was a windows clustering SME for a major tech company on account with a major financial company. One project I was tasked with was migrating a system which had a huge number of small files. I don’t recall the number. IIRC it took a couple weeks to do the initial file copy. Then the task was to sync the changes during the change window which was 8 hours. Every time I tried to sync the changes it took well more than 8 hours. I did find a tool that could do the job but of course no one wanted to pay for it. Fortunately for me I left that position before the migration and someone else had to figure it out. So yeah, I’ve encountered something similar in the real world. It was the result of poor software design choices but by the time I got there that path was chosen and I had to deal with the results.
- vlovich123 3y ago10k file creations per second does seem kind of slow no? A typical SSD today can do many millions of operations per second…
- yjftsjthsd-h 3y ago> and I'd be curious if anyone has ever encountered something like this in the real world. Not that many, no, but I can tell you from professional experience that real life Linux systems get really weird when you have merely tens of thousands of files in a directory, and it would be kind of nice if that were fixed.
- tyingq 3y agoIt does come up in real world situations. Lots of file based caches in things like Apache, or session stores, like in PHP, etc.