3 ms·
I've apparently used 1241165 out of 19005440, or about 6.53% on my development machine. It's not inconceivable to have 19 million files, though, if you had one
by bArray 10y ago
I've apparently used 1241165 out of 19005440, or about 6.53% on my development machine. It's not inconceivable to have 19 million files, though, if you had one per user. That's not great.
How is that number calculated? It's less than 2^25 but greater than 2^24, so I imagine it's even awkward to store. I guess that has something to do with the file table size for HDD - you can't have more nodes than you have addressable space for those nodes.
I knew Git did it for speed, but I was wondering whether it was also about the number of files that can exist in a single directory.
EDIT: Corrected percentage.
- LukeShu 10y agoIt's a property of the filesystem that is set when it was created. In the case of the mkfs.ext{2,3,4}, you can set the number of inodes with the -N flag, but by default, it is the size of the filesystem divided by inode_ratio (probably 4096).
- bArray 10y agoSeems to be about 16k for the ratio? 19,000,000 * 16,000 ~= 300GB - which is about the size of my disk partition. That's crazy. Thanks for that information.
- danieldk 10y agoYour filesystem's block size is probably 4KiB. So, it's one inode per for blocks. That's crazy. Why? Do you expect to store many files less than 16KiB?
- bArray 10y agoEach user has some small arbitrary information associated with them, with potentially tens of millions of users (and no doubt more in the future). There's not really a structure due to the nature of the project, each user needs personalised information and structure associated to them. I don't think that amount of information is really suitable for a database? It may be larger in some cases too, so there's not guarantees I can even make about size.
- danieldk 10y agoAh, sorry, I missed that you were still referring to the 'filesystem as a database'-project :).
- Xylakant 10y agoDocument stores shine in that case. Depending on the exact case something like postgres's HSTORE, couchdb, couchbase, mongodb or similar would work. They're all capable of storing arbitrary json docs under a given key and efficiently retrieve it. Partial indexing is possible as well.
- deleted 10y ago[deleted]
- caf 10y agoThis seems to be one of those bits of old UNIX lore that's being lost - it used to be fairly well-known in sysadmin circles that if you were going to be storing a lot of tiny files, you should create your filesystem with a higher than usual number of inodes (the classic case was an NNTP spool mount).
- bArray 10y ago"This seems to be one of those bits of old UNIX lore that's being lost" Hopefully not now, this is why we all appreciate people sharing these sorts of information :)
- gioele 10y ago> if you were going to be storing a lot of tiny files, you should create your filesystem with a higher than usual number of inodes (the classic case was an NNTP spool mount). I wonder how many people understand what the term "news server" of the partition wizard in the Debian installer means.
- technofiend 10y agoI literally just solved an issue this weekend that was inode exhaustion: there were years of snapshots and traces laying around and the DBAs were confused as to why the database couldn't create a directory to start and was exiting with "No space left on device" when there was 5% storage free. That lead to the second discussion of why you don't fix it by piping filenames one by one to rm, but either use find's inbuilt unlink/rm if there is one or xargs to rm if not.
- greggyb 10y agoIs that second discussion just a performance thing, or are there filesystem implications for one or the other?
- technofiend 10y agoJust speed: you don't spin up a new process for each file removed if you use xargs.
- fnj 10y agoThat's 6.53%, not 0.0653%. Bit of a difference.
- bArray 10y agoForgetting to times by hundred will eventually kill me, I swear. I'll update it!
- danieldk 10y agoHow is that number calculated? It's less than 2^25 but greater than 2^24 It depends on the filesytem, but it is typically a function of filesystem size (and in ext, the number of block groups and inodes per block group size). To take an example: sudo tune2fs -l <device> | egrep -i 'group' Blocks per group: 32768 Fragments per group: 32768 Inodes per group: 8192 So, we have got 1 inode per 4 blocks. This particular filesystem has 79472640 blocks (of 4KiB) and 19873792 inodes. This corresponds to the 4:1 ratio: 79472640 / 19873792 =~ 4.
- bArray 10y agoOkay, that makes sense - thanks! I guess there's an upper limit to that number too? It's probably squeezing into a 32 or 64 bit number?
- danieldk 10y agoThere is probably a upper limit, but practically you don't want more inodes than blocks. Since a non-empty file at least takes up one block, assuming that your files are all at least 1 byte, you couldn't use more inodes than blocks anyway. (The exception are empty files, because on most filesystems an empty file does not use any data blocks, but possibly an inode to store file size, permissions, etc.)
- majewsky 10y agoIf I'm not mistaken, some file systems store file contents smaller than a block directly inside the respective directory entry.