5 ms·
Can someone explain how tmpfs / /dev/shm track files if not by inode number?
by cpcallen 5y ago
Can someone explain how tmpfs / /dev/shm track files if not by inode number?
- esjeon 5y agoPlease read the article. TMPFS was internally using 32bit unsigned int for counting inode number, which got overflow-ed. Kernel already has 64bit wide ino_t, the dev migrated to it.
- puppet-master 5y agoThe question was clear enough: per the article tmpfs is able to track two distinct files with the same inode number, implying it does not use the inode internally to differentiate between files
- Tobu 5y agoI'm not sure where the confusion is coming from. If userspace wants to distinguish files, it normally does so with a (ino_t, dev_t) pair. With some nuance wrt inode generations if you aren't holding on to files yet want to guard against ino reuse, and some funkyness with overlayfs. Most filesystems don't expose a way to access files by ino only, so the internal implementation generally doesn't key anything by them or rely on them itself; they're just a kind of of uniqueness cookie. But keying by them somewhere in the kernel implementation is feasible, if you scope it properly.
- adrianmonk 5y agoIt may be that, in modern times, filesystems don't actually key things with i-node number. But wasn't that the original purpose? Otherwise, why the design where directories are basically a mapping of names to i-node numbers? Also, the article does say this: > On non-volatile filesystems, inode numbers are typically finite and often have some kind of implementation-defined semantic meaning (eg. implying where the inode is disk). If you've got other means to know where the i-node is on disk, then why would a filesystem bother encoding that into the i-node numbering scheme? Maybe the article is wrong here. I don't know. But if the question is where the confusion is coming from, the article definitely seems to be painting a picture where normal filesystems look up things based on i-node number and tmpfs is relatively unique in not doing so.
- formerly_proven 5y agoYes, the OG Unix file system simply had a fixed-size (at partition time, this is still the case for ext4 and XFS iirc) array of struct inode on the disk, and the inode number (iirc it was just idx or something like that in the code) is simply the index into that array. An OG Unix directory was simply a file with the directory flag set and the file was just an array of struct { char name[30]; int ino; } (~something like that). You actually used to be able to just open() a directory and directly read directory entries (dirents henceforth) from it, just like a file - exactly like a file, because it WAS a file. ext and XFS still largely work like this, though directories are now hashtables and I think XFS supports multiple independent arrays for inodes (or just for extent allocation? I don't remember). NTFS also looks a lot like a Unix file system with some weird growths on it, and not a whole lot like FAT.
- yyyk 5y agoThere's also some funkyness with btrfs subvolumes: https://news.ycombinator.com/item?id=28274054 https://news.ycombinator.com/item?id=28274054
- rav 5y agoSince the contents are in-memory, you could imagine a straightforward implementation where the directory hierarchy in the tmpfs is stored in a pointer-based in-memory tree. A file handle just needs a pointer to where in memory the contents are stored, and the inode numbers are computed on the side because it's part of the interface. You don't want to use the pointers as inode numbers directly, since for security reasons the kernel doesn't want to export valid kernel pointers to user space.
- yencabulator 5y agoTo play with these concepts, you can write a FUSE filesystem where each file has inode 42, while being wholly separate nodes & handles. Linux VFS doesn't really use the inode number internally for anything while a direntry is alive.
- pengaru 5y agoSince there's no need for these inode numbers to survive a reboot, the inodes are simply generated on-the-fly for a given tmpfs mount, and not persisted anywhere durable. TFA makes this clear...
- formerly_proven 5y agoinodes are actual objects in the kernel. Their inode number is supposed to be such that (device, inode, generation) uniquely identifies it, but the kernel itself doesn't really care about this (as long as your filesystem is not using the inode cache, which does use the inode number as a key). The way tmpfs works is that it uses the dirent cache as it's actual data structure (forcing the refcount of all cache entries >1 means they can't be evicted); those just directly contain pointers to inode objects. File data is stored in the page cache.