6 ms·
Does NILFS do checksums and snapshotting for every single file in the system? One of my biggest complaints about file systems in general is that they are all de
by didgetmaster 4y ago
Does NILFS do checksums and snapshotting for every single file in the system? One of my biggest complaints about file systems in general is that they are all designed to treat every file the exact same way.
We now have storage systems (even SSDs) that are big enough to hold hundreds of millions of files. Those files can be a mix of small files, big files, temp files, personal files, and public files. Yet every file system must treat your precious thesis paper the same way it treats a huge cat video you downloaded off the Internet.
We need some kind of 'object store' where each object can be given a set of attributes that govern how the file system treats it. Backup, encryption, COW, checksums, and other operations should not be wasted on a bunch of data that no one really cares about.
I have been working on a kind of object file system that addresses this problem.
- nix23 4y agoWell you can do that kind of with zfs filesystems, and the "object" is the recordsize.
- mustache_kimono 4y agoI was going to ask: "Is there any limit on the number of ZFS filesystems in a pool?" Google says 2^64 is the limit. Couldn't one just just generate a filesystem per object if snapshots, etc., on a per object level is what one cared about? Wonder how quickly this would fall over? > Backup, encryption, COW, checksums, and other operations should not be wasted on a bunch of data that no one really cares about. This GP comment is a little goofy though. There was a user I once encountered who wanted ZFS, but a la carte. "I want the snapshots but I don't need COW." You have to explain, "You don't get the snapshots unless you have the COW", etc.
- Conan_Kudo 4y agoOn Btrfs, you can mark a folder/file/subvolume to have nocow, which has the effect of only doing a COW operation when you are creating snapshots.
- mustache_kimono 4y agoAnd that may work for btrfs, but again at some cost: "When you enable nocow on your files, Btrfs cannot compute checksums, meaning the integrity against bitrot and other corruptions cannot be guaranteed (i.e. in nocow mode, Btrfs drops to similar data consistency guarantees as other popular filesystems, like ext4, XFS, ...). In RAID modes, Btrfs cannot determine which mirror has the good copy if there is corruption on one of them."[0] [0]: https://wiki.tnonline.net/w/Blog/SQLite_Performance_on_Btrfs#Test_2:_Using_nocow_option https://wiki.tnonline.net/w/Blog/SQLite_Performance_on_Btrfs...
- lazide 4y agoYup. It’s a pretty fundamental thing. COW and data checksums (and usually automatic/inline compression) co-exist that way because it’s otherwise too expensive performance wise, and potentially dangerous corruption wise. For instance, if you modify a single byte in a large file, you need to update the data on disk as well as the checksum in the block header, and other related data. Chances are, these are in different sectors, and also require re-reading in all the other data in the block to compute the checksum. Anywhere in that process is a chance for corruption of the original data and the update. If the byte changes the final compressed size, it may not fit in the current block at all, causing an expensive (or impossible) re-allocation. You could end up with the original data and update both invalid. Writing out a new COW block is done all at once, and if it fails, the write failed atomically, with the original data still intact.
- tjoff 4y ago> Chances are, these are in different sectors, and also require re-reading in all the other data in the block to compute the checksum. Anywhere in that process is a chance for corruption of the original data and the update. Not much different than any interrupted write though. And a COW needs to reread just as much. > If the byte changes the final compressed size, it may not fit in the current block at all, causing an expensive (or impossible) re-allocation. Something that you must always pay in a COW filesystem anyway? Is handled by other non-COW filesystems anyway. Just because a filesystem isn't COW doesn't mean every change needs to be in place either. Of course, a filesystem that is primarily COW might not want to maintain compression for non-COW edge-cases and that is quite reasonable.
- lazide 4y agoWhy is it ‘wasted’? Those things are mostly free on modern hardware. The challenge with your thesis here is that the only one who can know what is ‘that important’ is YOU, and your decision making and communication bandwidth is already the limiting factor. For many users, that cat video would be heartbreaking to lose, and they don’t have term papers to worry about. So having to decide or think what is or is not ‘important enough’ to you, and communicate that to the system, just makes everything slower than putting everything on a system good enough to protect the most sensitive and high value data you have.
- didgetmaster 4y agoNothing is free or even 'mostly free' when managing data. Data security (encryption), redundancy (backups), and integrity (checksums, etc.) all impose a cost on the system. Getting each piece of data properly classified will always be a challenge (AI or other tools may help with that), but it would still be nice to be able to do it. If I have a 50GB video file that I could easily re-download off the Internet, it would be nice to be able to turn off any security, redundancy, or integrity features for it. I wonder how many petabytes of storage space is being wasted by having multiple backups of all the operating system files that could be easily downloaded from multiple websites. Do I really need to encrypt that GB file that 10 million people also have a copy of? Am I worried if a single pixel in that high resolution photo has changed due to bit rot?
- Arnavion 4y ago>Do I really need to encrypt that GB file that 10 million people also have a copy of? Indeed you don't. Poettering has a similar idea in [1] (scroll down to "Summary of Resources and their Protections" for the tl;dr table), where he imagines OS files are only protected by dm-verity (for Silverblue-style immutable distros) / dm-integrity (for regular mutable distros). [1]: https://0pointer.net/blog/authenticated-boot-and-disk-encryption-on-linux.html https://0pointer.net/blog/authenticated-boot-and-disk-encryp...
- lazide 4y agoWorried if it gets lost or mangled? Not necessarily. Worried if it’s happening and I have no idea, and it’s spreading - due to hardware or software problems? And I’ll only discover it when something I do actually really care about and can’t easily replace? Absolutely! The big issue regarding duplication is really more a identification/‘supply chain’ issue. The last thing I want to be doing is trying to figure out how to get that other file from somewhere (that works), from whoever ‘has another copy’, when it gets mangled and I need a replacement ASAP. If you’re thinking of our local filesystems as potentially just a cache, we have no easy or secure way right now to fingerprint or recreate the other entries in the cache from sources (minus browser caches or the like, but even then, re-retrieving it may return different or not content). So backups are really keeping the computing equivalent of local ‘ability to manufacture’ handy as a mitigation against risk. It’s not waste, anymore than keeping the ability to manufacture it’s own tanks and weapons in-country is a waste for a nation. It’s an insurance cost against real world problems, and causes a moral hazard with other actors if not done. And CRC32/CRC32C (even in Java!) is > 16GB/s per core in modern processors, and more than adequate for block sized (typ. < 32MB) checksums. Blake3 is > 2GB/s per core, and is more than adequate for… well every use case we’re currently aware of. Modern OS’s really don’t have any excuses for not doing it.
- spookthesunset 4y agoIt might sound weird but the hard part of what you describe is not the technology but how to design the UX in a way that you aren’t babysitting everything. And doing that is not at all easy. For all anybody knows your cat video is “worth more” to you than your thesis paper. How can you get the system to determine the worth of each file without manually setting an attribute each time you create a file? And if you let the system guess, the cost of failure could be very high! What if it decided your thesis paper was worthless and stored it will a lower “integrity” (or whatever you call the metric)? I dunno. Storage is getting cheaper all the time and it might just be easier to fuck it and treat all files with the same high level of integrity. Maybe it would be so much work for a user to manually manage they’d just mark everything the same?
- didgetmaster 4y agoYou could always set the default behavior to be uniform for all files (e.g. protect everything or protect nothing) and just forget about it. But it would be nice to be able to manually set the protection level for specific files that are the exception. If I was copying an important file into an unprotected environment, I could change how it was handled (likewise if I was downloading some huge video I didn't care about into a system where the default protection was set to high). I agree that if you have 100 million files, then it could be nearly impossible to classify every single one of them correctly.
- spookthesunset 4y agoI’d think on a directory basis would be the ideal
- nintendo1889 4y agoA directory basis, or even better, a numerical priority that could be manually set in the application that generated them, or automatically, based on the user or application or in a hypervisor, based on the VM. Then it could be an opportunistic setting. I thought ZFS had some sort of unique settings like this.
- heavyset_go 4y agoYou can turn off CoW, checksumming, compression, etc at the file and directory levels using btrfs.
- Arnavion 4y agoIndeed. You can also make a directory into a subvolume so that that directory is not included in snapshots of the parent volume.
- llanowarelves 4y agoI have been spinning my wheels on personal backups and file organization the last few months. It is tough to perfectly structure it. I think directories or volumes having different properties and you having it split up as /consumer-media /work-media /work /docs /credentials etc may be the way to go. Then you can set integrity, encryption etc separately, either at filesystem level or as part of the software-level backup strategy.
- rodgerd 4y ago> Does NILFS do checksums and snapshotting for every single file in the system? NILFS is, by default, a filesystem that only ever appends until you garbage collect the tail. It doesn't really "snapshot" in the way that ZFS or btrfs do, because you can just walk the entire history of the filesystem until you run out of history. The snapshots are just bookmarks of a consistent state.