5 ms·
I've read the project's description and still failed to understand what it does and what it is useful for.
by mathfailure 4y ago
I've read the project's description and still failed to understand what it does and what it is useful for.
- ecnahc515 4y agoThe main advantage is the content addressable part. Existing overlay filesystems only handle the overlay aspect, and then tools like docker/containerd attempt to reuse layers efficiently, but it's not perfect. The same files from two different layers may have the same content, but it's still stored twice because the layers are the "unit" of storage roughly speaking. By making a single filesystem which handles both content addressability and the overlay aspect, you can avoid duplicating files that are the same, but in different layers.
- PlutoIsAPlanet 4y agoAnother way of describing it, rather than a Docker/Container image being a group of layered archives, each with changes, instead a list of file hashes is distributed, detailing where those files need to be in a mounted filesystem, with x permissions. Since everything is named based on hashes, content is naturally deduplicated if two images share the same files and all files are stored in the same place. If you boot on top of a composed filesystem, you also get easy file verification as long as the booted list is signed and unmodified. If you modify the local files, the hashes won't match.
- __MatrixMan__ 4y agoImagine a workflow: - clone a repo - run a command - run another command It runs several times daily. Maybe it's CI or something. Now suppose you want to cache the filesystem state for each command so that they can be rerun in a debug scenario where you'd expect them to behave the same as they did the first time because they have the same filesystem. (Having then recreated the bug, you could then start making changes towards a fix). You either end up with many many copies of that repo, or you use something like this to only store the unique files and instead have many many indices into that store.
- sluongng 4y agoGit has 2 object stores: the loose object store and the packed object store. What you said is applicable to the loose object store where a full copy of the file is stored as plain text. Those could be deduplicated quite nicely, and git does just that, loose object store is a CAS. However the packed object store is trickier. It stores duplicated object in a delta compressed format plus gzipped. So deduplicating packfiles on a file system level is almost never worth it. Git is moving toward using packed object store more and more. With some of the latest patches, you can effectively use git with very little loose object storage ultilization (zero if you are on a server hosting git repositories).
- __MatrixMan__ 4y agoHmm that's good to know, thanks. It still just applies to files that are handled by git though. If your workflow applies a patch and then invokes a compiler which generates intermediate files, both the post-patch file and the intermediate files will not end up in the packed object store. So if you're taking filesystem snapshots of those states you'll still want to deduplicate them some other way.