4 ms·
In short: Deduplication efforts frustrated by hardlink limits per inode — and a solution compatible with different file systems.
by replooda 6mo ago
In short: Deduplication efforts frustrated by hardlink limits per inode — and a solution compatible with different file systems.
- UltraSane 6mo agoThe real problem is they aren't deduplicating at the filesystem level like sane people do.
- otterley 6mo agoFrom the article: > [W]e shipped an optimization. Detect duplicate files by their content hash, use hardlinks instead of downloading each copy.
- UltraSane 6mo agoI meant TRANSPARENT filesystem level dedupe. They are doing it at the application level. filesystem level dedupe makes it impossible to store the same file more than once and doesn't consume hardlinks for the references. It is really awesome.
- mmh0000 6mo agoFilesystem/file level dedupe is for suckers. =D If the greatest filesystem in the world were a living being, it would be our God. That filesystem, of course, is ZFS. Handles this correctly: https://www.truenas.com/docs/references/zfsdeduplication/ https://www.truenas.com/docs/references/zfsdeduplication/
- UltraSane 6mo agoI was talking about block level dedupe.
- mmh0000 6mo agoI thought you might be. I just wanted to mention ZFS. Have I mentioned how great ZFS is yet?
- otterley 6mo agoZFS is great! However, it's too complicated for most Linux server use cases (especially with just one block device attached); it's not the default (root filesystem); and it's not supported for at least one major enterprise Linux distro family.
- vmilner 6mo agoIt's not as good as ed: https://www.gnu.org/fun/jokes/ed-msg.html https://www.gnu.org/fun/jokes/ed-msg.html
- burnt-resistor 6mo agoFile system dedupe is expensive because it requires another hash calculation that cannot be shared with application-level hashing, is a relatively rare OS-fs feature, doesn't play nice with backups (because files will be duplicated), and doesn't scale across boxes. A simpler solution is application-level dedupe that doesn't require fs-specific features. Simple scales and wins. And plays nice with backups. Hash = sha256 of file, and abs filename = {{aa}}/{{bb}}/{{cc}}/{{d}} where aa = hash 2 hex most significant digits bb = hash next 2 hex digits cc = hash next 2 hex after that d = remaining hex digits
- UltraSane 6mo agoAll good backup software should be able to do deduped incremental backups at the block level. I'm used to veeam and commvault
- burnt-resistor 6mo agoThat costs even more, unreuseable time and effort. It's simpler to dedupe at the application level rather than shift the burden onto N things. I guess you don't understand or appreciate simplicity.
- UltraSane 6mo agoThis article shows it really isn't that simple and is easy to mess up. Who cares if your storage and backup software both dedupe?
- otterley 6mo agoFor ZFS, at least, `zfs send` is the backup solution. And it performs incremental backups with the `-i` argument.
- UltraSane 6mo agozfs send is really awesome when combined with dedupe and incremental