3 ms·
ZFS Dedup has been wonderful for me : dedupratio = 7.05x (144 GB stored on a 25 GB volume, and still 1.3 GB left free). I use it for backups of versions of the
by mlok 5y ago
ZFS Dedup has been wonderful for me : dedupratio = 7.05x (144 GB stored on a 25 GB volume, and still 1.3 GB left free).
I use it for backups of versions of the same folders and files slowly evolving over a long period of time ( > 15 years) that gives a lot of duplication, of course. (I could also use compression on top of it)
- infogulch 5y agoBackups are the 99% case for duplicate files, but aren't snapshots a better replacement in every way? Snapshots are already deduplicated as soon as you take them, plus they're instant. Maybe if your backups are coming from a non-zfs system, but you could probably convert normal backups into snapshots without too much trouble. Why is dedup even present when the primary use case (backups) is better served in every way by snapshots?
- toast0 5y agoDedup (if it worked like it might have!) could solve the backup use case without needing to dictate your workflow. In theory, it could also really help with virtualized disk workloads where there may be a lot of duplicated data from the base OS, but you can't use a snapshot (easily) because windows won't run from a zfs filesystem. You could maybe do snapshotting on zfs volumes, but that's not as flexible as a dedupe that worked as imagined. Personally, I think online dedupe ends up being too expensive in memory and computation and ends up missing things because of divergent block sizes or offsets as another poster mentioned. ZFS doesn't support an offline dedupe, but I think btrfs does. That might be more interesting. It's still expensive to find duplicates, but it's possible, and it'd be neat to be able to rewrite the metadata to refer to a single copy and free some space, maybe.