4 ms·
Does anyone actually use dedup? I think even the OpenZFS documentation says compression is more useful in practice. If at all, dedup should be an offline feat
by FullyFunctional 5y ago
Does anyone actually use dedup? I think even the OpenZFS documentation says compression is more useful in practice. If at all, dedup should be an offline feature, to be run as scheduled by the operator.
My setup tries to get the absolute highest bandwidth and uses NVMe sticks in a stripe (I get my redundancy elsewhere), no compression, no dedup and yet can only hit ~ 3.5 GB/s reads (TrueNAS Core, EPYC 7443P, Samsung 980PRO, 256 GiB). I hope TrueNAS SCALE will perform better.
- bombcar 5y agoIt would be nice if ZFS was able to combine dedup and compression - basically be able to notice that a block/file/datastream was similar/identical to another one, and do compression along with a pointer ...
- willis936 5y agoZFS can have both features enabled at once. Though there is no clean way to disable either. Compression can be removed from files by rewriting them, but removing deduplication requires copying over all data to a fresh pool.
- nightfly 5y agoYou should be able to remove the deduplication data by zfs send/recv within the same pool too, I think
- willis936 5y agoNot according to people who know more about ZFS than me. https://www.truenas.com/community/threads/zfs-dedup-disable-cleanup.98917/ https://www.truenas.com/community/threads/zfs-dedup-disable-...
- nightfly 5y agoSo at a pool level you might not be able to turn it off once it's turned on, but you can also turn off deduplication per file system, including in properties you set when receiving a stream. I wasn't confident this would work, but a test proved it can. (chicken_test/dedup_source had deduplication enabled and 16 copies of the same 100MiB file) chicken:~# zpool list NAME SIZE ALLOC FREE CKPOINT EXPANDSZ FRAG CAP DEDUP HEALTH ALTROOT chicken_test 15G 144M 14.9G - - 0% 0% 16.00x ONLINE - chicken:~# zfs send chicken_test/dedup_source@send | zfs recv -o dedup=off chicken_test/nodedup_dest chicken:~# zpool list NAME SIZE ALLOC FREE CKPOINT EXPANDSZ FRAG CAP DEDUP HEALTH ALTROOT chicken_test 15G 2.29G 12.7G - - 0% 15% 16.00x ONLINE - chicken:~# zfs get dedup chicken_test/nodedup_dest NAME PROPERTY VALUE SOURCE chicken_test/nodedup_dest dedup off local chicken:~# zfs destroy -r chicken_test/dedup_source chicken:~# zpool list NAME SIZE ALLOC FREE CKPOINT EXPANDSZ FRAG CAP DEDUP HEALTH ALTROOT chicken_test 15G 2.29G 12.7G - - 0% 15% 1.00x ONLINE -
- lazide 5y agoZfs send/recv sends the blocks as written to the original filesystem (which is why it can be so fast, it doesn’t have to ‘understand’ what it happening or defragment things to read like reading a file does), but that also means undoing or applying dedup won’t work correctly unless it’s screwing with things you probably don’t want it too. One issue I had is that due to what I eventually tracked down as power issues, I had some corrupted data written to disk under my zfs pool (at the media write later), and I had dedup on. So dedup, unfortunately, actually made it REALLY suck to fix, because I couldn’t even copy a new version of the file to the same pool! It kept nuking the duplication, and keep the old bad data and I then couldn’t read the copy. :s It even did this after I deleted everything, because prune couldn’t remove the bad underlying entries because it was having a media failure. So delete files, scrub, put new files on resulted in them having the exact same failure. When I nuked the pool and recreated it, it was all fine though. So yeah, be careful with dedup.
- nightfly 5y agoZfs send/recv actually does send data at a logical level, unless instructed otherwise. There are options to send deduplicated streams, streams maintaining compression, and raw streams but none of those are the default. Also, see my reply to a sibling comment.
- lazide 5y agopractically speaking the tradeoffs to make that work are unlikely to make you or anyone else happy except in some VERY specific workloads.
- watersb 5y agoMy first ever large (> 4TB) ZFS pool is still stuck with dedup. It's a backup server, gets about 2x with deduplication. At the time, it was the difference between slow and impossible: I couldn't afford another 2x of disks. These days, the pool could fit on a portable SSD that would fit in my pocket. Careful, file-based dedup on top of ZFS might be more effective. Small changes to single, large files see some advantage with block based deduplication. You see this in collections disk images for virtual machines. You might see that in database applications, depending on log structure. I don't know, I don't have that experience. For most of us, file-based deduplication might work out better, and is almost certainly easier to understand. You can come with a mental model of what you're working with, dealing with successive collections of files. Even though files are just another abstraction over blocks, it's an abstraction that leaks less without the deduplication. I haven't used a combination of encryption and deduplication. That was Really Hard for ZFS to implement, and I'm not sure how meaningful such a combination is in practice.
- justinclift 5y ago> no compression, no dedup and yet can only hit ~ 3.5 GB/s reads (TrueNAS Core, EPYC 7443P, Samsung 980PRO, 256 GiB) Hmmm, that 3.5GB/s sound low. From rough memory of doing initial storage benchmarking of our "new" Hetzner dedicated boxes a few months ago (AX51-NVMe, https://www.hetzner.com/dedicated-rootserver/ax51-nvme https://www.hetzner.com/dedicated-rootserver/ax51-nvme), they were giving about 10GB/s with mirrored NVMe drives. Just logged into one of those boxes now, and it's running 2x 1TB Samsung PM9A1 drives (https://semiconductor.samsung.com/ssd/pc-ssd/pm9a1/ https://semiconductor.samsung.com/ssd/pc-ssd/pm9a1/), compression is on (lz4), and dedup is off. (Didn't do any real tuning at the time, as these specs were already far in excess of what's needed for these servers.) Would enabling lz4 compression be useful for your use case?
- FullyFunctional 5y agoThanks for the insight. I'll can try that. Note however that reading from the raw device on the same hardware & OS I hit 5+ GiB/s (IIRC).