3 ms·
You can probably achieve what you're looking for by stacking a few filesystems. For example, you could create a separate ZFS pool/vdev with a single full-disk z
by 555h 9y ago
You can probably achieve what you're looking for by stacking a few filesystems. For example, you could create a separate ZFS pool/vdev with a single full-disk zvol on each disk. Then use mdadm to create a RAID array of the zvols. Then put ext4 (or whatever) on the mdadm array.
I've done something similar for the purpose of getting FDE with ZFS in linux. It can be a little finicky, but it's definitely workable.
One ZFS-specific caveat (which may conflict with your desire to get high storage efficiency): you way need to prevent your ZFS pools from filling up too much [1]. You can either enable discard/TRIM on the whole stack, so the top level FS (e.g. ext4) can let ZFS know when a block is actually free. Or alternative just to limit your zvols to 85% (for example) of their respective pools. The latter is my preference, because there was originally a bug with discard in zfs and it's not immediately clear if it's totally fixed (although my fstrim tests seemed to work out fine).
[1] https://www.reddit.com/r/zfs/comments/3vtur4/what_exactly_happens_when_you_go_over_80_on/cxqu7uh/ https://www.reddit.com/r/zfs/comments/3vtur4/what_exactly_ha...
- Veratyr 9y agoHmm, that actually sounds workable. I could even format the mdadm device as ZFS too if I really wanted. I am somewhat worried about performance, have you had any issues with that?
- 555h 9y agoI haven't had any issues with performance, but then again my requirement was just "reasonable performance". I ran some quick benchmarks (data below). Obviously this is far from rigorous, but maybe it'll be useful. In previous tests I found that volblocksize=128K was optimal for my stack -- which is why the last benchmarks use that setting. Every additional ZFS filesystem in the stack may reduce storage efficiency (minimum free space requirements [1]; metadata & checksum overhead [2][3]) -- that's why I used ext4 as the top layer instead of another ZFS. [1] (as mentioned before) https://www.reddit.com/r/zfs/comments/3vtur4/what_exactly_happens_when_you_go_over_80_on/cxqu7uh/ https://www.reddit.com/r/zfs/comments/3vtur4/what_exactly_ha... [2] https://news.ycombinator.com/item?id=14756360 https://news.ycombinator.com/item?id=14756360 [3] https://forums.freenas.org/index.php?threads/what-is-the-exact-checksum-size-overhead.28187/#post-183802 https://forums.freenas.org/index.php?threads/what-is-the-exa... Test setup: debian stable kernel 4.9.0-3-amd64 zfs 0.6.5.9-5 ZFS "pool": mirror with 2x 7200rpm drives Benchmark command: for i in `seq 1 10`; do sync; dd if=/dev/zero of=DEST bs=1M count=1024 conv=fdatasync; done zfs mirror -> dataset Data (MB/s): 125,115,104,135,148,170,135,151,118,119 Mean (MB/s): 132.0 Std.dev.: 19.9 zfs mirror -> zvol (volblocksize=8K [default]) Data (MB/s): 150,115,127,125,122,118,105,118,124,128 Mean (MB/s): 123.2 Std.dev.: 11.6 zfs mirror -> zvol (volblocksize=128K) Data (MB/s): 68.5,112,115,114,94.3,85.1,83.1,98.4,120,108 Mean (MB/s): 99.8 Std.dev.: 16.9 zfs mirror -> zvol (volblocksize=128K) -> luks -> ext4 (my stack) Data (MB/s): 130,94.4,109,139,138,125,94.9,124,134,133 Mean (MB/s): 122.1 Std.dev.: 16.8 edit: formatting
- libx 9y agoCan you please elaborate what are the advantages of such configuration?
- 555h 9y agoThe advantage is that you can get all the features you want, even though they aren't all available in one filesystem. In my case, I wanted a reliable filesystem, RAID1/mirror support, block-level checksumming, and full-disk encryption. No filesystem provides these on linux right now. My solution was therefore to use ZFS to provide a mirrored, checksummed, reliable volume -- onto which I put a standard LUKS-encrypted ext4 filesystem. In the past I had tried the opposite (LUKS on the bare drives, then a ZFS mirror of the 2 decrypted volumes), but it was kinda annoying to manage, and I don't really need the other ZFS features (like snapshotting). In the grandparent post, the requirement was the ability to expand the RAID volume without rebuilding (something that ZFS doesn't offer), plus checksumming and reliability (which ZFS does offer). So one option would be to use mdadm to manage the RAID array, and then put ZFS on the resultant volume in order to get checksumming. The disadvantages are: extra complexity; extra overhead; more potential points of failure; more management hassle; etc. As soon as encryption for ZFSonlinux is stable, I'll be very happy to drop this filesystem stacking in favor of that!