4 ms·
I contributed a few patches to ZFS on Linux about 8 years ago - at a time when it was still very much in its infancy and panic'd when you looked at it in the wr
by nilsb 7y ago
I contributed a few patches to ZFS on Linux about 8 years ago - at a time when it was still very much in its infancy and panic'd when you looked at it in the wrong way.
It's incredible how far they've come. We're using ZFS on Linux on about 120 servers at work and it's rock solid. Snapshots are a life saver in our day-to-day ops.
- pletnes 7y agoWould you know how it compares to btrfs? I found btrfs easy to set up but with hard-to-debug rough edges.
- nilsb 7y agoUnfortunately I don't have much experience with btrfs, however I've found the CLI commands for ZFS to be much more intuitive to use - both in terms of how they're invoked and their output. There's one command for handling storage pools (zpool, https://manpages.debian.org/unstable/zfsutils-linux/zpool.8.en.html https://manpages.debian.org/unstable/zfsutils-linux/zpool.8....), e.g. adding/replacing disks, monitoring I/O utilization, etc. And then there's another command for dealing with ZFS datasets (zfs, https://manpages.debian.org/unstable/zfsutils-linux/zfs.8.en.html https://manpages.debian.org/unstable/zfsutils-linux/zfs.8.en...), e.g. setting properties for datasets (quotas, compression, delegation of privileges to non-root users), managing snapshots. Both CLI commands are scriptable (e.g. -H can be used to suppress human-readable headers and -p turns off "friendly" formatting for numbers) and there are libraries (libzfs, libzpool) which can be used to access their functionality (e.g. for managing snapshots) in your own programs. I don't think I can do ZFS on Linux (or btrfs for that matter) justice in a (short) comment on HN. There's just so much to it (e.g. SSDs as cache/log devices, sending/receiving snapshots over the network, compression, etc.) There are some rough edges, of course. Until a few years ago ZFS on Linux had massive problems with resource management, i.e. it would regularly panic in low-memory situations. These appear to be fixed though - at least I haven't seen any panics in the past two or so years. Due to its license ZFS on Linux will probably never be part of the upstream kernel. We use it on Debian which means we get to compile the ZFS kernel modules on each box when there's a kernel update (with DKMS). Plus there are some pitfalls when using ZFS as your root filesystem (on Debian stretch without systemd services would start before /var/log was mounted). We've had plenty of disk failures which ZFS detected well before SMART did because we run monthly scrubs for our ZFS pools (i.e. basically ZFS verifies its checksums for all the data that's stored in a pool). Recovery is as easy as popping in a new disk and running "zpool replace". Ok, this ended up much longer than I had planned. This has to suffice for now.
- braindeath 7y agozfs is really a study in cli ergonomics. It has a quite idiosyncratic CLI, which one might assume is a bad thing. But the CLI is so well thought out that learning this new "language" is quite intuitive. btrfs uses a more traditional command, subcommand, args setup, which by itself is not a problem, but the actual command and subcommand structure is a dumpster fire to be charitable. Try managing large groups of snapshots without custom tooling. The whole thing makes me sad really. btrfs has a lot going for it at the architecture and disk format level (dedup is actually something theoretically useful, while the zfs design was flawed from the start), but the implementation has just never made it.
- rkagerer 7y agoDumb question as I haven't used the feature: Why is dedupe flawed? Is it because it requires "enormous" amounts of RAM? Does it eventually slow down writes?
- braindeath 7y agoIt's not just the issue of RAM, but that dedupe is only attainable with the online DDT method. There is no offline deduping at the block level, and of course related annoying things like no COW between filesystems.
- rincebrain 7y agoDedup on ZFS is problematic because ZFS, in exchange for some of its core useful features, promises that block locations on disk are immutable. So the only place you can do dedup is inline, as the data is being written the first time, not after-the-fact. In addition, this requires you keep a huge indirection table (the DDT, or dedup table) that needs to be read for all writes, so either that has to be kept in memory or on fast storage, or you've just turned every write into one or more random reads, plus writes. This also means that even if you turn off dedup after turning it on, the performance implications remain until the DDT no longer contains any blocks (e.g. you rewrote all the data after turning dedup off). There are feature proposals to make the performance of dedup less pathological, but nobody's taken up implementing them so far. (Someone even did a proof of concept implementation of one of them, and it still hasn't been finished and integrated.)
- Youden 7y agoBTRFS is more flexible than ZFS (for example it works fine with RAIDs of differently sized drives and allows you to add and remove disks after the FS is built) but it really worries me that any filesystem has a big red warning on documentation for a major feature [0]. [0]: https://btrfs.wiki.kernel.org/index.php/RAID56 https://btrfs.wiki.kernel.org/index.php/RAID56
- danudey 7y agoMy experiences have been as such: With the exception of accidentally enabling deduplication on a 16 TB array on a system with 8 GB of RAM, ZFS has been fantastic. Everything works, and works much as you'd expect. The tools and concepts are clear and concise and it's obvious that a lot of work was put into everything from the design to the documentation. btrfs, on the other hand, has been a nightmare. I set it up on a new desktop machine earlier this week, and I've spent most of this afternoon trying to recover data from it. Due to a power outage, the filesystem is corrupted. I didn't know that right away because I enabled zstd compression in the filesystem but GRUB couldn't boot/read zstd files/filesystems, etc. A lot of btrfs is counterintuitive, or downright worrying. Most of the threads I've seen regarding filesystem errors end with "I reinstalled and everything is fine". The documentation suggests not trying to `btrfs check --repair`, because apparently that's the wrong way to do things? You should try to mount with the 'recovery' mount flag instead, which is not intuitive. In my case, the kernel is throwing an error about the checksum map, but it can't rebuild it or repair the filesystem. In a rescue image it errors because it tries to call pthread_cancel but can't load libgcc for whatever reason, so I can't rescue my system from a rescue system. Even when it did work it was confusing. Unlike ZFS, btrfs subvolumes seem... counterintuitive? On ZFS I can take a recursive snapshot of a volume or subvolume and any subvolumes it has, which allows me to divide a subtree up into multiple subvolumes but still treat it the same (e.g. if I want multiple entries in /var/lib/mysql/<mysql_instance> or something similar). On btrfs, you cannot do recursive snapshots, and I've even seen some people describe this as one of the (few) reasons you'd use subvolumes: to exclude something from snapshots. After ten years, it feels as though btrfs is 90% done; that is to say, it has 90% of the functionality it should, that those features are 90% done, and that what is done works about 90% of the time. Honestly, with the state that btrfs is in, I don't understand why it's in the kernel at all, and why distros support it (but I do understand why RHEL pulled it). Working with it directly makes me worry for the data I have stored on my synology, and now I'm wondering if I should have just built a FreeNAS box instead. TL;DR btrfs is a giant mess and you shouldn't use it. ZFS has more up-front overhead (getting the package versions configured for Ubuntu) and is more difficult to go all-in on (e.g. ZFS root) but at least you won't lose your data out of nowhere.
- rkeene2 7y agoI compared them in 2010 here: https://rkeene.org/projects/info/wiki/BtrFS https://rkeene.org/projects/info/wiki/BtrFS Overall similar. I ran into more critical bugs with BtrFS, but ZFS (on Solaris) was not perfect (most of our Solaris kernel panics were ZFS-related on Solaris).
- salamander014 7y agoI've used both in a very non-critical way (think basic dev workloads on linux, no DBs or any high performance applications). I found the same as the comment above, btrfs was a dream to set up and a nightmare when something went wrong, which was several times a year it seemed. ZFS on the other hand, I had to write scripts to handle my snapshots and snapshot expirations but the filesystem itself has been rock solid. I haven't put ZFS through it's paces. However with btrfs I had so many issues on a dead simple workload it just didn't seem stable to me. This sucks and I wish I could contribute because I loved their goals and thought they were designing everything right, but stability wins the long game, as it were.