4 ms·
Downvoted because I frequently hear this sentiment from both sides of the argument but rarely is anything more than an anecdote offered to support it. Feels mor
by click170 10y ago
Downvoted because I frequently hear this sentiment from both sides of the argument but rarely is anything more than an anecdote offered to support it. Feels more and more like a mud slinging competition instead of an assessment of merit.
- finnh 10y agoYou may wish to read the article we are discussing, then.
- rleigh 10y agoMany people use both filesystems, and you'll see many claims of how both work really well and that the user hasn't had any problems. The problem with these people's experiences is not that they are untrue, but that they are primarily from people who haven't had hardware failure/glitches, and who have never had their systems run the failure-case codepaths. Everything's fine and dandy up until the point you lose all your data. I've run Btrfs on many systems since just after it started to be usable, and written software with Btrfs-specific support which hammers it (and LVM) like nothing else creating and destroying tens of thousands of transient snapshots. I've also now run ZFS on several systems, admittedly over a smaller timeframe (3 years vs 7-8ish). I've had Btrfs totally trash a RAID1 mirror from a transient SATA cable connector glitch. On this test system, I had half the disk using Btrfs, half using mdraid/LVM. The mdraid half recovered and resynced transparently as soon as I reseated the connector; no service interruption or dataloss. Btrfs ceased to function, and on reboot toasted both mirrors resulting in total unrecoverable dataloss and repeated kernel panics. That's been fixed a while, but right here we're seeing the same thing. The failure codepaths, which are of critical importance, are untested and buggy. And even non-failure codepaths are still bad. Take the snapshotting case above, I had to take the system offline and do a full manual rebalance every 18 hours. The time from fresh new filesystem to read-only unbalanced disaster was just 18 hours when thrashed continuously, at most using 10% of the total space. And lastly, the performance of some things such as fsync are truly abysmal, to the extent that we had to use "eatmydata" to completely disable it for apt/dpkg operations! When under heavy parallel workloads, it could take many tens of minutes or hours(!) to complete writes which ext4 would complete in a minute or so. I've yet to experience any problems at all with ZFS. Now that might have been luck on my part, but it might also be down to better design and quality of implementation. It's certainly been battle tested in high end installations. That's not to say that Btrfs doesn't have some neat features; it does a few things ZFS doesn't, like rebalancing data over its devices while ZFS only does that on write. But Btrfs has let me down badly every time I've used it in anger, and those few neat features don't make up for its lack of robustness--the primary purpose of the filesystem is to reliably store data, and it fails at that. I don't like to see "mud slinging", since such fanboyism is unobjective and uninformed. I've reached my opinion based upon several years of practical intensive use of Btrfs for various things, the most demanding of which was repeated whole-archive rebuilds of the whole of Debian when I was maintaining the Debian build tools, and wrote btrfs snapshot support specifically for them, doing over 30 parallel builds on a single system using independent snapshots per build with over 20000 snapshots per run, creating and destroying several per second. The experiment was disastrous, and showed Btrfs to be unsuitable for such intensive workloads. When your filesystem is guaranteed be turned read-only at some unpredictable and unknown point in the future, you can't rely on it. Regular rebalancing mitigates but doesn't solve this, and has a terrible performance impact. Not dataloss per se (unless it makes you lose writes when it turns read-only), but it's a serious design or implementation flaw. I did all this testing and adding of Btrfs support to various tools because I had high hopes for its potential; unfortunately they exposed serious shortcomings, many of which exist to this day. Today I'm using ZFS, not because of any irrational prejudice against Btrfs, but because Btrfs has never managed to deliver a robust and well tested filesystem!
- nisa 10y agoI'm not having nearly your experience but I just want to say I'm agreeing 100% to you conclusions based on my experience. We ran at Uni a Hadoop Cluster that had disks slowly dying (some bad sectors every few days, but otherwise fine) and lacked the money to replace them. We replaced ext4 with ZFS (no raid, just plain zpools with failmode=continue) and ZFS ran mostly fine and scrubs kept the metadata sane. Never had data loss (HDFS has it's own replication and checksumming, we just need sane metadata for Hadoop to run and intermediate MapReduce outputs where send to directories with zfs set copies=2) and we only replaced the botched disks that had longer scrub times or couldn't survive a scrub. I'm still surprised how ZFS managed to pull that off. Only bugs I've found where related to ZoL at this time but could be worked around. btrfs switched to readonly as fast as ext4 (which is probably the correct thing to do) but was useless for this problem. On new hardware with new enterprise disks we choose btrfs and we had a painful tour of crashes, data corruption, metadata corruption (undeletable files), deadlocks until kernel 4.4 where things got a little bit better. Here disks and server where enterprise class and fully working. Just btrfs bugs. Also no RAID. I'm not doing that anymore so I don't know if any new bugs appeared but the whole experience will keep me from ever using btrfs. This was ~2years ago and you could easily find lot's of slides Fujitsu or Suse that btrfs is stable and you can use it (around kernel 3.13-3.16). It's probably fine for your notebook or even your backup HDD but don't think you can stress it without experiencing pain (be it corruption, hangups or dataloss) or just abysmal performance. That beeing said ZFS on Linux is also a far cry from rock solid but I'm optimistic that they flesh out the problems and tackle them in a solid way. As a Linux fanboy for years this gave me some solid appreciation for Solaris engineering.