39 ms·
Having just heard of bcachefs when reading this article, I tried to understand what makes it better than other existing FS but couldn't quite find a clear answe
by Ecco 3y ago
Having just heard of bcachefs when reading this article, I tried to understand what makes it better than other existing FS but couldn't quite find a clear answer. It feels like it's feature set is equivalent to ZFS.
Do you guys know why someone should get excited by bcachefs?
- fodkodrasz 3y agoIt is (soon) officially in kernel, as opposed to zfs. It has a more limited feature set and said to have simpler codebase than zfs/btrfs. It has a single outstanding non-stable feature. It seems to be in active development, while btrfs seems to have become stagnant/abandonware before it was finished/stabilised completely. I have read several horror stories about data loss, so I have avoided it so far. On the other hand it is not widely deployed yet, there is less accumulated knowledge than in case of zfs. I'm looking forward to trying it in my NAS when buying new disks next year. The COW snapshots would fit my needs (automatic daily snapshots, weekly backups). (Now using LUKS+LVM+ext4, this would give a better, more integrated, deduplicated solution, I have lots of duplicated data right now)
- viraptor 3y ago> while btrfs seems to have become stagnant/abandonware before it was finished/stabilised completely Why would you think so? I can't remember the last time a kernel was released without something at least a bit exciting about btrfs https://kernelnewbies.org/LinuxChanges#Linux_6.5.File_systems https://kernelnewbies.org/LinuxChanges#Linux_6.5.File_system...
- Geezus_42 3y agoBecause they still haven't fixed the write hole issues in certain RAID configurations,which have been known for a decade or more.
- viraptor 3y agoA project not finishing some feature is not the same as being abandoned. It seems lots of people are happy to use btrfs in production without that raid mode. In other words, for all the complaining about raid5 that happens every time btrfs is mentioned, you'd think there would be at least one person who cares enough to implement it. Yet people use it in production and keep improving the other parts of that project instead.
- Geezus_42 3y agoIt's fine for single a single disk or something that presents as a single disk, like a SAN. RAID1 seems fine also. I really wanted to love it, but after it ate my data a couple of times, I gave up trying to use it for mass storage. At that time they didn't have warnings in the documentation and had stated that RAID5/6 were basically complete. I found using mirrored vdevs in ZFS much easier to manage and much more stable.
- wtallis 3y ago> I found using mirrored vdevs in ZFS much easier to manage and much more stable. That's not exactly a fair comparison. If you restricted your usage of btrfs to a similarly narrow range of features, you would probably have had a much better experience.
- creatonez 3y agoRAID1 in Btrfs is not entirely fine. It won't nuke your data and there is no write-hole issue, but if a disk fails you'll have to go into a read-only mode during the rebuild and deal with various hurdles in getting it rebuilt.
- wtallis 3y ago> but if a disk fails you'll have to go into a read-only mode during the rebuild and deal with various hurdles in getting it rebuilt. I can't say for sure that this never happens, but that's certainly not been the failure mode for any of the drive failures my btrfs RAID1 has experienced. I don't think I've ever needed to reboot or even remount my filesystem, just replace the failed drive (physically, then in software). But I always have more than two drives in the filesystem, so a single drive failure only puts a fraction of my data at risk, not everything.
- creatonez 3y ago> I can't remember the last time a kernel was released without something at least a bit exciting about btrfs They are fixing the fixable issues, but the on-disk format still makes some gotchas inevitable. It sounds like there's never going to be a great solution to live rebuilding of redundancy.
- pantalaimon 3y agoZFS will never be integrated into the Linux kernel due to it's licence. btrfs is complicated to use and has many pitfalls that can lead to it eating your data.
- ndsipa_pomu 3y ago> btrfs is complicated to use and has many pitfalls that can lead to it eating your data I use btrfs in preference over ext4 for Linux filesystems and turn on zstd compression for performance and a bit of space saving. It seems simple enough for my use case, though I'm not doing any snapshots etc. What are some of the potential pitfalls?
- Ridj48dhsnsh 3y agoJust an anecdote, but when I used gocryptfs on a btrfs partition, I'd always end up with a few corrupted files on power failure. After switching to gocryptfs on ext4, I never have any corruption.
- redneb 3y agoI have used heavily[1] the combination btrfs+gocryptfs for many years and had no problems. The only quirk is that I have to pass the -noprealloc flag to gocryptfs, otherwise the performance is really bad. [1] by heavily, I mean that I use it for my home directory
- Ridj48dhsnsh 3y agoI used it for my home directory too for several years up until a few months ago, but with default settings. The corrupted files were almost always Firefox cookies db and cache files. Maybe there's something specific about how Firefox writes them that makes them prone to corruption.
- pantalaimon 3y agoI was very excited about btrfs' advanced features, but that meant that btrfs would bite me multiple times when I expected it to 'just work™': - RAID5/6 are still not stable - it will not mount a RAID in degraded mode automatically, failing the high availability promise that might tempt you towards RAID. - swapfile support exists, but it breaks snapshots (and I don't want to snapshot the swapfile) - Just an Ubuntu/Debian thing, but snapshots are not integrated into the update process unless you install `apt-btrfs-snapshot` (and know that package exists)
- c0balt 3y agoThe main advantage, if I understood it correctly, is supposed to be performance. The promise is to have similar speeds to ext4/xfs with the feature set of btrfs/ZFS. While that sounds nice it took a lot of time to get it stable and upstreamed. Like any FS you might not want to go with the latest shiny thing but there are some that are willing to risk it, similar to debates around Btrfs vs ZFS. The last benchmarks from Phoenix are a few years old but look promising: https://www.phoronix.com/review/bcachefs-linux-2019 https://www.phoronix.com/review/bcachefs-linux-2019
- ndsipa_pomu 3y agoFrom that benchmark article > The design features of this file-system are similar to ZFS/Btrfs and include native encryption, snapshots, compression, caching, multi-device/RAID support, and more. But even with all of its features, it aims to offer XFS/EXT4-like performance, which is something that can't generally be said for Btrfs. I was surprised at that as I believed that btrfs is generally faster than ext4. Looking ahead to the last page, the geometric mean or the benchmarks supports that view too.
- BenjiWiebe 3y agoThe geometric mean on the last page shows ext4 to be faster than btrfs.
- ndsipa_pomu 3y agoOops - I was thinking that smaller was better.
- yencabulator 3y agoI worked on Ceph, a distributed storage system, for a while. Here's what we learned from benchmarks (over a decade ago): Btrfs goes very fast at first but slows down when it has to start pruning/compacting its on-disk btree structures, and then the performance suffers bad. Thus, btrfs works best when you have a spiky workload that lets it "catch up", and you never fill the disk. Concrete example: historically, removing a snapshot while under load was a disaster, with IO waits over 2 minutes. So, both sides are correct: btrfs is very fast and btrfs is very slow. XFS was never crazy fast, and has the smallest feature set of bunch, but it just kept chugging at the same pace with almost no change, regardless of what the workload did. In more complex use, you had to avoid triggering bad behavior; e.g. there was a fixed number of write streams open, something like 8, and if you had more than that many concurrent writes going your blocks got fragmented. It was very much a freight train; not particularly fast but very predictable performance, and no serious degradation ever. ext4 was sort of in between those; mostly very fast, with some hiccups in performance. Great as long as your storage is 100% reliable -- we had scrubbing in-product on top of the filesystem. We ended up recommending xfs to most customers, at the time. Predictability trumped minor gain in performance, for most uses.
- 2OEH8eoCRo0 3y ago> It feels like it's feature set is equivalent to ZFS. It does something that ZFS can't- be merged into the kernel.
- linsomniac 3y agoWorking deduplication would be amazing! ZFS has deduplication, but every time I've tried it has ended in a world of pain. Maybe they've fixed it in a more recent release, but the amount of RAM required for deduplication always outstripped the amount of RAM I had available to give it (the deduplication tables have to reside in RAM).
- ptman 3y agoZFS now has reflink support, which doesn't require lots of RAM, but isn't done automatically while writing. You need to run something like https://github.com/markfasheh/duperemove https://github.com/markfasheh/duperemove
- mastax 3y agoThe deduplication tables can be put in a special vdev now, I think.
- kevincox 3y agoIMHO bcachefs has important advantages over ZFS. It is far more flexible. ZFS is really similar to traditional block-based RAID. You can get pretty flexible configurations but 1. They are largely fixed after creation and 2. They only operate at "dataset" level granularity. bachefs has a really flexible design here where you basically add all of your disks to the storage pool and then you can pick redundancy and performance settings per folder (arbitrary subtrees, not just datasets decided at setup time) or even file. For example you can configure a default of 2 replicas for all data, but for your cache directory set it to 1 replica. If you have an important documents folder you can set that to 3 replicas, or 4.2 erasure coding. Similarly you can tell it to put your cache folder on devices labeled "ssd" but your documents folder should write to "ssd" but then be migrated to "hdd" when they are cold. And again, all of this can be set at any time on any subtree. Not just when you initially set up your disks or create the directories.
- soupdiver 3y agothat actually sounds quite neat
- FullyFunctional 3y agoThe #1 point of ZFS is protections again bitrot, ie. checksums on all data. Does bcachefs do this?
- lizknope 3y agohttps://bcachefs.org/ https://bcachefs.org/ It's literally the second item listed on the main web site Full data and metadata checksumming
- FullyFunctional 3y agoThanks, that's good news. Unfortunately it seems bcachefs isn't using pools like ZFS but instead, like btrfs, creates file systems directly on a collection of disks. That's a bummer if true.
- lproven 3y agoI tried to explain in this piece: https://www.theregister.com/2022/03/18/bcachefs/ https://www.theregister.com/2022/03/18/bcachefs/