13 ms·
Btrfs Coming to Fedora 33
- noodlesUK 6y agoIt’s funny that fedora is moving to supporting btrfs as the default FS, when red hat has stopped supporting it altogether. I’m a big fan. In a world where we can’t have GPL compatible ZFS, btrfs is the next best thing.
- ge0rg 6y agobtrfs has made some headlines in the past about its severely broken checksum computation in RAID modes, which rules them out for production use, i.e. https://phoronix.com/scan.php?page=news_item&px=Btrfs-RAID-56-Is-Bad https://phoronix.com/scan.php?page=news_item&px=Btrfs-RAID-5... The wrong parity and unrecoverable errors has been confirmed by multiple parties. The Btrfs RAID 5/6 code has been called as much as fatally flawed It would be interesting to know if this has been addressed already. The big fat red warning on btrfs' own wiki page is not inspiring trust either: https://btrfs.wiki.kernel.org/index.php/RAID56 https://btrfs.wiki.kernel.org/index.php/RAID56 For data, it should be safe as long as a scrub is run immediately after any unclean shutdown
- greatgib 6y agoI don't understand why you were downvoted. That is huge and I did not see this detailed in the original article and not even heard about it before.
- nix23 6y agoFanboys...there was even a time when peoples defended Windows for being the best server-OS, and no one needs check-summing in the FS because HW-raid is the "professional" way.
- jorritposthuma 6y agoIt's interesting to see how Synology "solves" this problem: https://www.synology.com/en-global/knowledgebase/DSM/tutorial/Storage/What_was_the_RAID_implementation_for_Btrfs_File_System_on_SynologyNAS https://www.synology.com/en-global/knowledgebase/DSM/tutoria...
- deleted 6y ago[deleted]
- berkut 6y agoExactly the same way ReadyNAS solves it...
- josteink 6y agoI observed that on my ReadyNAS too and always thought it was a weird setup. Now I know there are proper reasons for it. Good to know.
- nix23 6y agoI really don't trust Synology anymore, had so many destroyed Raid's in the past mainly with RS407. I buy a old Workstation slap FreeBSD with ZFS (RaidZ1 or 2) on it, and i am much happier.
- tobias3 6y agoSynology has solved it by running btrfs on top of mdraid, then patching it so that btrfs reports errors to mdraid and putting repair code into mdraid. This reliably works. I'm sure some would be interested in this, but Synology seems to not take GPLv2 too seriously and nobody seems to care about this. RAID 5/6 in btrfs is relatively neglected because the big users that pay developers to work on btrfs (and put the work up-stream) don't care that much about it. E.g. Facebook is probably using it for their HDFS storage, so HDFS gives them the RAID functionality.
- throwaway8941 6y agoThey use it for everything except MySQL databases, which are stored on XFS. There was a recent discussion on LKML where Chris Mason (I think) described all this.
- petre 6y agoDoes XFS have better performance than ext4 for MySQL DBs?
- ibotty 6y agoThere are some modes that are horribly broken in btrfs (e.g. most of RAID except RAID1), but the common options are safe and in continuous use.
- nix23 6y ago>but the common options are safe and in continuous use Which would be Raid5/6 and 1
- ansible 6y agoWe've recently suffered a catastrophic failure with btrfs RAID-1 on two fileservers recently. They were both RAID-1 with two 4TB drives, stock Ubuntu 16.04. With one of the two, the filesystem was allowed to fill to 99% capacity (no automatic monitoring), so obviously that is operator error, and I'm not aware of any filesystem that can handle such situations gracefully. The system became unresponsive, with btrfs-transaction taking up an increasing percentage of CPU time. Removing files and snapshots did not see an increase in free space. So what is super-curious is that the other server, which only ever got up to about 30% full, also started exhibiting the same symptoms: unresponsive, high load from btrfs-transaction. I was able to mount the filesystems in read-only mode, and recover the files, and checked it against the offsite backup. So no data loss, only a service loss. Both systems had 10 or so subvolumes, and read-only snapshots were taken three times per day for each subvolume. After 4 years, that's close to 15000 snapshots, maybe more. I searched around, but didn't find anything particularly relevant to this issue.
- eptcyka 6y agoI suffered a similar breakdown with no raid. But this was about 5 years ago. The filesystem filled up, had plenty of snapshots, but removing snapshots did not actually clean up any space. Had this happen twice. So I don't think this is related to RAID. Just regular old butter filesystem. On a slightly unrelated note, I once suffered a complete failure of BTRFS where after a shutdown it just wouldn't mount anything again. Interestingly, on IRC, I was told that his is because the firmware on my Samsung NVME SSD was buggy. It might be, but ext4 has not failed me once in that regard.
- baaym 6y agoI configured BTRFS for my data a couple of years ago on my Debian machine. It's using RAID10 - and the RAID 5/6 issue was widely known back then so I did not dare to touch that. I must say that (fingers crossed) until now I haven't had any issues with it. There have been a couple of unclean shutdowns that haven't led to any corruptions. Scrub runs every couple of weeks and on top of that the data has an offline (external disk) and off-site (rsync.net through borgbackup) backup. I'm not expecting BTRFS to let me down, but if it does I have recovery options. I tend to be very careful with my 17+ years of photo archives, especially since I'm generating hundreds of megabytes of new content every week since my daughter was born. It definitely doesn't look good that RAID 5/6 is broken, but I feel very safe using the other stable RAID modes. Sometimes I'm dreaming of setting up ZFS - however the downside of that is that it doesn't live in the kernel and you're forced to work with kernel modules. It's certainly doable, but since I had a kernel module issue with Wireguard a few months ago I'll be "looking the cat out of the tree" for a bit more before I decide if I actually want to make the move. For now BTRFS feels stable for me, so until that changes - or the benefits of ZFS increase by a fair amount - I'll probably stay on this.
- cmurf 6y agoRaid 5/6 isn't exactly relevant to this article, given its about laptop/desktop default file system. And the installer won't let you do raid5/6. Anyway a current write-up on btrfs raid 5 here: https://lore.kernel.org/linux-btrfs/20200627032414.GX10769@hungrycats.org/ https://lore.kernel.org/linux-btrfs/20200627032414.GX10769@h... Includes this observation: "btrfs raid5 is quantitatively more robust against data corruption than ext4+mdadm (which cannot self-repair corruption at all), but not as reliable as btrfs raid1 (which can self-repair all single-disk corruptions detectable by csum check)." But worth reading the whole thing.
- jiggawatts 6y agoI'm reading through the links, and at first I thought this was a historical problem that was last reported in 2016/2017, but then gems like this from mid 2020 are popping up: "When 'btrfs scrub' is used for a raid5 array, it still runs a thread for each disk, but each thread reads data blocks from all disks in order to compute parity. This is a performance disaster, as every disk is read and written competitively by each thread" I have no words. Why would anyone ever think that this is a good idea? Who sat down at their computer and typed the code that does this!? How was this never tested? You know what, thinking about it, I do actually have some rather choice words to describe the situation: This boggles the mind to a level that requires further explanation, because the casual observer would likely fail to grasp the enormity of the failure that has occurred here. This isn't like, "oops, I forgot to up-shift gears in my car when going on the onramp", this is more like "the pilot forgot about the flaps after takeoff and the plane ran out of fuel.". There's a fundamental difference in the expectation of quality between, say, a random command line utility and a RAID filesystem. To give some context: BTRFS was developed largely concurrent with, and in direct competition to Sun's ZFS. Unlike all previous SAN arrays, RAID cards, and filesystems, ZFS was explicitly designed for reliability. Sun famously had a 'test rig' where they abused each new build to death. Physically pulling disks. Randomly corrupting blocks. Running multiple operation types in parallel, while pulling disks. That kind of thing. When I read ZFS whitepapers, I was amazed at how many fundamental flaws in RAID integrity protection they discovered, and then solved. Rigorously. Meanwhile, BTRFS literally says, in 2020: Don't trust is, especially not for metadata, or data, or while scrubbing, which you had better baby-sit, otherwise say goodbye to your production environment! More fun quotes: - plan for the filesystem to be unusable during recovery. - be prepared to reboot multiple times during disk replacement. - btrfs raid5 does not provide as complete protection against on-disk data corruption as btrfs raid1 does. - scrub and dev stats report data corruption on wrong devices in raid5. - scrub sometimes counts a csum error as a read error instead on raid5. - errors during readahead operations are repaired without incrementing dev stats, discarding critical failure information. This is not just a raid5 bug, it affects all btrfs profiles. You'd have to be nuts to use BTRFS for RAID 5 or 6, and I would question its use for any form of RAID. PS: To the people downvoting this, please explain how you like people to be uninformed about catastrophic data corruption going ignored for 4 years below in the comments.
- deleted 6y ago[deleted]
- petre 6y agoSynology and SuSE are using it in production. I've been avoiding it even though we use openSuSE and it selects btrfs by default.
- cjbconnor 6y agoSynology uses a combination of btrfs and md to avoid btrfs's flaws
- ansible 6y ago> For data, it should be safe as long as a scrub is run immediately after any unclean shutdown Note that btrfs-scrub (according to the docs) is looking for on-disk block errors, comparing the data to the CRC (or whatever algo) to see if that matches. If it can recover the original data, the bad block is re-written with the correct CRC. So btrfs-scrub does not find or fix any errors with the filesystem structure. For that you need to run btrfs-check, and that can only be done offline.
- 8fingerlouie 6y ago> For that you need to run btrfs-check, and that can only be done offline. You almost never want to run btrfs-check. In itself, Btrfs is fairly "self healing", and only if you run into "weird things" in the log like invalid index files should you run check, and almost certainly never with the --repair option. The thing about Btrfs (and ZFS) is that it is Copy-on-Write, and as such, new data is written to an unused part of the filesystem. It doesn't need to modify structures of existing data, thereby reducing the risk of bad metadata. A file in Btrfs is either written or it is not. If a power outage occurs after a file has been written, but before the metadata is updated, the file is not written, and in case you're updating a file, the old file is still preserved. If a crash occurs while writing metadata, Btrfs has multiple copies of metadata, and is able to repair itself online.
- chasil 6y agoThe latest mkfs.btrfs on my Oracle Linux 7 supports "crc32c, xxhash, sha256 or blake2." The relative benefits are discussed in "man 5 btrfs" in the "CHECKSUM ALGORITHM" section. man 5 btrfs | col -b | sed -n '/^CHECKSUM/,/^FILESYSTEM/p'
- wtallis 6y agoThe reasoning for running a scrub after an unclean shutdown is not to catch corruption of the filesystem structure, but to catch data corruption resulting from the RAID5 write hole. If you're not using RAID5 or RAID6 for metadata, the write hole won't affect the filesystem structure and there shouldn't be any issues for btrfs-check to find.
- cmurf 6y ago
- znpy 6y agoI just herd my roommate saying "i don't want to play too much with my amd cpu overclocking because if i crash my system too much btrfs might get corrputed again". laughs in zfs
- vetinari 6y agoDon't worry, zfs has it's own share of problems. I have one CentOS machine where I can't update ZFS from 0.7 to 0.8, if I want to boot again (zfs#8885).
- bromonkey 6y ago*zfs on Linux
- DiabloD3 6y agoOr switch distros. I'm pretty sure I can't replicate this bug on Debian or Ubuntu.
- diffeomorphism 6y agoThat is like asking to move to a different house because you don't like the color of the walls.
- DiabloD3 6y agoTo use that analogy, no, its more like moving to a different house because lead paint remediation is too costly and error prone.
- diffeomorphism 6y agoThat seems like an entirely different analogy and basically just says this apartment is bad, you should move. Then why did you move in in the first place? The point is, you chose your distro for a reason and one little piece of software is far from enough to override that.
- bkor 6y agoI've quickly skimmed the discussion on fedora-devel regarding btrfs. I wondered mainly how they'd handle the various cases where btrfs does not work well, e.g. files that change often inline (databases, VMs, etc). Apparently an application can tell to treat those files differently. So it's basically a matter of fixing various software to work nicely with btrfs as well as any similar filesystem. As mentioned in the thread, openSUSE already uses btrfs for loads of years. I do wonder why it isn't more supported by upstream software (postgres, VM things, etc). I was also surprised how good the discussion was on fedora-devel. It seems that people didn't break into camps of "over my dead body", and "if we do not do this I'll leave", etc. It seemed more like a "this is my really thought out proposal, but is it actually feasible or did I forget anything". Then people raised problems, sometimes the proposer(s) already had a solution, sometimes it was something they didn't think of. What's nice is that despite seeing various problems everyone seemed to understand that people are trying to improve things. I didn't read fedora-devel for many years so that was a welcome change over the past.
- vetinari 6y ago> I wondered mainly how they'd handle the various cases where btrfs does not work well, e.g. files that change often inline (databases, VMs, etc). It least for default libvirt locations fedora disables CoW (chattr +C). No change in software needed.
- prussian 6y agoThis is what I was thinking as well. It just means packagers have to be mindful of what kind of files their maintained software makes and to appropriately carry these metadata changes where they are needed. I don't really see an issue with this mindset.
- nix23 6y ago>databases, VMs, etc You should disable CoW and Caching on the filesystem/-set where VM and Databases-files reside, but that counts for ZFS as well...well for all CoW-FS's in fact.
- knorker 6y agoIs btrfs still super slow at deleting large files? It can take multiple seconds (like 5-10) on my fileserver to delete just one ~10GB file.
- cmurf 6y agoThe 'rm' command should complete pretty quickly. Actually freeing up space takes time since a delete is subject to delayed allocation. The default transaction commit time is 30 seconds. And if there are snapshots or reflink copies, a backref walk is needed before extents can be freed.
- knorker 6y agoI'm talking about the time it takes for the `rm` command to finish. I've started the habit of running all `rm` commands in the background on btrfs.
- cmurf 6y agoIf it's reproducible, my suggestion is to strace the rm command and find out what it's doing that's taking so long; and post it to the mailing list: https://btrfs.wiki.kernel.org/index.php/Btrfs_mailing_list https://btrfs.wiki.kernel.org/index.php/Btrfs_mailing_list
- throwme45353464 6y agoDo you have snapshots enabled by any chance, or any other btrfs features?
- Tsiklon 6y agoInteresting - As RHEL is downstream of Fedora, I thought BTRFS was not being explored further by Red Hat, on account of their deprecation of it downstream in RHEL. If I recall this was because they didn't have the developers able to work on the software, preferring to use XFS + LVM to accomplish some of the goals of BTRFS as their STRATIS project. I wonder what this means for RHEL going forward in RHEL 9?
- rwmj 6y agoIt means nothing. This is only for the Fedora Desktop spin, and it's lead by the Fedora community not Red Hat.
- freedomben 6y agoDisclaimer: I work for Red Hat but I have zero internal insight into this. I'm just a happy Fedora desktop user I wouldn't say it means nothing because Fedora is looked at as a proving ground for inclusion in RHEL, but I would agree that one shouldn't read much into it. There are plenty of software packaged/supported on Fedora that isn't and won't be shipped in RHEL. BTRFS may or may not just be yet another one like that. I've heard/seen more excitement about Stratis (which does seem awesome so far) than I have btrfs.
- cpuguy83 6y agoThey stopped supporting it on RHEL (even experimentally) because the code churn was (is?) too high. The thing to understand about RHEL is they backport... everything.... to ancient kernels. Code that has lots of churn can be very difficult to backport, particularly to such an old codebase.
- blaser-waffle 6y agoAs a full-time RHEL admin, there is a reason: we have some old, like ooooooooold, kernels and servers floating around. I can't see needing BTFS on my granddaddy boxes, but we've definitely made use of backported code.
- 6y ago
- bleepblorp 6y agoFor desktop use, what problem does btrfs solve better than lvm+ext4? Btrfs is slower than lvm+ext4, doesn't like working with large files, requires more ongoing maintenance (scrub, rebalancing), and is more prone to data corruption. Given that lvm can do snapshots under ext4, the only real benefit of btrfs is btrfs send, but for most use cases that doesn't seem like a large enough benefit to be worth the rest of brtfs' drawbacks.
- viraptor 6y ago> requires more ongoing maintenance (scrub, rebalancing) I'm not sure that's a good description. Scrubbing is an option that you get extra. You don't need to use it and the behaviour won't be different than for example ext4 with regards to bad data detection. It's purely an extra feature. If rebalancing is useful for desktop users (wasn't really in my experience), I'm sure it will get a system-provided job that balances the resource use and amount of reclaimed space.
- chasil 6y agoThere are concerns with ZFS of the "scrub of death" on a system lacking ECC ram: https://jrs-s.net/2015/02/03/will-zfs-and-non-ecc-ram-kill-your-data/ https://jrs-s.net/2015/02/03/will-zfs-and-non-ecc-ram-kill-y... There is some debate on this question: https://arstechnica.com/civis/viewtopic.php?f=2&t=1235679&p=26303271#p26303271 https://arstechnica.com/civis/viewtopic.php?f=2&t=1235679&p=... I'm curious if the situation is improved for BtrFS.
- pseudalopex 6y agoThe first link actually explains why ZFS is no more dangerous than any other filesystem.
- prussian 6y agohonestly, I've only ever needed to rebalance on a desktop system when statfs() returned an f_bavail=0 and some program decided to take this information seriously and refused to write at all. There are still quirks with statfs info coming from btrfs volumes only solved with intense rebalancing.
- KaiserPro 6y agoI have no issues with BTRFS, apart from the documentation is really poor. I don't mind it becoming popular, but please for the love of $Deity can we please make the docs as good as ZFS. I will contribute cash if needs be. Facebook are utterly shit at documenting things, so yes it might work for their usecase, but they essentially store knowledge through Shamanism, which is terrible unless you're inducted into the world of the spirits.
- nix23 6y ago>documentation is really poor Absolutely, the best one is at Archwiki and even that one is meeh.
- eptcyka 6y agoI for one support Facebook's push to make more problems solvable solely by entering a mud hut and ingesting hallucinogenics.
- blaser-waffle 6y agoThat won't make them less evil, just less coherent.
- KaiserPro 6y agoIts only less coherent if youre not off your tits.
- rwmj 6y agoA few points: - No this doesn't affect RHEL. - It's only for Fedora Desktop spin (which for various reasons including this, but also others, you shouldn't use even on a Desktop - I install Fedora Server on my laptop). - Only a subset of btrfs features will be used, especially avoiding the ones which are known to be problematic.
- rkangel 6y agoCan you elaborate on the reasons for using Fedora server instead? I use Fedora Workstation on several laptops and desktops (without issue). I'm curious if I'm missing some problem/opportunity.
- rwmj 6y agoDefaults to GNOME, firewall disabled, ext4 (and now btrfs) instead of XFS. (These spins are all about defaults - so installing Server doesn't make it any less useful for desktops, but you may have to dnf install a few things the first time you use it.)
- rkangel 6y agoWhat do you mean "defaults to GNOME"? Server defaults to no DE at all I thought (been a year or two since I've used it) and Fedora defaults to GNOME anyway. If it by default doesn't install email client and music players I don't use that might be nice though.
- boredishBoi 6y agoI think the defaults they’re referring to are for fedora desktop, not server
- rkangel 6y agoYes you're right - thank you. On a re-read I see that I had the assertions the wrong way round.
- AdmiralAsshat 6y agoThere's probably no way to convert the legacy ext4+LUKS filesystem over, so, I guess I'll decide when F33 comes out whether I want BTRFS enough to do a clean install rather than an upgrade.
- filmor 6y agoYou can convert an ext4 filesystem to btrfs using btrfs-convert.
- symlinkk 6y agoIs that stable?
- LanternLight83 6y agoAlright, so, disclaimer, I have lost data doing this but it was purely operator error! Closed my SSH session at the worst possible time. One of the podcast hosts at Linux Unplugged made the same mistake. The btrfs-convert tool hypothetically leaves the ext filesystem all but untouched, and COW's the needed filesystem metadata onto the end of it, with data modifications coming thereafter (or intelligently stored within the free space of the ext system). You wait until you feel comfortable with BTRFS, then delete the preserved ext system and run a balance, which rewrites all data to disk in the usual structure. Alternatively, the preserved system can be restored, although I don't actually see instructions for that. https://btrfs.wiki.kernel.org/index.php/Conversion_from_Ext3 https://btrfs.wiki.kernel.org/index.php/Conversion_from_Ext3 I love btrfs and would trust it but am glad to have backups of irreplaceable data.
- AdmiralAsshat 6y agoI suppose I wouldn't need to worry about an SSH disconnect while working on localhost, but I'd still be a little concerned about trying it. I recently had a Pop!_OS/Ubuntu system irreparably damaged during an upgrade because the lock screen come on during the five minutes I had stepped away from monitoring to use the bathroom. The wiki also doesn't mention how well this process plays with LUKS, so, given how sensitive headers and such can be, I'll probably wait to hear from some F33 early-adopters on the conversion tool to see how it turns out. I'm reasonably happy with my ext4 system as it is, and most of my core data is on backup disks anyway, so I'm not sure how much the integrity checking of BTRFS would help. Seems like it would be more critical to have on the backup disks to make sure I'm not propagating the bit-rot.
- flas9sd 6y ago> The switch to Btrfs will use a single-partition disk layout, and Btrfs’ built-in volume management. The previous default layout placed constraints on disk usage that can be a difficult adjustment for novice users. Btrfs solves this problem by avoiding it. good call, subvolumes came to the rescue of the /home partitioning. It's useful, but the inflexibility was in the proper split of available disk space. Novice users then saw warnings for /home - when there was still plenty of space on the root partition.
- kev009 6y agoSurprising, I thought RH had drawn pretty clear lines in the sand in favor of evolving XFS to be the universal disk FS for Linux.
- kevin_b_er 6y agoSo how long until systemd-somethingmajor requires btrfs?
- 48bb-9a7e-dc4f2 6y agoIt's painful to see how much disinformation (even in subtle ways or to go off the rails) some topics are getting in here. If you want to get a better picture, also about Zfs and Fedora, read the previous Btrfs threads where the developers took the time to discuss it. And to kill some FUD. I have no association with Btrfs or Fedora but I'd like to have a modern FS in-tree as battle tested as it can be.