4 ms·
In BTRFS, RAID1 simply means 2 copies on different devices. There is RAID1C3 and RAID1C4 if you want more redundancy/copies. This whole critique boils down to
by tashbarg 5y ago
In BTRFS, RAID1 simply means 2 copies on different devices. There is RAID1C3 and RAID1C4 if you want more redundancy/copies.
This whole critique boils down to poorly chosen naming and bad documentation. Both VERY valid critique. But feature-wise, BTRFS delivers „classic“ RAID1 and more/better. It’s just hard to find and easy to misread.
- morning_gelato 5y agoI would argue btrfs does not deliver "classic" RAID1. Last I checked if you lose 1 disk in a 2 disk btrfs RAID1 setup and then reboot, it will be unable to mount the filesystem until you change the mount options so that it is set to 'degraded' mode. This is very different from other RAID1 setups (e.g. hardware raid controllers, mdadm, openzfs), and a big problem if you don't have lights out management or fast physical access to the machine.
- mst 5y agoI would argue that the 'degraded' stuff is a valid but different critique - and in fact is covered in a completely separate part of the article at some length.
- morning_gelato 5y agoI think we are in agreement. I was responding to the comment that stated "BTRFS delivers „classic“ RAID1 and more/better.", which is what I am disagreeing with. Requiring that mount options be changed whenever there is a drive failure (despite having sufficient redundancy) is definitely an anti-feature in my book.
- tashbarg 5y agoThere’s no need for changing mount options. If you want to allow mounting of degraded arrays, just put the degraded option there from the start.
- mst 5y agoThat sounds to me like it was originally set up that way early in development because they wanted people to give immediate manual attention to a system before booting it in that state. If btrfs is mature enough that it's "safe" to boot missing a disk now I think either the defaults or the documentation probably want changing to make that clearer. Like, I get "oh just add this option" as a response but in this case the fact distros don't default to adding it and the docs don't say "sure, do that" somewhere prominent mean I'm allowed to be a bit worried about how safe it actually is.
- wtallis 5y agoWhether to include the degraded option by default is a policy choice that's far beyond the purview of filesystem developers, and not something that most distros can give a clear answer to, either. It boils down to a question of the end user's use cases and risk tolerance. But it seems pretty reasonable to state that a loss of redundancy should either be handled by the user, or by a piece of software sitting between the user and the filesystem itself and acting in accordance with the user's preferences. Silently continuing to operate but with less safety than the user originally requested is the kind of dangerous that should be an opt-in feature, not a default. Moving the decision into the filesystem itself only makes sense if the filesystem is equipped to enact mitigating actions such as claiming a hot spare as the replacement device, notifying the user/sysadmin through whatever logging/reporting mechanism is actually monitored by a human, signalling applications like load balancers to stop relying on this particular machine if a healthy alternative is available, etc. (There's also an implementation detail that can trip up users who are trying to live dangerously: you're not supposed to mount a degraded btrfs array as writable until you're prepared to fix the problem making it degraded—such as by providing the devices needed to restore redundancy, or converting it to not be a redundant array anymore.)
- mst 5y agoThe critique is more about the fact that it'll stripe across any two devices in the pool. So in an 8+4+2+2 setup, you could run into serious trouble after only losing the two 2Tb drives (which are likely the oldest and flakiest two in a bodged-together-from-leftovers style array). I do agree that this is still strictly better than not being able to do it at all so long as you understand the risks, but being able to do RAID1 over 8+(4+2+2) would be even nicer for a bunch of uses and I've yet to figure out how to do that.
- wtallis 5y agoThe btrfs allocation policy is to put data on the device with the most free space, after respecting requirements for redundancy. So if your drives are sized 8+4+2+2 and you're using the RAID1 profile, it will in fact operate as 8+(4+2+2), using only the 8 and 4TB devices until the 4TB device is half full, at which point it starts using the 2TB devices. You can still end up in the situation where there's data that is duplicated across the two 2TB devices, if those are the two you started with and you added the larger devices later. But that can be fixed by doing a rebalance operation at any time after adding more devices. That's usually a good idea, though sometimes you might want to avoid a full rebalance so as not to put too much IO load on a single disproportionately large device. (If you want a hard guarantee that no data will ever be mirrored across the two 2TB drives, use dm/md to concatenate them into a 4TB block device, and add that to the btrfs array.) I have a btrfs array currently consisting of a mix of 1TB, 2TB and 4TB devices; 10 drives with a nominal total capacity of 26TB. This is using the RAID1 profile for data, and has been using the RAID1c3 profile for metadata since that feature became available. The number of drives in this array has fluctuated up and down over the years and it has survived drive failures and the accidental removal of the wrong drive from the hot-swap bay, and hasn't lost data in that time. But the current allocation is a bit uneven, because I don't always rebalance after adding or removing devices.