3 ms·
If you can't even boot because of a mystery error how are you supposed to resolve the issue? It should instead be giving the user error messages written to the
by keep_reading 3y ago
If you can't even boot because of a mystery error how are you supposed to resolve the issue?
It should instead be giving the user error messages written to their terminal, in logs, etc instead of breaking the entire system until the user finds the manual
- wtallis 3y agoIf your NAS has a hot spare drive installed then it should probably include the degraded mount option by default and automatically add the hot spare drive to the filesystem in the event of a failure. Alternatively, if there's enough free space on the surviving drives to rebalance the array and restore redundancy without replacing the failed drive, that operation could be kicked off automatically. Or the filesystem can be (hopefully temporarily) set to not store new data redundantly, if that is an acceptable risk for the user. But the filesystem cannot know which method the user would prefer; automatically rebuilding the array involves policy decisions that are outside the scope of the filesystem and requires userspace tooling. If the system doesn't have spare capacity ready, the only sane response is to not boot/mount normally. "giving the user error messages written to their terminal, in logs, etc" isn't a real solution for something like a NAS with no terminal connected and nobody looking at the logs as long as they can still establish a SMB connection; it's too likely to be a silent failure in practice. Mounting the filesystem degraded but read-only makes sense if it's necessary to boot the system so that the user (or their pre-configured userspace tooling) can decide how to deal with the problem, but a lot of Linux distros aren't happy with the root filesystem being read-only. In summary: there's no single right answer to the problem of a failed drive, and btrfs defaults to what is the safest behavior based on the information available to the filesystem itself. Userspace tooling with more information can make other, less universal choices. A distro that tries to simply adopt btrfs as a drop-in replacement for ext4 probably doesn't have all the tooling necessary to make good use of the unique features of btrfs.
- keep_reading 3y ago> If the system doesn't have spare capacity ready, the only sane response is to not boot/mount normally. It doesn't need the spare to "boot normally" and the system can turn on a scary LED, ring bells, call you, text you, hit you up on WhatsApp, DM you on Instagram, or whatever method you want your NAS to use to notify you there's a degradation. (You're monitoring it right??) This explanation of "it's dangerous to boot off a degraded array" is lunacy. I will not take this terrible advice from armchair experts when I've been doing this for over 25 years
- wtallis 3y agoI didn't say it's dangerous to boot off a degraded array. I said it's dangerous to boot off a degraded array normally. Mounting it degraded but read-only is reasonable, because that prevents silently writing new data without the level of redundancy the user previously requested. There's nothing terrible about advice against responding to a drive failure by putting the system into an even more precarious state without user interaction.
- nwmcsween 3y agoJust wondering have you worked on any large DCs or large NAS or SAN systems? Drive failures are a daily occurrence in places with a lot of spinning metal, having things fail to boot by default would be a nightmare.
- wtallis 3y ago> having things fail to boot by default would be a nightmare. Having things fail to boot would just mean you haven't configured your system appropriately for your environment. If you are using a btrfs RAID filesystem for your root filesystem, and you need that fs to be writeable in order to boot, and you want it to boot even if it's missing a drive, then you need to add an extra mount option and a few lines to your init scripts to persist new downgraded RAID settings in the event a degraded mount was necessary. But that's hardly the only valid use case for btrfs; plenty of users want strong guarantees about the redundancy of their data rather than silent downgrading. Also, do you really expect me to believe that any of the large shops still running enough spinning rust to have daily drive failures are still booting off those arrays instead of having separate SSDs as their boot drives? Separate storage of the OS from storage of the important data is such a common and long-ingrained practice that it is embodied in the physical layout of typical server systems, and the primary reason for it is the need for different tradeoffs between performance, redundancy, capacity and cost.