10 ms·
Grub2 security update renders system unbootable
- jacquesm 6y agoIf you are wondering why many people hate updating working systems, no matter what the security implications, look no further than this. Time and again it is an innocent security update that would end up in a reinstall, finding a bunch of bloat ware on a system, losing critical functionality, data loss and time lost. Updates should be restricted to the absolute minimum and tested to the point that deploying them does not put customer data at risk.
- dkdk8283 6y agoThis is especially concerning since the entire value proposition of RedHat is stability and meticulously maintained packages
- rubber_duck 6y agoI thought their value proposition was enterprise support for large companies that want to outsource that ?
- Wowfunhappy 6y agoWell, and they still do a better job than... a lot of companies.
- fluential 6y agoRarely such updates are deployed into production without testing in non-prod envs
- MattGaiser 6y agoBut that is part of the resistance to updating. A whole pile of testing must be done again.
- tankenmate 6y agoBut the testing should be automated so that incremental changes are just that incremental. Stop the world massive changes are an anathema as it so much harder to pinpoint the breaking change.
- tanzann 6y agoWhat can you recommend as a balanced linux distro between the constant updates of arch on one hand and mostly very outdated package versions of debian/centos on the other?
- michaelcampbell 6y agoSo you want stable + cutting edge?
- mattdm 6y agoYes. That's the "simultaneously too fast and too slow" that all distros have to grapple with.
- _ph_ 6y agoThere is of course Fedora. The latest is quite cutting edge, but you can choose when to upgrade to the next release. And there are security updates for older releases, so you don't have to immediately upgrade. Still, keeps you in the Red Hat universe.
- freedomben 6y agoThis is exactly why I settled on Fedora after a few years on Arch, a few on Ubuntu/Debian/etc. Fedora is the best blend I've found. New releases about twice a year but you can skip a release if you want and upgrade every other release (or always stay on release n-1 which some people I know do). Arch and Fedora follow the same conventions so everything on Fedora is where you expect it from Arch (this also makes much of the great Arch wiki applicable directly). There will always be a soft spot in my heart for Arch, but Fedora is the perfect balance for me. I can't have my work machine getting broken and forcing me to read forums and mailing lists, etc.
- forgotpwd16 6y agoInstead of being afraid of updates you should be able to rollback any problematic updates and be impossible to end up in a broken system because updates interrupted. The only distros offering that are NixOS and GuixSD.
- hinkley 6y agoThis is also why I pay someone to change my oil. It's a simple procedure, but the list of things that can go wrong is a list. At a small shop I ran our VCS on my workstation (because otherwise we had none) and at the end of a day I rebooted to do some maintenance on the machine, and it didn't come back online until around 1pm the next day. On the plus side we finally got a dedicated server. That said, you still gotta touch systems. If you get off the upgrade treadmill, you will find yourself one day facing a CVE issue that have no patches for the versions you use. And then not only do you have to do a security update, you have to fast-track a version upgrade that your code might not work with anymore. Last time we had this, I got the team to agree to putting one upgrade of OS or libraries per month onto the task list. We even let the assignee choose what they were going to upgrade (a few had no preferences, and so we'd have them work on whatever was most overdue). We put some caches in place at work recently, to improve perf and give us some protection against partial outages. Already we are seeing services that cannot handle the same workload they could three months ago. To paraphrase Jim Highsmith: there are no answers here, and we waste enormous amounts of time looking for them. There are only tradeoffs, with their own sets of consequences to be managed. The tradeoff I prefer is to drill it into people that shit happens. Something is going to be down, and the more things we have, the higher the likelihood that one of them will be down at any given moment. You can put all of your eggs in one basket and have one outage a year, in which everyone not involved in fixing the damned thing might as well go home, or you have a bunch of systems where I can do 80% of my duties if any one of them is offline. Systems that are never seen to break (or at least stumble) tend to become siloed, insular. Being able to talk about the 'bigger issue' around an outage is often a time when teams finally communicate with each other (unless you have dysfunctional managers, in which case: run) And if all else fails, there is always some set of tasks you want people to do but they never seem to have the time for. Having a captive audience isn't the worst thing that could happen to you.
- kevin_thibedeau 6y agoI won't let anyone change my oil because a monkey once severely crossthreaded the drain plug by feeding it on with an impact wrench rather than hand starting it.
- cesarb 6y agoNotably, this is the second time in a few months that this happens. The previous one was a microcode update, though that one was Intel's fault (according to Intel, it happened only when the microcode was loaded by the kernel; they probably had tested only loading it from the firmware).
- theevilsharpie 6y ago> according to Intel, it happened only when the microcode was loaded by the kernel; they probably had tested only loading it from the firmware I followed the Spectre situation quite closely, and never heard anything along those lines. Do you have a source? Many system vendors pulled the faulty firmware from their own sites, and Dell went so far as to recommend rolling back if you had installed it, so the partner ecosystem didn't want people running the firmware at all.
- cesarb 6y ago> > according to Intel, it happened only when the microcode was loaded by the kernel; they probably had tested only loading it from the firmware > I followed the Spectre situation quite closely, and never heard anything along those lines. Do you have a source? Sure, it was this one: https://github.com/intel/Intel-Linux-Processor-Microcode-Data-Files/issues/31#issuecomment-644885826 https://github.com/intel/Intel-Linux-Processor-Microcode-Dat... "Intel identified an issue when OS loading microcode update revision 0xDC for cpuid 406E3 and 506E3. The microcode update has been reverted to revision 0xD6. This issue does not affect the microcode update when loaded from BIOS."
- theevilsharpie 6y agoAh, I see. This is for a different firmware issue. They have so many these days, it's hard to keep track. :/ Thanks!
- jchw 6y agoIn my opinion, the true antidote is reproducible, and ideally immutable, systems. Maybe this bug would’ve been harder since it’s all the way back to the bootloader, but assuming a more minimal bootloader (perhaps EFIstub directly, or systemd-boot,) you could have something like NixOS where you can, at boot time, jump to a previous configuration with all of the old software. (Works even better for docker or VM images where you have a supervisor that can move around disk images. Can just construct immutable root images and roll back by booting with an older one.)
- necovek 6y agoThis one is actually quite interesting: on a long running server with kernel livepatching, you might not even notice it if it gets fixed before you reboot the next time.
- tzs 6y agoI'd also add that you should reboot after updates--even if they haven't touched anything that requires a reboot to put in use. The reason is that there is plenty of configuration information that is only read during startup, and plenty of programs either only used during startup or that do things during startup that they do not do later. If an upgrade broke any of that you might not find out until the next reboot. It is a whole lot easier to deal with if you find out right away that the update broke booting than if your first reboot is six months and several intervening updates later. We hit something like that once. An extended power failure took down several servers. When they came back up several of our things were messed up. We eventually tracked it down to some update months earlier had made a change to how network filesystems got mounted, making them show up later in the boot process--which turned out to be late enough to be after some of our stuff that expected the network drives to be available.
- kd913 6y agoI will state the same comment as last time. Can distros maybe consider moving to systemd-boot at some point? Systemd is already built in and can handle things like mounting pretty easily and simply. It is a hell of a lot leaner than grub, doesn't use a billion superfluous modules. That and it is a lot easier to prevent tampering compared with the cumbersome nonsense that is grub passwords. Oh and it enables distros to gather accurate boot times and enables booting into UEFI direct from the desktop. It works with secureboot/shim/Hashtool. Also each distro has it's bootloader entries in separate folders to avoid accidental conflicts.
- storedbox 6y agoUnfortunately, I think this move would receive quite a lot of flack due to it being systemd. I totally agree though.
- dijit 6y agosystemd-boot (AKA, gummiboot) is leaner, for sure- much in the same way a bicycle is leaner than a truck. It might surprise you to know this but systemd-boot does not support, among other things: "BIOS"/MBR boot, EXT4, XFS, mdraid, LUKS (hidden boot), OPAL (aka hardware FDE, self-encrypting drives) or even btrfs! mdraid being pretty critical to a lot of server work. Then there's the topic of altering bootloaders for distro's which value stability, it wouldn't happen in a point release.
- michaelcampbell 6y agoI don't use it, and I don't respond in the least bit of snark (honestly!), but it DID surprise me to learn all the things it doesn't support.
- Spivak 6y agoI feel like you're making this sound worse than it actually is * systemd-boot doesn't support those filesystems as the /boot partition, all of those things for your / are totally fine. For most people having your /boot partition be XFS vs FAT32 is firmly in the "who cares?" realm. * Nobody really supports OPAL, even grub. The only reliable solution is using https://github.com/sedutil/sedutil https://github.com/sedutil/sedutil to unlock the disk and then soft reboot into any normal bootloader. * You can do mdraid, you just have to stick the metadata at the end of the drive rather than the beginning. The utilities that set up your raid even warn you about this because this setup is so standard.
- tmikaeld 6y agoAgain? I remember when zfs didn't init in time so grub found no root. It was fun panic when it's on a virtualization server..
- protomyth 6y agoAm I reading this correctly that "yum update" is all I have to do to screw up an 8.2 minimal install?
- fctorial 6y agosudo yum update
- wtracy 6y agoAny word on whether the equivalent update on Debian has similar issues?
- psanford 6y agoYes, there are reports of failures on Debian: https://bugs.launchpad.net/ubuntu/+source/grub2/+bug/1889509/comments/6 https://bugs.launchpad.net/ubuntu/+source/grub2/+bug/1889509...
- psanford 6y agoThe link is to a redhat bug report, but the issue is also affecting other distros: https://bugs.launchpad.net/ubuntu/+source/grub2/+bug/1889509 https://bugs.launchpad.net/ubuntu/+source/grub2/+bug/1889509
- deleted 6y ago[deleted]
- RedShift1 6y agoGrub2 is not a bootloader. I'm not even really sure what it is. In Grub 1 you had a configuration file with the operating systems you wanted to boot. Simple, effective. With grub2 all I'm seeing is a bunch of sh scripts that are impossible to write by hand.
- peteri 6y agoI assume this is related to this from yesterday https://news.ycombinator.com/item?id=23990075 https://news.ycombinator.com/item?id=23990075 Which is about revoking secure boot keys
- Wowfunhappy 6y agoThere was a story on HN last night from Debian where they laid out this issue, and basically stated "Yes, this security update is going to render some systems unbootable, here is why we're doing it anyway." https://www.debian.org/security/2020-GRUB-UEFI-SecureBoot/ https://www.debian.org/security/2020-GRUB-UEFI-SecureBoot/ Stability is important, especially when it comes to unbootable machines—but I don't quite know what anyone was supposed to do here. If a user has secure boot enabled, the OS has to assume that the user wants/needs security at that level of the chain—and it is therefor responsible for ensuring the chain's integrity. In this case, there was no way to do that without some machines (temporarily) failing to boot. What would have been a better way to handle this?
- RedShift1 6y agoIs there proof that secure boot actually prevented an attack?
- wbkang 6y agoA huge number of devices today (mobile or not) have their disk encryption keys protected by secure boot or similar mechanism where the TPM vends the keys only when securely booted, and they are not breakable trivially. So I'd say the answer is yes.
- Wowfunhappy 6y agoI don't know, but if you don't think secure boot isn't valuable and/or worth the trade-off, you should turn it off on your machine. Anyone who does so will not be affected by this update. I know there are some environments that don't allow turning off secure boot, but this really isn't the OS vendor's fault.
- g_p 6y agoSecure boot has, for the most part, prevented the underlying system from being compromised by malware that can run before boot (and therefore hide itself from the OS). I don't think we can really prove a negative here, but it means no more "boot sector" viruses, either on the HDD itself, or on external bootable media. That certainly is a good first step to preventing foot canon incidents, and giving the OS a decent chance at preventing malware running before the OS has control of the system. Couple in, as another post said, a tpm which only releases the decryption key if the system is in a valid state with secure boot running, it helps a lot with basic cold boot protection. TPM is far from perfect, but it is arguably better than a user not encrypting their disk at all, and this prevents attacks like replacing the utilman.exe with cmd.exe etc.
- gojomo 6y agoWell, a system that can't be booted is totally secure against many kinds of remote & local attacks. So, "mission accomplished", I guess?
- znpy 6y agoThe stablest systems I had the pleasure to manage were two identical rhel6 clusters into different geographical locations for higher availability and fault tolerance. Such systems were installed, turned on and never touched again. Kernel 2.6.32 that managed about six-seven years of uptime up to mid-2019. We operated a lot onto such systems, mounting and unmounting iscsi devices, starting and stopping stuff, turning on and off network interfaces and clustered filesystem (thanks to Veritas cluster manager). The key move was never updating. Such systems were literally mission critical, without that cluster the whole company was unable to produce its main products. Considering how much stuff they ran and how many simultaneous users were connected, I was humbled by their stability (and by rhel's stability). If you're getting angry at this post: the customer was not in the it field and was completely okay with buying new hardware and doing a full reinstall every X years.
- MaxBarraclough 6y agoUsually security is one of the top reasons to update software systems. Would a security flaw in these clusters have been an issue? Were they aggressively firewalled?
- znpy 6y agoYes, there were multiple firewall layers. Also, services were not accessed from outside the corporate network.
- awill 6y agoThis is why I dislike grub. It's really, really bloated. A bootloader just needs to pick the partition to boot, and little else. I switched to gummiboot ages ago, and it's so simple. There's far less to go wrong (gummiboot got absorbed by systemd, so it's now called systemd-boot)
- jkingsbery 6y ago"This update enhances security by making the system unbootable, which is the most secure a computer can be."
- deleted 6y ago[deleted]
- bjornedstrom 6y agoAfter reading this I decided to downgrade my Ubuntu machine for now until it's figured out. There are instructions here: https://wiki.ubuntu.com/SecurityTeam/KnowledgeBase/GRUB2SecureBootBypass https://wiki.ubuntu.com/SecurityTeam/KnowledgeBase/GRUB2Secu... under the heading "DOWNGRADE `GRUB2`/`GRUB2-SIGNED` TO THE PREVIOUS VERSION FOR RECOVERY" Under the heading is a small shell script that will download the old debs for you. Note that for it to work and not have wget spam 404:s, you have to update the entire GRUB2_LP_URL and GRUB2_SIGNED_LP_URL to the links in the little table. At first glance it looks like you only have to change GRUB2_VERSION and GRUB2_SIGNED_VERSION.
- sgt 6y agoOn this subject, a few months ago I had Ubuntu servers suddenly failing due to a bizarre automatic snap update that basically took most of our Docker containers down. Did anyone experience this with production systems?