8 ms·
This is some very poor journalism. The linux issues are so, so very different from the windows BSOD issue. The redhat kernel panics were caused by a bug in the
by roblabla 2y ago
This is some very poor journalism. The linux issues are so, so very different from the windows BSOD issue.
The redhat kernel panics were caused by a bug in the kernel ebpf implementation, likely a regression introduced by a rhel-specific patch. Blaming crowdstrike for this is stupid (just like blaming microsoft for the crowdstrike bsod is stupid).
For background, I also work on a product using eBPFs, and had kernel updates cause kernel panics in my eBPF probes.
In my case, the panic happened because the kernel decided to change an LSM hook interface, adding a new argument in front of the others. When the probe gets loaded, the kernel doesn’t typecheck the arguments, and so doesn’t realise the probe isn’t compatible with the new kernel. When the probe runs, shit happens and you end up with a kernel panic.
eBPF probes causing kernel panics are almost always indication of a kernel bug, not a bug in the ebpf vendor. There are exceptions of course (such as an ebpf denying access to a resource causing pid1 to crash). But they’re very few.
- josefx 2y ago> likely a regression introduced by a rhel-specific patch. Blaming crowdstrike for this is stupid (just like blaming microsoft for the crowdstrike bsod is stupid). Yeah, it isn't as if crowdstrike was specifically advertising certified support for RedHat Linux and related products. https://www.crowdstrike.com/partners/falcon-for-red-hat/ https://www.crowdstrike.com/partners/falcon-for-red-hat/
- dtx1 2y agoBut being certified for RedHat Linux doesn't protect you from Bugs in the RedHat Kernel. That's on RedHat.
- michaelt 2y agoBack in The Good Old Days, an OS vendor would release a beta version and software vendors would test against it and fix problems before the stable OS version was released. Obviously OS updates come out a lot more often these days than they used to - but we're also better at test automation than ever before, and beta software is easier to get than ever. It sure would be nice if companies that decide to produce kernel modules and to support certain OSes could test those kernel modules against those OSes at the beta stage.
- roblabla 2y ago1.An eBPF probe is not a kernel module. An eBPF probe should never cause kernel panics. 2. RHEL didn't provide beta kernels before very recently, as far as I can tell. 3. Even if you caught an error then, you're still at the mercy of RHEL to fix it. If RHEL breaks a feature, you report it to them, and they decide to ship anyways... well, your product will still kpanic. I'm not talking hypotheticals: I haven't seen RHEL do that, but I've seen other distros do it.
- CRConrad 2y ago> An eBPF probe is not a kernel module. But if it runs on the same privilege level as the rest of the kernel, then isn't it, for the purposes of this discussion, in effect "a kernel module"?
- kragen 2y agoit doesn't—the semantics of ebpf confine it—so it isn't
- CRConrad 2y agoEither 1) Those Crowdstrike unit files aren't ebpf probes, so the whole subject of ebpf probes is irrelevant here; or 2) They're obviously able to stop the rest of the kernel from even booting up (as Crowdstrike so convincingly demonstrated millions of times over[1]), so yes, they do indeed have at least as much power as any other bit of the kernel. Either way, hunting around for nits to pick is a bit pathetic. [1]: In July 0000002024...
- kragen 2y agodenial of service is not the same thing as arbitrary code execution, and that goes double in kernel mode, but yes, it does seem that the linux implementation of ebpf had buggy sandboxing; i don't think allowing clownstrike to prevent booting was part of the intended objective i wasn't hunting around for nits to pick; i was hunting around to see if you'd ever contributed any useful comments to the site. instead i found you making authoritative pronouncements about ebpf that were so wrong that you had evidently never read so much as a one-line summary of what ebpf was for. do you have a more promising historical comment to offer? perhaps something where people complimented your contribution as being informative? have you ever made a worthwhile comment on hn? on thursday, wahern posted this comment https://news.ycombinator.com/item?id=41061179 https://news.ycombinator.com/item?id=41061179 where they traced through the illumos/opensolaris source code to track down how a peculiar solaris interprocess communication mechanism worked, an investigation i had started but gotten stuck on. why can't you make comments like that instead of harassing me about how i format my comments? the reason i'm asking is because i'd like to be able to talk to more people like wahern, but most of them avoid this site. a major reason why is that comments here frequently receive vacuous, aggressive responses like the comment you made the day before in https://news.ycombinator.com/item?id=41056718 https://news.ycombinator.com/item?id=41056718, where you launched a personal attack on me because you didn't like how i was formatting my comments i'd like you to ⓐ apologize for doing that (this is not the first time you've done that to me personally; so far i haven't looked through your comment history far enough to find out how many other people you have a history of repeatedly harassing) and ⓑ commit to not doing it again because i'm sure you're capable of making comments that make the site better instead of worse
- roblabla 2y agoYes, and? They probably do test their software on RHEL. But how are they supposed to prevent a bug in a newly released kernel update? You can't test your software on future updates that aren't out yet. If RHEL breaks some core functionality you depend on, in a newly released update, you can't really do much to prevent breakage, even with the best QA in the world. At best, they could have caught it as soon as RHEL published the new kernel... but by then it's already too late, all your currently-deployed probes now have a ticking time bomb, and need to be updated before the RHEL kernel update is applied, lest you kernel panic.
- hsbauauvhabzb 2y agoMaybe by not loading the module into unknown kernels in the first place? If say you support a distro, you can’t turn around and complain that supporting the newest version is hard, no matter who caused the problem. Plenty of products say ‘this works on $x but it’s not officially supported’.
- broknbottle 2y agoThis was their newer eBPF falcon sensor that was trying to load a bpf program in the kernel and triggered kernel panic. This shouldn’t have happened and was definitely a bug in the kernel. For the kernel mode, their software will flag an unknown kernel as unsupported and go into a reduced functionality mode (rfm). The idiots didn’t know that RH E4S was a thing for like 3+ years.. I’m still baffled by how clueless most of the security people and vendors are when it comes to backporting and different streams / channels that are offered by multiple Linux OS vendors. https://access.redhat.com/solutions/7001909 https://access.redhat.com/solutions/7001909
- roblabla 2y agoAgain: this is not a kernel module. eBPF probes are meant to be Compile Once, Run Everywhere, that's their whole point! https://facebookmicrosites.github.io/bpf/blog/2020/02/19/bpf-portability-and-co-re.html https://facebookmicrosites.github.io/bpf/blog/2020/02/19/bpf... If you expect software to be future-bug-proof, well, I guess you live in a far better world than I do. If you advertise your software to be compatible with RHEL, but a glibc bug gets in and causes your sw to crash for a couple of days, before RHEL realises the problem and fixes it, does that mean your software should instantly no longer be advertised as RHEL compatible? That'd make things a lot more confusing, if you ask me.
- mbesto 2y ago> just like blaming microsoft for the crowdstrike bsod is stupid Wait, how is this stupid? Unless I'm missing something, wasn't the patch part of a Microsoft payload that included an update to Crowdstrike? Surely Crowdstrike is culpable, but that doesn't completely absolve Microsoft of any responsibility, as its their payload.
- roblabla 2y agoDo you have a source for this? It's the first time I hear of this. From what I've understood (perhaps wrongly), the error came from the CrowdStrike driver (csagent.sys) having bugs in their configuration parser that could cause it to BSOD. CrowdStrike pushed a corrupted configuration (the CS-000whatever.sys we're told to delete) that hit that bug. I'm not sure how Microsoft fits in this story.
- mbesto 2y agoJust read more into it. You're correct. I think it would be dumb to solely blame MS, but I don't think you can completely absolve them. this comment right here sums it up: > Sure, but Windows shares some portion of the blame for allowing third-party security vendors to “shit in the kernel”. https://news.ycombinator.com/item?id=41006176 https://news.ycombinator.com/item?id=41006176
- roblabla 2y agoYeah, the fact that Windows requires kernel-level access to be able to do EDR stuff is really unfortunate. MacOS has been very successful with their userspace EndpointSecurity Framework for this purpose. On the other hand, Linux is similarly crippled: eBPF LSM are fairly recent and don't work everywhere (I'm looking at you Ubuntu[0]), and the only real alternative if you want to be able to block processes is a kernel module. Which comes with the same dangers as Windows. [0]: https://bugs.launchpad.net/ubuntu/+source/linux/+bug/2054810 https://bugs.launchpad.net/ubuntu/+source/linux/+bug/2054810
- sgift 2y ago
- watt 2y agoThere should not be any software caused crashes during operation of software. Every NPE that is not caused by hardware issue, is a null pointer not properly handled. Software needs to handle their null checks. Missing a check (or precondition, or validation) is squarely on Microsoft. > always indication of a kernel bug and before but: > blaming microsoft for the crowdstrike bsod is stupid and who owns the kernel in windows land? Microsoft. how is it stupid to blame Microsoft for not making kernel safe?
- Diggsey 2y agoMicrosoft don't own the kernel in that sense: anyone can write kernel drivers for windows... While there are some things that the kernel can do to protect against a bad driver, it's not a security boundary, so ultimately bad code can cause crashes. AIUI, Microsoft actually has good tooling for validating drivers before they are deployed, but it requires that you actually run the validation...
- mebeim 2y agoLet me give you an analogy: Volvo is known to manufacture very safe cars. Now let's say I drive a Volvo car with a box of dynamite on the passenger seat. I stop at a red light but hit the brake a bit too hard and the box of dynamite falls and causes an explosion, disintegrating everything in a 20-foot radius. So whose fault was it? Volvo? > Missing a check (or precondition, or validation) is squarely on Microsoft. Missing a check for presence of dynamite before allowing me to start the car is squarely on Volvo! You see how silly that sounds? Now, back to being serious: MS cannot possibly control and validate everything you decide to install and run on your system, specially if the things you install are kernel drivers. It is simply impossible. If you install a kernel driver developed by a 3rd party company, and that driver crashes your system because the devs at that company forgot to perform proper validation of data, well... that's on them. Even if MS wanted, they wouldn't be able to verify the soundness of any piece of code that is installed as a driver and runs with kernel level privileges. That'd require solving the halting problem.
- roblabla 2y ago@watt there's a big difference here. eBPF is a bytecode that is interpreted in the kernel, with the explicit goal to allow writing code that executes at the kernel-level in a safe way. Any kernel panic (again, short of pid1 kills) is considered a bug, and could even potentially be exploited to gain capabilities in some cases. Here, the kernel explicitly says "this is safe", so any problem within is a bug in the kernel. In contrast, a kernel module/driver is just some third-party code that is loaded in the kernel. Here, all bets are off: it is up to the third-party to do their job properly and make sure their code is correct. In this case, CrowdStrike explicitly opted into writing a kernel module, and then failed to, as you say, "handle their null check". The segfault wasn't in Windows code, it was in CrowdStrike code that lives in the kernel. Crowdstrike should have handled their nullcheck, failed, and that will lead to a BSOD. To be clear: the only way microsoft could make the kernel safer here is by disallowing kernel modules entirely. While there is an argument to be made that this could be a good idea, it is a bit beside the point.
- xyzzy123 2y agoIt's not clear to me they are so different but maybe I am not "sufficiently smart". To me this feels like a complicated question - both Linux and Windows organisations are quite good at kernel reliability engineering even though quite different organisational structures and engineering approaches are involved. Yes "the wrong people were trusted" but I don't see how we can completely solve this with engineering.
- roblabla 2y ago> It's not clear to me they are so different but maybe I am not "sufficiently smart". They're different because linux promises "eBPF are safe and cannot crash the kernel", and failed to deliver on that, while Microsoft says "drivers are all-powerful and as such must be written with care", and CrowdStrike did not heed this warning. > Yes "the wrong people were trusted" but I don't see how we can completely solve this with engineering. I mean, we could solve the "third party software fucks the kernel up" problem easily with engineering: providing userspace APIs to do stuff that currently need kernelspace access. There's no inherent reason security products (or, really, any products) needs to live in the kernel, it's just that there are no APIs to do this job, so security products have to go there. If Microsoft provided a good API doing what the custom drivers currently do, most security products would drop their driver in a heartbeat. For instance, macOS fixed this exact issue a couple years ago by introducing Endpoint Security Framework, a userspace API that allows watching a bunch of events, and authorizing whether they should be allowed or blocked. It's a well-designed API that should obsolete the need for kernelspace access in security products.
- j2bryson 2y agoSo what happened with the linux bug? Presumably people fixed the OS side problem straight away?
- roblabla 2y agokernel-5.14.0-427.13.1.el9_4 broke it. It was released in Apr 30, 2024, with RHEL 9.4 (this was the RHEL 9.4 release kernel). According to the comments on https://access.redhat.com/solutions/7068083 https://access.redhat.com/solutions/7068083, RHEL became aware of the issue on May 3, 2024. A workaround was identified (configuring CS to use the kernel module backend instead of the ebpf backend) on May 9, 2024. RHEL then fixed it in kernel-5.14.0-427.18.1.el9_4, in May 23, 2024. So the bug was fixed in ~20 days from the moment it was reported. It's unclear whether this issue was caused by a RHEL-specific backport/patch or was also present in mainline kernels.