4 ms·
I'm curious what people think, but while obviously CrowdStrike caused the breakage, does the Operating System not have some responsibility in not allowing such
by ehsankia 2y ago
I'm curious what people think, but while obviously CrowdStrike caused the breakage, does the Operating System not have some responsibility in not allowing such outages to happen? Especially if it's an enterprise product?
Ideas:
1. Microsoft themselves could potentially enforce a gradual rollout on updates (did the update go through windows updates?)
2. Have better automatic recovery options, could windows have detected the bad driver and reverted it automatically?
3. just generally be more resilient, why should one bad driver take everything down?
- gruez 2y ago>1. Microsoft themselves could potentially enforce a gradual rollout on updates (did the update go through windows updates?) No. The crash was caused by a "configuration update", not through code pushed through windows update. https://www.crowdstrike.com/blog/falcon-update-for-windows-hosts-technical-details/ https://www.crowdstrike.com/blog/falcon-update-for-windows-h... >2. Have better automatic recovery options, could windows have detected the bad driver and reverted it automatically? see: https://news.ycombinator.com/item?id=41019743 https://news.ycombinator.com/item?id=41019743 >3. just generally be more resilient, why should one bad driver take everything down? It's kernel mode. With great power comes great responsibility. Blaming the OS is like blaming linux that you can sudo rm -rf /.
- xboxnolifes 2y agoI'd expect the HN crowd to not give Microsoft beef here or want them to lock down the OS more. At least, that would be consistent with usual comments about letting users control their OS. Companies chose (with some regulatory pressure) to install CrowdStrike at the kernel level after all.
- ehsankia 2y agoI didn't talk about locking down, but just generally building more resilience/recovery into the OS. Also, from the little I saw, CrowdStrike got around certifying every driver update by pushing updates through "update" files that are read and executed by the driver. That seems like a huge hack to bypass certification and likely shouldn't be allowed?
- jkrejcha 2y ago> does the Operating System not have some responsibility in not allowing such outages to happen? No. External drivers are acting as part of the operating system, their responsibility is to follow the rules for kernel drivers. Don't use-after-free or dereference otherwise bad memory (as appears to be the case here[1]), make sure you only access pageable memory at the correct IRQL, etc etc. The kernel's job is to provide services and to make sure the other components are running smoothly. The problem is, by the time that something bad like a bad dereference has occurred, other issues may start to arise. And unloading a driver or something may cause data loss and may not actually fix the underlying problem (especially if you pin the error on the wrong driver[2]). If third party software does not follow the contract, there's... really not much they can or really should do. In user mode, an access violation is given to the program when you access bad memory. This usually results in a process crash, which while annoying, may be fine. User mode programs can't[3] bring down the operating system. In kernel mode, there's no way for the OS to know that you're not going to start overwriting the disk accidentally so they made what they believe to be the safest choice--stop[4]. In any case, there's really no way for 1 to happen (since driver updates can be done externally to Microsoft), 2 is nebulous, and 3 is potentially dangerous. --- [1]: https://learn.microsoft.com/en-us/windows-hardware/drivers/debugger/bug-check-0x50--page-fault-in-nonpaged-area https://learn.microsoft.com/en-us/windows-hardware/drivers/d... [2]: For example, in the case of stack corruption. [3]: Technically they can in a couple of limited cases (but this really isn't the point). The first being killing CSRSS or another process with the "critical" kernel flag, and another by using NtShutdownSystem from the NT API. [4]: Other operating systems have taken a different philosophy. Notably Linux can be configured to allow the machine to run after a driver or other external event causes a kernel oops.
- josephcsible 2y ago> does the Operating System not have some responsibility in not allowing such outages to happen? No. The OS is supposed to guarantee that userspace programs can't crash the system like this, but CrowdStrike is an invasive kernelspace driver, not a userspace program.
- rstuart4133 2y agoPhone OS's are counter examples. Their app isolation is so strong they don't need anti virus software. Effectively all a virus gets to access to without explicit user permission is itself and the user. In particular it doesn't get to screw with the OS itself, nor with the features the OS reserves for the user like installing and uninstalling, and turning permissions on and off. Granted, just giving the virus access to the user creates problems. While the virus can't directly change it's permissions, it can socially manipulate the user into doing it for them. But even so it's fixable in the sense it doesn't take a wipe and install and the user can always just uninstall the virus. So yes, the OS can prevent this problem entire by implementing strong access controls. The problem isn't that it can't be done. The problem is that Windows doesn't do it (although the appear to be moving in that direction). Neither do all Desktop Linux's I'm familiar with, nor does macos. I think it's important to distinguish between the desktop and the kernel, because Android and ChromeOS are build with Linux and they enforce an access control system that is secure by default. I'm sure a secure OS could be built on the WinNT kernel too, but Microsoft only gives you one and it's insecure by default.
- AstralStorm 2y agoThey have tried some experiments in this area, like CLR drivers and altogether different kernel. Nothing stuck for a variety of reasons... Reverting security to an insecure state is called a downgrade attack. Can't allow that. What could be done is a better Safe Mode with Networking that would allow for secure remoting over an AD enterprise configuration... And the machines configured to automatically enter that mode. Still bit of a security issue potential.
- shadowgovt 2y agoWhat responsibility does the Linux bear if you write your own kernel module, install it, and it bricks your machine until you boot into a different OS? There are limits to the sensible responsibility of the operating system vendor / maintainer. They stopped somewhere south of "mutating core configuration of the OS itself," because if they don't, the owner doesn't control their own computer.