5 ms·
It's really difficult to evaluate the risk the CrowdStrike system imposed. Was this a confluence of improbable events or an inevitable disaster waiting to happe
by nickm12 2y ago
It's really difficult to evaluate the risk the CrowdStrike system imposed. Was this a confluence of improbable events or an inevitable disaster waiting to happen?
Some still-open questions in my mind:
- was the broken rule in the config file (C-00000291-...32.sys) human authored and reviewed or machine-generated?
- was the config file syntactically or semantically invalid according to its spec?
- what is the intended failure mode of the kernel driver that encounters an invalid config (presumably it's not "go into a boot loop")?
- what automated testing was done on both the file going out and the kernel driver code? Where would we have expected to catch this bug?
- what release strategy, if any, was in place to limit the blast radius of a bug? Was there a bug in the release gates or were there simply no release gates?
Given what we know so far, it seems much more likely that this was a "disaster waiting to happen" but I still think there's a lot more to know. I look forward to the public post-mortem.
- nickm12 2y agoTo answer some of my questions based on the "Preliminary Post Incident Review", the config file was indeed invalid, it was only checked with a (buggy) validator, and then released to the whole world at once. Never was the config file ever tested with the actual software that would read it in an actual environment like the customer machines that got this update. They don't say why it was invalid or really what the file is, but it seems like it is some kind of relatively complex set of rules that are evaluated by the kernel module. Presumably they are manually authored and reviewed and it seems possible the bug was missed in review because this was a relatively new type of rule. So this isn't a case of an incident that slipped through a rigorous testing and release process process following industry best practice, but rather a "disaster waiting to happen". Further, CrowdStrike CEO George Kurtz should have known better, considering an analogous incident happened under his watch as CTO of McAfee in 2010. https://www.crowdstrike.com/falcon-content-update-remediation-and-guidance-hub/ https://www.crowdstrike.com/falcon-content-update-remediatio...
- refulgentis 2y agoWould any of these, or even a collection of these, resolving in some direction make it highly improbable that it'll never happen again? Seems to me 3rd party code, running in the kernel, on parsed inputs, that can be remotely updated is enough to be disaster waiting to happen gestures breezily at Friday That's, in the Taleb parlance, a Fat Tony argument, but barring it being a cosmic ray causing a uncorrected bit flop during deploy, I don't think there's room to call it anything but "a disaster waiting to happen"
- slt2021 2y agokernel driver could have data check on the channel file and fail gracefully/ignore wrong file instead of BSOD. this code is executed only once during the driver initialization, so shouldn't be much overhead, but will greatly improve reliability against broken channel file
- refulgentis 2y agoThis is going to code as radical, but I always assumed it was derivable from bog-standard first principles that would fit in any economics class I sat in for my 40 credits: the natural cost of these bits we sell is zero, so in the long run, if the bar is "just write a good & tested kernel driver", there will always be one more subsequent market entrant who will go too cheap on engineering. Then, they touch the hot wire and burn down the establishment. That doesn't mean capitalism bad, but it does mean I expect only Microsoft is capable of writing and maintaining this type of software in the long run. Ex. The dentist and dental hygienist were asking me who was attacking Microsoft on Friday, and they were not going to get through to the the subtleties of 3rd kernel driver release gating strategy. MS has a very strong incentive to fix this. I don't know how they will. But I love when incentives align and assume they always will, in the long run.
- deleted 2y ago[deleted]
- nickm12 2y agoYes, if CrowdStrike was following industry best practices and this happened, it would teach us something novel about industry practices that we could learn from and use to reduce the risk of a similar scale outage happening again. If they weren't following these practices, this is kind of a boring incident with not much to be learned, despite how dramatic the scale is. Practices like staged rollout of changes exist precisely because we've learned these lessons before.
- YZF 2y agoWell, kernel code is kernel code, and kernel code in general takes input from outside the kernel. An audio driver takes audio data, a video driver might take drawing instructions, a file system interacts with files, etc. Microsoft, and others, have been releasing kernel code since forever and for the most part, not crashlooping their entire install base. My Tesla remote updates ... hmph. It doesn't feel like this is inherently impossible. It feels more like not enough design/process to mitigate the risks.
- deleted 2y ago[deleted]
- hdhshdhshdjd 2y agoWas somebody trying to install an exploit or back door and fucked up?
- TechDebtDevin 2y agoEverything is a conspiracy now eh?
- choppaface 2y agoTo be fair, the xd backdoor wasn’t immediately obvious https://www.wired.com/story/xz-backdoor-everything-you-need-to-know/ https://www.wired.com/story/xz-backdoor-everything-you-need-...
- hdhshdhshdjd 2y agoYou do remember Solarwinds right? This is an obvious high value target, so it is reasonable to entertain malicious causes. Given the number of systems infected, if you could push code that rebooted every client into a compromised state you’d still have run of some % of the lot until it was halted. That time window could be invaluable. Now, imagine if you screw up the code and just boot loop everything. I’d say business wise it’s better for crowd strike to let people think it’s an own-goal. The truth may be mundane but a hack is as reasonable a theory as “oops we pushed boot loop code to world+dog”.
- saagarjha 2y ago> The truth may be mundane but a hack is as reasonable a theory as “oops we pushed boot loop code to world+dog”. No it's not. There are many signs that point to this being a mistake. There are very few that point to it being a hack. You can't just go "oh it being a hack is one of the options therefore it is also something worth considering".
- azinman2 2y agoEspecially because if it was crowdstrike wouldn’t be apologizing and accepting blame.
- deleted 2y ago[deleted]
- Guthur 2y agoThe glaring question is how and why it was rolled out everywhere all at once? Many corporations have pretty strict rules on system update scheduling so as to ensure business continuity in case of situations like this but all of those were completely circumvented and we had fully synchronised global failure. It really does not seem like business as usual situation.
- chii 2y ago> strict rules on system update scheduling which crowdstrike gets to bypass because they claime themselves as an antivirus and malware detection platform - at least, this is what the executives they've wined and dined into the purchase contracts have been told. The update schedule is independently controlled by crowdstrike, rather than by a system admin i believe.
- xvector 2y agoCrowdStrike's reasoning is that an instantaneous global rollout helps them protect against rapidly spreading malware. However, I doubt they need an instantaneous rollout for every deployment.
- slenk 2y agoI feel like they need to at least first rollout to themselves
- kijin 2y agoWell, millions of PCs bluescreening at the same time does help stop a rapidly spreading malware. Only this time, crowdstrike itself has become indistinguishable from malware.
- imtringued 2y agoWhe I first saw news about the outage I was wondering what this malware "CrowdStrike" was. I mean, the name kind of sounds hostile.
- TeMPOraL 2y ago
- YZF 2y agoIt seems like a none of the above situation because each of those should have really minimized the chances of something like this happening. But this is pure speculation. Even the most perfect organization engineering culture can still have one thing get through... (Wasn't there some Linux incident a little back though?) Quality starts with good design, good people, etc. the process parts come much after that. I'd like to think that if you do this "right" then this sort of stuff simply can't happen. If we have organization/culture/engineering/process issues then we're likely not going to get an in-depth public most-mortem. I'd love to get one just for all of us to learn from it. Let's see. Given the cost/impact having something like the Challenger investigation with some smart uninvolved people would be good.
- 7952 2y agoIn a world of complex systems a "confluence of improbable events" is the same thing as "a disaster waiting to happen". Its the swiss cheese model of failure. Y
- k8sToGo 2y agoEvery system can only survive so many improbable events. Even in aviation.