18 ms·
Preliminary Post Incident Review
- cataflam 2y agoBesides missing the actual testing (!), the staged rollout (!), looks like they also weren't fuzzing this kernel driver that routinely takes instant worldwide updates. Oops.
- l00tr 2y agocheck their developer github, "i write kernel-safe bytecode interpreters" :D, [link redacted]
- brcmthrowaway 2y agoHe Codes With Honor(tm)
- l00tr 2y ago[dead]
- Scaevolus 2y ago"problematic content"? It was a file of all zero bytes. How exactly was that produced?
- Zironic 2y agoIf I had to guess blindly based on their writeup, it would seem that if their Content Configuration System is given invalid data, instead of aborting the template, it generates a null template. To a degree it makes sense because it's not unusual for a template generator to provide a null response if given invalid inputs however the Content Validator then took that null and published it instead of handling the null case as it should have.
- jiggawatts 2y agoReturning null instead of throwing an exception when an error occurs is the quality of programming I see from junior outsourced developers. “if (corrupt digital signature) return null;” is the type of code I see buried in authentication systems, gleefully converting what should be a sudden stop into a shambling zombie of invalid state and null reference exceptions fifty pages of code later in some controller that’s already written to the database on behalf of an attacker. If I peer into my crystal ball I see a vision of CrowdStrike error handling code quality that looks suspiciously the same. (If I sound salty, it’s because I’ve been cleaning up their mess since last week.)
- whoknowsidont 2y ago>Returning null instead of throwing an exception when an error occurs is the quality of programming I see from junior outsourced developers. This is kernel code, most likely written in C (and regardless of language, you don't really do exceptions in the kernel at all for various reasons). Returning NULL or ERR_PTR (in the case of linux) is absolutely one of the most standard, common, and enforced ways of indicating an error state in kernel code, across many OS's. So it's no surprise to see the pattern here, as you would expect.
- chrisjj 2y agoThe've said the crash was not related to those zero bytes. https://www.crowdstrike.com/blog/falcon-update-for-windows-hosts-technical-details/ https://www.crowdstrike.com/blog/falcon-update-for-windows-h...
- romwell 2y agoThis reads like a bunch of baloney to obscure the real problem. The only relevant part you need to see: >Due to a bug in the Content Validator, one of the two Template Instances passed validation despite containing problematic content data. Problematic content? Yeah, this is telling exactly nothing. Their mitigation is "ummm we'll test more and maybe not roll the updates to everyone at once", without any direct explanation on how that would prevent this from happening again. Conspicuously absent: — fixing whatever produced "problematic content" — fixing whatever made it possible for "problematic content" to cause "ungraceful" crashes — rewriting code so that the Validator and Interpreter would use the same code path to catch such issues in test — allowing the sysadmins to roll back updates before the OS boots — diversifying the test environment to include actual client machine configurations running actual releases as they would be received by clients This is a nothing sandwich, not an incident review.
- Zironic 2y ago>Add additional validation checks to the Content Validator for Rapid Response Content. A new check is in process to guard against this type of problematic content from being deployed in the future. >Enhance existing error handling in the Content Interpreter. They did write that they intended to fix the bugs in both the validator and the interpreter. Though it's a big mystery to me and most of the comments on the topic how an interpreter that crashes on a null template would ever get into production.
- romwell 2y ago>They did write that they intended to fix the bugs I strongly disagree. Add additional validation and enhance error handling say as much as "add band-aids and improve health" in response to a broken arm. Which is not something you'd want to hear from a kindergarten that sends your kid back to you with shattered bones. Note that the things I said were missing are indeed missing in the "mitigation". In particular, additional checks and "enhanced" error handling don't address: — the fact that it's possible for content to be "problematic" for interpreter, but not the validator; — the possibility for "problematic" content to crash the entire system still remaining; — nothing being said about what made the content "problematic" (spoiler: a bunch of zeros, but they didn't say it), how that content was produced in the first place, and the possibility of it happening in the future still remaining; — the fact that their clients aren't in control of their own systems, have no way to roll back a bad update, and can have their entire fleet disabled or compromised by CrowdStrike in an instant; — the business practices and incentives that didn't result in all their "mitigation" steps (as well as steps addressing the above) being already implemented still driving CrowdStrike's relationship with its employees and clients. The latter is particularly important. This is less a software issue, and more an organizational failure. Elsewhere on HN and reddit, people were writing that ridiculous SLA's, such as "4 hour response to a vulnerability", make it practically impossible to release well-tested code, and that reliance on a rootkit for security is little more than CYA — which means that the writing was on the wall, and this will happen again. You can't fix bad business practices with bug fixes and improved testing. And you can't fix what you don't look into. Hence my qualification of this "review" as a red herring.
- deleted 2y ago[deleted]
- romwell 2y agoCopying my content from the duplicate thread[1] here: This reads like a bunch of baloney to obscure the real problem. The only relevant part you need to see: >Due to a bug in the Content Validator, one of the two Template Instances passed validation despite containing problematic content data. Problematic content? Yeah, this is telling exactly nothing. Their mitigation is "ummm we'll test more and maybe not roll the updates to everyone at once", without any direct explanation on how that would prevent this from happening again. Conspicuously absent: — fixing whatever produced "problematic content" — fixing whatever made it possible for "problematic content" to cause "ungraceful" crashes — rewriting code so that the Validator and Interpreter would use the same code path to catch such issues in test — allowing the sysadmins to roll back updates before the OS boots — diversifying the test environment to include actual client machine configurations running actual releases as they would be received by clients This is a nothing sandwich, not an incident review. [1] https://news.ycombinator.com/item?id=41053703 https://news.ycombinator.com/item?id=41053703
- dang 2y ago> Copying my content from the duplicate thread[1] here Please don't do this! It makes merging threads a pain because then we have to find the duplicate subthreads (i.e. your two comments) and merge the replies as well. Instead, if you or anyone will let us know at hn@ycombinator.com which threads need merging, we can do that. The solution is deduplication, not further duplication!
- rurban 2y agoThey bypassed the tests and staged deployment, because their previous update looked good. Ha. What if they implemented a release process, and follow it? Like everyone else does. Hackers at the workplace, sigh.
- CommanderData 2y agoThey know better obviously, transcending process and bureaucracy.
- rurban 2y agoSame thing happened with Falcon on Debian before. Later they admitted that they didn't test some platforms they were releasing. Never heard of Docker? How can you keep on with such a Q&R manager? He'll cost them billions
- shepherdjerred 2y agoDocker wouldn't help with testing kernel modules. You'd need a VM.
- fulafel 2y agoAlso it must have been a manual testing effort, otherwise there would be no motive to skip it. IOW, missing test automation.
- throwaway7ahgb 2y agoWhere do you see that, it looks like there was a bug in the template tester? Or you mean the manual tests?
- kasabali 2y ago> Based on the testing performed before the initial deployment of the Template Type (on March 05, 2024), trust in the checks performed in the Content Validator, and previous successful IPC Template Instance deployments, these instances were deployed into production.
- lopkeny12ko 2y ago[flagged]
- nemetroid 2y agoNo, the Twitter poster is still wrong.
- bdjsiqoocwk 2y agoIt's called Twitter.
- joenot443 2y agoNo, the name’s been changed.
- bdjsiqoocwk 2y agoNo it hasn't.
- justusthane 2y agoI'm not a fan of Musk or of the-platform-formerly-known-as-Twitter, but I'm not sure how you can insist that the name hasn't been changed.
- mc32 2y agoHow can these companies be certified and compliant, etc., and then in practice have horrible SDLC? What was the impact of diverse teams (offshoring)? Often companies don’t have necessary checks to ensure disparateness of teams does not impact quality. Maybe it was zero or maybe it was more.
- hulitu 2y ago> How can these companies be certified and compliant, etc., and then in practice have horrible SDLC? Checklists ?
- nodesocket 2y agoWhy do they insist on using what sounds like military pseudo jargon throughout the document? ex. sensors? I mean how about hosts, machines, clients?
- com 2y agoIt’s endemic in the tech security industry - they’ve been mentally colonised by ex-mil and ex-law enforcement (wannabe mil) folks for a long time. I try to use social work terms and principles in professional settings, which blows these people’s minds. Advocacy, capacity evaluation, community engagement, cultural competencies, duty of care, ethics, evidence-based intervention, incentives, macro-, mezzo- and micro-practice, minimisation of harm, respect, self concept, self control etc etc It means that my teams aren’t focussed on “nuking the bad guys from orbit” or whatever, but building defence in depth and indeed our own communities of practice (hah!), and using psychological and social lenses as well as tech and adversarial ones to predict, prevent and address disruptive and dangerous actors. YMMV though.
- phaedrus 2y agoEven computer security itself is a metaphor (at least in its inception). I often wonder what if instead of using terms like access, key, illegal operation, firewall, etc. we'd instead chosen metaphors from a different domain, for example plumbing. I'm sure a plumbing metaphor could also be found for every computer security concern. Would be so quick to romanticize as well as militarize a field dealing with "leaks," "blockages," "illegal taps," and "water quality"?
- coremoff 2y agoSuch a disingenuous review; waffle and distraction to hide the important bits (or rather bit: bug in content validator) behind a wall of text that few people are going to finish. If this is how they are going to publish what happened, I don't have any hope that they've actually learned anything from this event. > Throughout this PIR, we have used generalized terminology to describe the Falcon platform for improved readability Translation: we've filled this PIR with technobable so that when you don't understand it you won't ask questions for fear of appearing slow.
- notepad0x90 2y ago> "behind a wall of text that few people are going to finish." heh? it's not that long and very readable.
- coremoff 2y agoI disagree; it's much longer than it needs to be, is filled with pseudo-technoese to hide that there's little of consequence in there, and the tiny bit of real information in there is couched with distractions and unnecessary detail. As I understand it, they're telling us that the outage was caused by an unspecified bug in the "Content Validator", and that the file that was shipped was done so without testing because it worked fine last time. I think they wrote what they did because they couldn't publish the above directly without being rightly excoriated for it, and at least this way a lot of the people reading it won't understand what they're saying but it sounds very technical.
- notepad0x90 2y agono, it's one of most well written PIR's I've seen. It establishes terms and procedures after communicating that this isn't an RCA, then they detail the timeline of tests and deployments done and what went wrong. They were not excessively verbose or terse. This is the right way of communicating to the intended audience. It is both technical people, executives and law makers alike that will be reading this. They communicated their findings clearly without code, screenshots, excessive historical details and other distractions.
- CommanderData 2y ago"We didn't properly test our update." Should be the tldr. On threads there's information about CrordStrike slashing QA team numbers, whether that was a factor should be looked at.
- hulitu 2y agoThey write perfect software. Why should they test it ? /s
- Ukv 2y agoA summary, to my understanding: * Their software reads config files to determine which behavior to monitor/block * A "problematic" config file made it through automatic validation checks "due to a bug in the Content Validator" * Further testing of the file was skipped because of "trust in the checks performed in the Content Validator" and successful tests of previous versions * The config file causes their software to perform an out-of-bounds memory read, which it does not handle gracefully
- Narretz 2y ago* Further testing of the file was skipped because of "trust in the checks performed in the Content Validator" and successful tests of previous versions that's crazy. How costly can it be to test the file fully in a CI job? I fail to see how this wasn't implemented already.
- modestygrime 2y agoJust reeks of incompetence. Do they not have e2e smoketests of this stuff?
- denton-scratch 2y ago> How costly can it be to test the file fully in a CI job? It didn't need a CI job. It just needed one person to actually boot and run a Windows instance with the Crowdstrike software installed: a smoke test. TFA is mostly an irrelevent discourse on the product architecture, stuffed with proprietary Crowdstrike jargon, with about a couple of paragraphs dedicated to the actual problem; and they don't mention the non-existence of a smoke test. To me, TFA is not a signal that Crowdstrike has a plan to remediate the problem, yet.
- hrpnk 2y agoThey mentioned they do dogfooding. Wonder why it did not work for this update.
- 2y ago
- red2awn 2y ago> How Do We Prevent This From Happening Again? > Software Resiliency and Testing > * Improve Rapid Response Content testing by using testing types such as: > * Local developer testing So no one actually tested the changes before deploying?!
- Narretz 2y agoAnd why is it "local developer testing" and not CI/CD. This makes them look like absolute amateurs.
- belter 2y ago> This makes them look like absolute amateurs. This applies also to all Architects and CTO's at all these Fortune 500 companies, who allowed these self updating systems into their critical systems. I would offer a copy of Antifragile to each of these teams: https://en.wikipedia.org/wiki/Antifragile_(book) https://en.wikipedia.org/wiki/Antifragile_(book) "Every captain goes down with every ship"
- acdha 2y agoArchitects likely do not have a choice. These things are driven by auditors and requirements for things like insurance or PCI and it’s expensive to protest those. I know people who’ve gone full serverless just to lop off the branches of the audit tree about general purpose server operating systems, and now I’m wondering whether anyone is thinking about iOS/ChromeOS for the same reason. The more successful path here is probably demanding proof of a decent SDLC, use of memory-safe languages, etc. in contract language.
- belter 2y ago> Architects likely do not have a choice. Architects don't have a choice, CTO are well paid to golf with the CEO and delegate to their teams, Auditors just audit but are not involved with the technical implementations, Developers just develop according to the Spec, and Security team just are a pain in the ass. Nobody owns it... Everybody get's well paid, and at the end we have to get lessons learned...It's a s*&^&t show...
- nine_zeros 2y agoWill managers continue to push engineers even when engineers advise to go slower or no?
- bobwaycott 2y agoAlways.
- Cyphase 2y agoLots of words about improving testing of the Rapid Response Content, very little about "the sensor client should not ever count on the Rapid Response Content being well-formed to avoid crashes". > Enhance existing error handling in the Content Interpreter. That's it. Also, it sounds like they might have separate "validation" code, based on this; why is "deploy it in a realistic test fleet" not part of validation? I notice they haven't yet explained anything about what the Content Validator does to validate the content. > Add additional validation checks to the Content Validator for Rapid Response Content. A new check is in process to guard against this type of problematic content from being deployed in the future. Could it say any less? I hope the new check is a test fleet. But let's go back to, "the sensor client should not ever count on the Rapid Response Content being well-formed to avoid crashes".
- hun3 2y agoIs error handling enough? A perfectly valid rule file could hang (but not outright crash) the system, for example.
- ReaLNero 2y agoPerhaps set a timeout on the operation then? Given this is kernel it's not as easy as userspace, but I'm sure you could request to set a interrupt on a timer.
- throwanem 2y agoIf the rules are Turing-complete, then sure. I don't see enough in the report to tell one way or another; the way rules are made to sound as if filling templates about equally suggests either (if templates may reference other templates) and there is not a lot more detail. Halting seems relatively easy to manage with something like a watchdog timer, though, compared to a sound, crash- and memory-safe* parser for a whole programming language, especially if that language exists more or less by accident. (Again, no claim; there's not enough available detail.) I would not want to do any of this directly on metal, where the only safety is what you make for yourself. But that's the line Crowdstrike are in. * By EDR standards, at least, where "only" one reboot a week forced entirely by memory lost to an unkillable process counts as exceptionally good.
- Cyphase 2y agoDirect link to the PIR, instead of the list of posts: https://www.crowdstrike.com/blog/falcon-content-update-preliminary-post-incident-report/ https://www.crowdstrike.com/blog/falcon-content-update-preli...
- Cyphase 2y agoThe article link has been updated to that; it used to be the "hub" page at https://www.crowdstrike.com/falcon-content-update-remediation-and-guidance-hub/ https://www.crowdstrike.com/falcon-content-update-remediatio... Some updates from the hub page: They published an "executive summary" in PDF format: https://www.crowdstrike.com/wp-content/uploads/2024/07/CrowdStrike-PIR-Executive-Summary.pdf https://www.crowdstrike.com/wp-content/uploads/2024/07/Crowd... That includes a couple of bullet points under "Third Party Validation" (independent code/process reviews), which they added to the PIR on the hub page, but not on the dedicated PIR page. > Updated 2024-07-24 2217 UTC > ### Third Party Validation > - Conduct multiple independent third-party security code reviews. > - Conduct independent reviews of end-to-end quality processes from development through deployment.
- squirrel 2y agoThere’s only one sentence that matters: "Provide customers with greater control over the delivery of Rapid Response Content updates by allowing granular selection of when and where these updates are deployed." This is where they admit that: 1. They deployed changes to their software directly to customer production machines; 2. They didn’t allow their clients any opportunity to test those changes before they took effect; and 3. This was cosmically stupid and they’re going to stop doing that. Software that does 1. and 2. has absolutely no place in critical infrastructure like hospitals and emergency services. I predict we’ll see other vendors removing similar bonehead “features” very very quietly over the next few months.
- hello_moto 2y ago> I predict we’ll see other vendors removing similar bonehead “features” very very quietly over the next few months. Absolutely this is what will happen. I don't know much about the practice of AV definition-like feature across Cybersecurity but I would imagine there might be a possibility that no vendors do rolling update today because it involves Opt-in/Opt-out which might influence the vendor's speed to identify attack which in turns affect their "Reputation" as well. "I bought Vendor-A solution but I got hacked and have to pay Ransomware" (with a side note: because I did not consume the latest critical update of AV definition) is what Vendors worried. Now that this Global Outage happened, it will change the landscape a bit.
- XlA5vEKsMISoIln 2y ago>Now that this Global Outage happened, it will change the landscape a bit. I seriously doubt that. Questions like "why should we use CrowdStrike" will be met with "suppose they've learned their lesson".
- hello_moto 2y agoI'm referring to the landscape how current Cybersecurity vendors deliver "detection definition" (for lack of better phrase) to their customers. If you don't send them fast to your customer and your customer gets compromised, your reputation gets hit. If you send them fast, this BSOD happened. It's more like damn if you do, damn if you don't.
- brianmback 2y ago[flagged]
- deleted 2y ago[deleted]
- duped 2y agoHere is my summary with the marketing bullshit ripped out. Falcon configuration is shipped with both direct driver updates ("sensor content"), and out of band ("rapid response content"). "Sensor Content" are scripts (*) that ship with the driver. "Rapid response content" are data that can be delivered dynamically. One way that "Rapid Response Content" is implemented is with templated "Sensor Content" scripts. CrowdStrike can keep the behavior the same but adjust the parameters by shipping "channel" files that fill in the templates. "Sensor content", including the templates, are a part of the normal test and release process and goes through testing/verification before being signed/shipped. Customers have control over rollouts and testing. "Rapid Response Content" is deployed through a different channel that customers do not have control over. Crowdstrike shipped a broken channel file that passed validation but was not tested. They are going to fix this by adding testing of "rapid response" content updates and support the same rollout logic they do for the driver itself. (*) I'm using the word "script" here loosely. I don't know what these things are, but they sound like scripts. --- In other words, they have scripts that would crash given garbage arguments. The validator is supposed to check this before they ship, but the validator screwed it up (why is this a part of release and not done at runtime? (!)). It appears they did not test it, they do not do canary deployments or support rollout of these changes, and everything broke. Corrupting these channel files sounds like a promising way to attack CS, I wonder if anyone is going down that road.
- hello_moto 2y ago> Corrupting these channel files sounds like a promising way to attack CS, I wonder if anyone is going down that road. Would have happened long time ago if it was that easy no?
- duped 2y agoHow do we know it hasn't?
- hello_moto 2y agoIf it happened, the industry would have known by now. The group behind it will come out to the public.
- EricE 2y agoA file full of zeros is an "undetected error"? Good grief.
- dmazzoni 2y agoIt wasn't a file full of zeros that caused the problem. While some affected users did have a file full of zeros, that was actually a result of the system in the process of trying to download an update, and not the version of the file that caused the crash.
- deleted 2y ago[deleted]
- jvreeland 2y agoI really dislike reading website that take over half the screen and make me read off to the side like this. I can fix it by zooming in but I don't understand why they thought making the navigation take up that much of the screen or not be collapsable was a good move.
- sgammon 2y ago1) Everything went mostly well 2) The things that did not fail went so great 3) Many many machines did not fail 4) macOS and Linux unaffected 5) Small lil bug in the content verifier 6) Please enjoy this $10 gift card 7) Every windows machine on earth bsod'd but many things worked
- mikequinlan 2y agoRegarding the gift card, TechCrunch says "On Wednesday, some of the people who posted about the gift card said that when they went to redeem the offer, they got an error message saying the voucher had been canceled. When TechCrunch checked the voucher, the Uber Eats page provided an error message that said the gift card “has been canceled by the issuing party and is no longer valid.”" https://techcrunch.com/2024/07/24/crowdstrike-offers-a-10-apology-gift-card-to-say-sorry-for-outage/ https://techcrunch.com/2024/07/24/crowdstrike-offers-a-10-ap...
- sgammon 2y agoThere's a KB up about this now. To use your voucher, reboot into safe mode and...
- mikequinlan 2y agoOn another forum a person replied… >The system to redeem the card is probably stuck in a boot loop
- rm445 2y agoFun post, but I'll state the obvious because I think many people do believe that every Windows machine BSOD'd. It was only ones with Crowdstrike software. Which is apparently very common but isn't actually pre-installed by Microsoft in Windows, or anything like that. Source: work in a Windows shop and had a normal day.
- sgammon 2y agoTrue, and definitely worth a mention. This is only Microsoft's fault insofar as it was possible at all to crash this way, this broadly, with so little recourse via remote tooling.
- gostsamo 2y agoCowards. Why don't you just stand up and admit that you didn't bother testing everything you send to production? Everything else is smoke and the smell of sulfur.
- NoPicklez 2y agoBecause producing smoke and the smell of sulfur is how you keep your business afloat after an incident like this Getting on your knees and admitting terrible fault with apologies galore isn't going to garner you any more sympathy.
- deleted 2y ago[deleted]
- yuliyp 2y ago> Why don't you just stand up and admit that you didn't bother testing everything you send to production? The "What Happened on July 19, 2024?" section combined with the "Rapid Response Content Deployment" make it very clear to anyone reading that that is the case. Similarly, the discussion of the sensor release process in "Sensor Content" and lack of discussion of a release process in the "Rapid Response Content" section solidify the idea that they didn't consider validated rapid response content causing bad behavior as a thing to worry about.
- aeyes 2y agoDo you see how they only talk about technical changes to prevent this from happening again? To me this was a complete failure on the process and review side. If something so blatantly obvious can slip through, how could ever I trust them to prevent an insider from shipping a backdoor? They are auto updating code with the highest privileges on millions of machines. I'd expect their processes to be much much more cautious.
- m3kw9 2y agoStill have kernel access
- anonu 2y agoIn my experience with outages, usually the problem lies in some human error not following the process: Someone didn't do something, checks weren't performed, code reviews were skipped, someone got lazy. In this post mortem there are a lot of words but not one of them actually explains what the problem was. which is: what was the process in place and why did it fail? They also say a "bug in the content validation". Like what kind of bug? Could it have been prevented with proper testing or code review?
- yuliyp 2y ago> In my experience with outages, usually the problem lies in some human error not following the process Everyone makes mistakes. Blaming them for making those mistakes doesn't help prevent mistakes in the future. > what kind of bug? Could it have been prevented with proper testing or code review? It doesn't matter what the exact details of the bug are. A validator and the thing it tries to defend being imperfect mates is a failure mode. They happened to trip that failure mode spectacularly. Also saying "proper testing and code review" in a post-mortem is useless like 95% of the time. Short of a culture of rubber-stamping and yolo-merging where there is something to do, it's a truism that any bug could have been detected with a test or caught by a diligent reviewer in code review. But they could also have been (and were) missed. "git gud" is not an incident prevention strategy, it's wishful thinking or blaming the devs unlucky enough to break it. More useful as follow-ups are things like "this type of failure mode feels very dangerous, we can do something to make those failures impossible or much more likely to be caught"
- flanked-evergl 2y ago> Everyone makes mistakes. Blaming them for making those mistakes doesn't help prevent mistakes in the future. You can't reliably fix problems you don't understand.
- gwd 2y ago> ...what was the process in place and why did it fail? It appears the process was: 1. Channel files are considered trusted; so no need to sanity-check inputs in the sensor, and no need to fuzz the sensor itself to make sure it deals gracefully with corrupted channel files. 2. Channel files are trusted if they pass a Content Validator. No additional testing is needed; in particular, the channel files don't even need to be smoke-tested on a real system. 3. A Content Validator is considered 100% effective if it has been run on three previous batches of channel files without incident. Now it's possible that there were prescribed steps in the process which were not followed; but those too are to be expected if there is no automation in place. A proper process requires some sort of explicit override to skip parts of it.
- 1970-01-01 2y ago>When received by the sensor and loaded into the Content Interpreter, problematic content in Channel File 291 resulted in an out-of-bounds memory read triggering an exception. Wasn't 'Channel File 291' a garbage file filled with null pointers? Meaning it's problematic content in the same way as filling your parachute bag with ice cream and screws is problematic.
- hyperpape 2y agoThey specifically denied that null bytes were the issue in an earlier update. https://www.crowdstrike.com/blog/falcon-update-for-windows-hosts-technical-details/ https://www.crowdstrike.com/blog/falcon-update-for-windows-h...
- 1970-01-01 2y agoNull pointers, not a null array
- hyperpape 2y agoI'm not sure what you're saying, but note that a file fundamentally cannot contain a null pointer, a file can just contain various bytes.
- openasocket 2y agoI work on a piece of software that is installed on a very large number of servers we do not own. The crowd strike incident is exactly our nightmare scenario. We are extremely cautious about updates, we roll it out very slowly with tons of metrics and automatic rollbacks. I’ve told my manager to bookmark articles about the crowdstrike incident and share it with anyone who complains about how slow the update process is. The two golden rules are to let host owners control when to update whenever possible, and when it isn’t to deploy very very slowly. If a customer has a CI/CD system, you should make it possible for them to deploy your updates through the same mechanism. So your change gets all the same deployment safety guardrails and automated tests and rollbacks for free. When that isn’t possible, deploy very slowly and monitor. If you start seeing disruptions in metrics (like agents suddenly not checking in because of a reboot loop) rollback or at least pause the deployment.
- taspeotis 2y agoI don’t have much sympathy for CrowdStrike but deploying slowly seems mutually exclusive to protecting against emerging threats. They have to strike a balance.
- yardstick 2y agoIn CrowdStrikes case, they could have rolled out to even 1 million endpoints first and done an automated sanity/wellness check before unleashing the content update on everyone. In the past when I have designed update mechanisms I’ve included basic failsafes such as automated checking for a % failed updates over a sliding 24-hour window and stopping any more if there’s too many failures.
- zavec 2y agoEven a staged rollout over a few hours would have made a huge difference here. "Slow" in the context of a rollout can still be pretty fast.
- bboygravity 2y agoBut it can also still be way too slow in the context of an exploit that is being abused globally.
- dmitrygr 2y ago> How Do We Prevent This From Happening Again? > * Local developer testing Yup... now that all machines are internet connected, telemetry has replaced QA departments. There are actual people in positions of power that think that they do not need QA and can just test on customers. If there is anything right in the world, crowdsuck will be destroyed by lawsuits and every decisionmaker involved will never work as such again.
- mianos 2y agoI see a path to this every day. An actual scenario: Some developer starts working on pre deployment validation of config files. Let's say in a pipeline. Most of the time the config files are OK. Management says: "Why are you spending so long on this project, the sprint plan said one week, we can't approve anything that takes more than a week." Developer: "This is harder than it looks" (heard that before). Management: "Well, if the config file is OK then we won't have a problem in production. Stop working on it". Developer: Stops working on it. Config file with a syntax error slips through, .. The rest is history
- mdriley 2y ago> Based on the testing performed before the initial deployment of the Template Type (on March 05, 2024), trust in the checks performed in the Content Validator, and previous successful IPC Template Instance deployments, these instances were deployed into production. It compiled, so they shipped it to everyone all at once without ever running it themselves. They fell short of "works on my machine".
- meshko 2y agoI so hate it when people fill these postmortems with marketing speak. Don't they know it is counterproductive?
- yashafromrussia 2y agoWell I'm glad they at least released a public postmortem on the incident. To be honest, I feel naive saying this, but having worked at a bunch of startups my whole life, I expected companies like CrowdStrike to do better than not testing it on their own machines before deploying an update without the ability to roll it back.
- notepad0x90 2y agoOne lesson I've learned from this fiasco is to examine my own self when it comes to these situations. I am so befuddled by all the wild opinions, speculations and conclusions as well as observations of the PIR here. You can never have enough humility.
- thayne 2y agoSo this event is probably close to a worst case scenario for an untested sensor update. But have they never had issues with such untested updates before, like an update resulting in false positives on legitimate software? Because if they did, that should have been a clue that these types if updates should be tested too.
- novia 2y agoCrowdstrike issues false positives allll the time. They'll fix them and then they'll come back in a future update. One such false positive is an empty file. Crowdstike hates empty files.
- aenis 2y ago"Based on the testing performed before the initial deployment of the Template Type (on March 05, 2024), trust in the checks performed in the Content Validator, and previous successful IPC Template Instance deployments, these instances were deployed into production." So they did not test this update at all, even locally. Its going to be interesting how this plays out in courts. The contract they have with us limits their liability significantly, but this - surely - is gross negligence.
- onetokeoverthe 2y ago[dead]
- throwaway2037 2y agoAs I understand, it is incredibly difficult to prove "gross negligence". It is better to pressure them to settle in a giant class action lawsuit. I am curious what the total amount of settlements / fines will be in the end. I guess ~2B USD.
- aenis 2y agoSame here. Our losses were quite significant - between lost productivity, inability to provide services, inability of our clients to actually use contracted services, and having to fix their mess - its very easily in the millions. And then there will be the costs of litigation. It was crazy in the IT department over the weekend, but not much less crazy in our legal teams, who were being bombarded with pitches from law firms offering help in recovery. It will be a fun space to watch, and this 'we haven't tested because we, like, did that before and nothing bad happened' statement in the initial report will be quoted in many lawsuits.
- throwaway2037 2y agoTo be clear: I do not expect the settlement to bankrupt them, but I do expect it to be painful. And, when you say "easily in the millions" -- good luck to demonstrate that in a class action lawsuit, and have the judge believe you. It is much harder than people think. You will be lucky to recoup 10% of those expenses after a settlement. Also, your company may also have cyber-security insurance. (Yes, the insurance companies will join the class action lawsuit, but you cannot get blood from a stone. There will be limits about the settlement size.)
- amluto 2y agoCrowdStrike is more than big enough to have a real 2000’s-style QA team. There should be actual people with actual computers whose job is to break the software and write bug reports. Nothing is deployed without QA sign off, and no one is permitted to apply pressure to QA to sign off on anything. CI/CD is simply not sufficient for a product that can fail in a non-revertable way. A good QA team could turn around a rapid response update with more than enough testing to catch screwups like this and even some rather more subtle ones in an hour or two.
- SirMaster 2y agoI feel like for a system that is this widely used and installed in such a critical position that upon a BSOD crash due to a faulting kernel module like this, the system should be able to automatically roll back to try the previous version on subsequent boot(s).
- taylerjane 2y ago[dead]
- Martin_kRussell 2y ago[dead]
- whlh 2y ago[dead]