6 ms·
If it's true that a bad patch was the reason for this I assume someone, or multiple people, will have a really bad day today. Makes me wonder what kind of testi
by ThePhysicist 2y ago
If it's true that a bad patch was the reason for this I assume someone, or multiple people, will have a really bad day today. Makes me wonder what kind of testing they have in place for patches like this, normally I wouldn't expect something to go out immediately to all clients but rather a gradual rollout. But who knows, Microsoft keeps their master keys on a USB stick while selling cloud HSM so maybe Crowdstrike just yolos their critical software updates as well while selling security software to the world.
- nilsb 2y agoWho needs testing when apologizing to your customers is cheaper?
- falcor84 2y agoI would assume that its enterprise customers have an uptime SLA as part of their contract, and that breaching it isn't very cheap for Crowdstrike.
- perbu 2y agoSoftware doesn't have uptime guarantees. They might have time-to-fix on critical issues, though. I assume this is gross negligence, which would leave them open to claims made through courts, though.
- jsiepkes 2y agoI highly doubt their SLA says something about compensating for damages. At most you won't have to pay for the time they were down. And even more ironically; A botched update doesn't mean they are down. It means you are down. So I don't even think their SLA applies to this.
- InsideOutSanta 2y agoYeah, they'll pay with "credits" for the downtime, if what is currently happening even technically qualifies as downtime.
- agrajag 2y agoReputational damage from this is going to be catastrophic. Even if that’s the limit of their liability it’s hard not to see customers leaving en masse.
- dandanua 2y agoThe company will perish, there is no doubt in that.
- icelancer 2y agoExtremely unlikely. This isn't the first blowup Crowdstrike has had; though it's the worst (IIRC), Crowdstrike is "too big to fail" with tons of enterprise customers who have insane switching costs, even after this nonsense. Unfortunately for all of us, Crowdstrike will be around for awhile.
- zik 2y agoBusinesses would be crazy to continue with Crowdstrike after this. It's going to cause billions in losses to a huge number of companies. If I was a risk assessment officer at a large company I'd be speed dialling every alternative right now.
- ajscanlan 2y agoit would be crazy not to at least investigate migration paths away from Crowdstrike, or better redundancies for yourself
- hello_moto 2y agoCybersecurity industry has regular and annual security testing/competitions done by various Organizations that simulates tons of attacks. Vendors are tested against these cases and graded with their effectiveness. I heard Crowdstrike is "best-in-market" for good reasons as others who have more deep knowledge of the industry have shared in this thread.
- dclowd9901 2y agoAnd when it’s more costly for customers to walk back the mistake of adopting your service. Yeah, I get the impression a lot of SaaS companies operate on this model these days. We just signed with a relatively unknown CI platform, because they were available for support during our evaluation. I wonder how available they’ll be when we have a contract in place…
- helsinkiandrew 2y agoAs at 4am NY time CRWD has lost $10Bn (~13%) in marketcap. Of course they've tested, but just not enough for this issue (as is often the case). This is probably several seemingly non consequential issues coming together. I'm not sure why though, when the system is this important that even successfully tested updates aren't rolled out piecemeal though (or perhaps it has and we're only seeing the result of partial failures around the world)
- tehlike 2y agoTesting is never enough. In fact, it won't catch 99% of issues by the virtue of them often testing happy paths only, or that they test what humans can think of, and by no means they are exhaustive. A robust canarying mechanism is the only way you can limit the blast radius. Set up A/B testing infra at the binary level so you can ship updates selectively and compare their metrics. Been doing this for more than 10 years now, it's the ONLY way. Testing is not.
- wwtrv 2y agoDepends on what you mean by enough. It should be more than enough to catch issues like this one specifically. If they can't even manage that they'll fail at your approach as well.
- tehlike 2y agoCanary offers more bang for the buck, and is much easier to set up. So I kind of disagree.
- wwtrv 2y ago> Canary offers more bang for the buck I'm not sure that justifies potentially bricking the devices of hundreds(?) of your clients by shipping untested updates to them. Of course it depends... and would require deeper financial analysis.
- 2y ago
- kjkjadksj 2y agoExactly. They knocked half the world offline probably killed thousands in ERs and the stock is only down to about June lows.
- krspnda 2y agohah that tweet was one heck of an apology. "we deployed a fix to the issue, speak with your customer rep"
- hello_moto 2y agoUnfortunately cybersecurity still revolves around obscurity.
- jachee 2y agoHere's hoping they start from the top. They won't, but hope springs eternal.
- cen4 2y agoDoesn't matter what testing exists. More scale. More complexity. More Bugs. Its like building a gigantic factory farm. And then realizing that environment itself is the birthing chamber and breeding ground of superbugs with the capacity to wipe out everything. I used to work at a global response center for big tech once upon a time. We would get hundreds of issues, we couldn't replicate cause we literally have to set up our own govt or airline or bank or telco to test certain things. So I used to joke with the corporate robots to just hurry up and take over govts, airlines, banks and telcos already, cause thats the only path to better control.
- jonathanstrange 2y agoTesting + a careful incremental rollout in stages is the solution. Don't patch all systems world-wide at once, start with a few, add a few more, etc. Choose them randomly.
- roemerb 2y ago> Its like building a gigantic factory farm. And then realizing that environment itself is the birthing chamber and breeding ground of superbugs with the capacity to wipe out everything. Factorio player detected
- tempaway4575144 2y agoSounds like it was a 'channel file' which I think is akin to an av definition file that caused the problem rather than an actual software change. So they must have had a bug lurking in their kernel driver which was uncovered by a particular channel file. Still, seems like someone skipped some testing. https://x.com/George_Kurtz/status/1814235001745027317 https://x.com/George_Kurtz/status/1814235001745027317 https://x.com/brody_n77/status/1814185935476863321 https://x.com/brody_n77/status/1814185935476863321
- JonChesterfield 2y agoThe parser crashing the system on a malformed input file strongly suggests their software stack in general is trash
- Sohcahtoa82 2y agoSounds like something a fuzzer likely would have found pretty quickly.
- camdenreslink 2y agoHow about a try-catch block? The software reading the definition file should be minimally resilient against malformed input. That's like programming 101.
- caput770 2y agoA badpage fault in a kernel driver doesn't exactly recover from exceptions like that