6 ms·
But how come they didn't catch it in the testing deployments? what was the difference that caused it to happen when they deployed to the outside world. I find i
by golemiprague 2y ago
But how come they didn't catch it in the testing deployments? what was the difference that caused it to happen when they deployed to the outside world. I find it hard to believe that they didn't test it before deployment. I also think companies should all have a testing environment before deploying 3rd party components. I mean, we all install some packages during development that fails or cause some problems but nobody think it is a good idea to do it directly in their production environment before testing, so how is this different?
- someonehere 2y agoThat’s what a lot of us are wondering. There’s a lot of outside thinking of the box right now about this in certain circles.
- IAmGraydon 2y agoThere’s no point in leaving vague allusions. Can you expand on this?
- kbar13 2y agosecurity industry's favorite language is nothingspeak
- jmb99 2y ago> I find it hard to believe that they didn't test it before deployment. I’m not sure why you find that hard to believe - based on the (admittedly fairly limited) evidence we have right now, it’s highly unlikely that this deployment was tested much, if at all. It seems much more likely to me that they were playing fast and loose with definition updates to meet some arbitrary SLAs[1] on zero-day prevention, and it finally caught up with them. Much more likely than somehow every single real-world pc running their software being affected but their test machines somehow all impervious. [1] When my company was considering getting into endpoint security and network anomaly detection, we were required on multiple occasions by multiple potential clients to provide a 4-hour SLA on a wide number of CVE types and severities. That would mean 24/7 on-call security engineers and a sub-4-hour definition creation and deployment. Yes, that 4 hours was for the deployment being available on 100% of the targets. Good luck writing and deploying a high-quality definition for a zero day in 4 hours, let alone running it through a test pipeline, let alone writing new tests to actually cover it. We very quickly noped out of the space, because that was considered “normal” (at least to the potential clients we were discussing). It wouldn’t shock me if CS was working in roughly the same way here.
- drooopy 2y agoThis whole f*up was a failure of management and processes at Crowdstrike. "Intern Steve" pushing faulty code to production on a Friday is only a couple of cm of the tip of an enormous iceberg.
- chronid 2y agoI wrote this in another thread already, but the fuck up was both at crowdstrike (they borked a release) but also and more importantly their customers. Shit happens even with the best testing in the world. You do not deploy anything, ever on your entire production fleet at the same time and you do not buy software that does that. It's madness and we're not talking about small companies with tiny IT departments here.
- perbu 2y agoShit might happen with the best testing, but with decent testing it would not be this serious.
- wazzaps 2y agoApparently CrowdStrike bypassed clients' staging areas with this update. Source: https://x.com/patrickwardle/status/1814367918425079934 https://x.com/patrickwardle/status/1814367918425079934
- owl57 2y ago> you do not buy software that does that Note how the incident disproportionally affected highly regulated industries, where businesses don't have a choice to screw "best practice".
- TeMPOraL 2y agoOnly highlighting that "best practice" of cybersecurity is, charitably, total bullshit; less charitably, a racket. This is apparent if you look at the costs to the day-to-day ability of employees to do work, but maybe it'll be more apparent now that people got killed because of it.
- treflop 2y agoI’ve seen places where failed releases are just “part of normal engineering.” Because no one is perfect, they say.
- slenk 2y agoI really dislike this mentality. Don't even get me started on celebrating when your rocket blows up
- galangalalgol 2y agoIf it is a standard production rocket, I agree. If it is a first of kind or even third of kind launch, celebrating the lessons learned from a failure is a healthy attitude. This production software is not the same thing at all.
- Heliosmaster 2y agospaceX celebrating when their rocket blows up after a certain milestone it's like us devs celebrating when our branch with that new big feature only fails a few tests. Did it pass no? Are you satisfied as first try? Probably
- photonthug 2y agoEven on hn, comments advocating engineering excellence or just quality in general are frequently looked down on, which probably also tells you a lot about the wider world. This is why we can’t have nice things, but maybe we just don’t want them anyway? “Mistakes will be made” is way less true if you actually put the effort in to prevent them, but I am beginning to think this has become code for quiet-quitters to telegraph a “I want to get paid for no effort and sympathize with others who feel the same” sentiment and appear compassionate and grimly realistic all at the same time. yes, billion dollar companies are going to make mistakes, but almost always because of cost cutting, willful ignorance, or negligence. If average people are apologizing for them and excusing that, there has to be some reason that it’s good for them.
- treflop 2y ago
- usrusr 2y agoOne possible explanation could be automated testing deployments for definitions updates that don't run the current version of the definition consumer, and the old one they do run is unaffected.
- itronitron 2y agofor all we know, the deployment was the test
- owl57 2y agoAs the old saying goes, everyone has a test environment, and some also have a separate production one.
- albert_e 2y agoMy guess -- there are two separate pipelines one for code changes and one for data files. Pipeline 1 -- Code updates to their software are treated as material changes that require non-production and canary testing before global roll-out of a new "Version". Pipeline 2 -- Content / channel updates are handled differently -- via a separate pipeline -- because only new malware signatures and the like are distrubuted via this route. The new files are just data files -- they are supposed to be in a standard format and only read, not "executed". This pipeline itself must have been tested originally and found tobe working satisfactorily -- but inside the pipeline there is no "test" stagethat verifies the integrity of the data fine so generated, nor - more importantly - checking if this new data file works without errors when deployed to the latest versions of the software in use. The agent software that reads these daily channel files must have been "thoroughly" tested (as part of pipeline 1) for all conceivable data file sizes and simulated contents before deployment. (any invalid data files should simply be rejected with an error ... "obviously") But the exact scenario here -- possibly caused by a broken pipeline in the second path (pipeline 2) -- created invalid data files with some quirks. And THAT specific scenario was not imagined or tested in the software version dev-test-deploy pipeine (pipeline 1). If this is true -- The lesson obviously is that even for "data" only distributions and roll-outs, however standardized and stable their pipelines may be, testing is still an essential part before large scale roll-outs. It will increase cost and add latency sure, but we have to live with it. (similar to how people pay for "security" software in the first place) Same lesson for enterprise customers as well -- test new distributions on non-production within your IT setup, or have a canary deployment in place before allowing full roll-outs into production fleets.
- sateesh 2y agoSame lesson for enterprise customers as well -- test new distributions on non-production within your IT setup, or have a canary deployment in place before allowing full roll-outs into production fleets. It was mentioned in one of the HN threads, that the update was pushed overriding the settings customer had [1]. What recourse any customer can have in in such a case ? 1. https://news.ycombinator.com/item?id=41003390 https://news.ycombinator.com/item?id=41003390
- 2y ago
- masfuerte 2y agoI find it hard to believe they didn't do any testing. I wonder if they tested the virus signatures against the engine, but didn't check the final release artefact (the .sys file) and the bug was somehow introduced in the packaging step. This would have been poor, but to have released it with no testing would have been the most staggering negligence.