5 ms·
>>So basically we have nothing. Except the biggest IT outage ever. And a postmortem showing their validation checks were insufficient. And a rollout process
by nyc_data_geek1 2y ago
>>So basically we have nothing.
Except the biggest IT outage ever. And a postmortem showing their validation checks were insufficient. And a rollout process that did not stage at all, just rawdogged straight to global prod. And no lab where the new code was actually installed and run prior to global rawdogging.
I'd say there's smoke, and numerous accounts of fire, which this can be taken in the context of.
- nickcocker 2y ago[dead]
- mewpmewp2 2y agoThere definitely was a huge outage, but based on the given information we still can't know for sure how much they invested in testing and quality control. There's always a chance of failure even for the most meticulous companies. Now I'm not defending or excusing the company, but a singular event like this can happen to anyone and nothing is 100%. If thorough investigation revealed poor quality control investment compared to what would be appropriate for a company like this, then we can say for sure.
- hulitu 2y ago[flagged]
- dang 2y agoCould you please stop posting unsubstantive comments and/or flamebait? Posts like this one and https://news.ycombinator.com/item?id=41542151 https://news.ycombinator.com/item?id=41542151 are definitely not what we're trying for on HN. If you wouldn't mind reviewing https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html and taking the intended spirit of the site more to heart, we'd be grateful.
- daedrdev 2y agoTwo things are clear though Nobody ran this update The update was pushed globally to all computers With that alone we know they have failed the simplest of quality control methods for a piece of software as widespread as theirs. This is even excluding that there should have been some kind of error handling to allow the computer to boot if they did push bad code.
- busterarm 2y agoAlso it's the _second_ time that they had done this in a few short months. They had previous bricked linux hosts earlier with a similar type of update. So we also know that they don't learn from their mistakes.
- rblatz 2y agoThe blame for the Linux situation isn’t as clear cut as you make it out to be. Red hat rolled out a breaking change to BPF which was likely a regression. That wasn’t caused directly by a crowdstrike update.
- IcyWindows 2y agoAt least one of the incidents involved Debian machines, so I don't understand how Red Hat's change would be related.
- rblatz 2y agoSorry, that’s correct it was Debian, but Debian did apply a RHEL specific patch to their kernel. That’s the relationship to red hat.
- busterarm 2y agoIt's not about the blame, it's about how you respond to incidents and what mitigation steps you take. Even if they aren't directly responsible, they clearly didn't take proper mitigation steps when they encountered the problem.
- idkwhatimdoin 2y ago> If thorough investigation revealed poor quality control investment compared to what would be appropriate for a company like this, then we can say for sure. We don't really need that thorough of an investigation. They had no staged deploys when servicing millions of machines. That alone is enough to say they're not running the company correctly.
- dartos 2y agoTotally agree. I’d consider staggering a rollout to be the absolute basics of due diligence. Especially when you’re building a critical part of millions of customer machines.
- wlonkly 2y agoI also fall on the side of "stagger the rollout" (or "give customers tools to stagger the rollout"), but at the same time I recognize that a lot of customers would not accept delays on the latest malware data. Before the incident, if you asked a customer if they would like to get updates faster even if it means that there is a remote chance of a problem with them... I bet they'd still want to get updates faster.
- dartos 2y agoThere must be balance
- mewpmewp2 2y agoI would say that canary release is an absolute must 100%. Except I can think of cases where it might still not be enough. So, I just don't feel comfortable judging them out of the box. Does all the evidence seem to point against them? For sure. But I just don't feel comfortable giving that final verdict without knowing for sure. Specifically because this is about fighting against malicious actors, where time can be of essence to deploy some sort of protection against a novel threat. If there's deadlines that you can go over, and nothing bad happens, for sure. Always have canary releases, and perfect QA, monitoring everything thoroughly, but I'm just saying, there can be cases where damage that could be done if you don't act fast enough, is just so much worse. And I don't know that it wasn't the case for them. I just don't know.
- deleted 2y ago[deleted]
- quietbritishjim 2y agoThe sentence you quoted clearly meant, from the context, "clearly we have nothing [to learn from the opinions of these former employees]". Nothing in your comment is really anything to do with that.
- tomrod 2y agoTriangulation versus new signal.
- sundvor 2y ago"Everyone" piles on Tesla all the time; a worthwhile comparison would be how Tesla roll out vehicle updates. Sometimes people are up in arms "where's my next version" (eg when adaptive headlights was introduced), yet Tesla prioritise a safe, slow roll out. Sometimes the updates fail (and get resolved individually), but never on a global scale. (None experienced myself, as a TM3 owner on the "advanced" update preference). I understand the premise of Crowdstrike's model is to have up to date protection everywhere but clearly they didn't think this through enough times, if at all.
- kccqzy 2y agoYou can also say the same thing about Google. Just go look at the release notes on the App Store for the Google Home app. There was a period of more than six months where every single release said "over the next few weeks we're rolling out the totally redesigned Google Home app: new easier to navigate 5-tab layout." When I read the same release notes so often I begin to question whether this redesign is really taking more than six months to roll out. And then I read the Sonos app disaster and I thought that was the other extreme.
- cesarb 2y ago> Just go look at the release notes on the App Store for the Google Home app. [...] When I read the same release notes so often I begin to question whether this redesign is really taking more than six months to roll out. Google is terrible at release notes. Since several years ago, the release notes for the "Google" app on the Android app store always shows the exact same four unchanging entries, loosely translating from Portuguese: "enhanced search page appearance", "new doodles designed for app experience", "offline voice actions (play music, enable Wi-Fi, enable flashlight) - available only in the USA", "web pages opened directly within the app". I heavily doubt it's taking these many years to roll out these changes; they probably simply don't care anymore, and never update these app store release notes.
- hello_moto 2y ago> And no lab where the new code was actually installed and run prior to global rawdogging. I thought the new code was actually installed, the running part depends on the script input...?