4 ms·
You’re not wrong about the end result, but the breakdown of systems this complex goes deeper than placing the blame on some CrowdStrike employee. Whoever thoug
by RCitronsBroker 2y ago
You’re not wrong about the end result, but the breakdown of systems this complex goes deeper than placing the blame on some CrowdStrike employee.
Whoever thought up
the great idea to allow auto-update-able kernel modules for something as mission critical as emergency response or healthcare deserves just as much blame.
I’ve worked in healthcare for my whole career, this is madness. Not that their process is without flaw, but can we remind ourselves of how stringently we assess medical devices? I cannot imagine it’s controversial to say that emergency response equipment is every bit as critical as a insulin pump. If they fail, someone dies.
- hsbauauvhabzb 2y agoBut at the same time, auditing every update to an assurance level beyond ‘it didn’t bsod in test’ is incredibly hard. I don’t disagree with anything you’ve said, but I’d be very interested in solving the problem of actually auditing constant updates from vendors.
- withinboredom 2y agoEven a rudimentary “delay autoupdate by two weeks” would have saved lives here. Let everyone else update first.
- exe34 2y agostaged releases. don't cripple all your systems in one go. hot backups that you only update after the main system isn't dead from an update.
- com 2y agoAutomated CI/CD - many of us already do this hundreds of times a day. If you’re an emergency call centre, join a consortium of similar orgs and standardise tech and do it properly. Defer updates. Most things can wait 8-12 hours. Even more can wait 3 weeks (did this for all but security-critical npm package updates in one place). Demand legal changes to ensure fair liability for failure to undertake basic measures by service providers for paid software and services. Demand proper liability for C-suites not ensuring that actual risk management is in place instead of stupid box-ticking. Design better software. Seriously, the kinds of half-baked stuff that costs so much is incredible. It doesn’t take longer, and it doesn’t cost more to do things right, the only change is that management needs to be engaged with outcomes and have skin in the game. Execs should run the risk of going to jail for egregious failures.
- soneil 2y ago> Whoever thought up the great idea to allow auto-update-able kernel modules What's made this whole thing so "interesting" is that the whole point of these "channel files" was to decouple the risk from updating the kernel driver. Accepted best practice for this product has been to stagger rollout of the kernel driver, so a pilot group gets the current release, the herd get n-1, and sensitive machines get n-2. The product provides for this, and most sites either use it, or admit they should. So when your pilot group start bluescreening with "DRIVER OVERRAN STACK BUFFER" (actual example from last year), it's caught (by the customer, still) and triaged before it reaches n-1, let alone n-2 & front page of The Times. But the whole 'sell' of the product is that they get 0-day definitions. So endpoints running the relatively trusted n-2 release still get the same protection against active threats. n-2 have a stable driver running today's "channel data". I'm not clear if Friday's "channel file" is the issue in itself, or whether it triggered a less-explored code path in the kernel driver - but the result is the same. The best practice of staggering the kernel driver releases, didn't save us from a logic bomb in the "channel file". I just think the distinction is interesting because following accepted best practices, vendor recommendations, and conservative deployment recommendations did not protect from this. It's not the customers that were yolo'ing this.
- RCitronsBroker 2y agothis was a very valuable insight, I’m a med student at the moment, my interest in networking and tech in general is a tad more shallow, but i appreciate your perspective nonetheless! Additionally, would you mind sharing your thoughts on the following observations? Afaik, similarly to medical devices, we recognize the criticality of software for applications such as ATC or microcontroller-based railway switchyards; for obvious reasons ofc. Alright, but ensuring the availability of barebones emergency response or Hospital IT shouldn’t be far off in terms of criticality, no? Yet, ATC, avionics, rail DMIs/infrastructure and similar go through the effort of building ultra-available, purpose built systems that are very different from Windows instances running CS kernel tools, even thoughtful ones. In contrast, apparently said healthcare/emergency related applications seemingly are okay with relying on mission critical windows boxes. I hope that info is factual, otherwise mea culpa. I don’t mind healthcare using less elaborate tech for non-critical purposes, the equivalent of the service responsible for providing train delay updates, stuff far away from operating signals type ops. But if its mission critical or able to impede critical services, that’s really worrying to me.