5 ms·
This event is predicted in Sydney Dekker’s book “Drift into Failure”, which basically postulates that in order to prevent local failure we setup failure prevent
by tbatchelli 2y ago
This event is predicted in Sydney Dekker’s book “Drift into Failure”, which basically postulates that in order to prevent local failure we setup failure prevention systems that increase the complexity beyond our ability to handle, and introduce systemic failures that are global. It’s a sobering book to read if you ever thought we could make systems fault tolerant.
- COGlory 2y agoWe need more local expertise is really the only answer. Any organization that just outsources everything is prone to this. Not that organizations that don't outsource aren't prone to other things, but at least their failures will be asynchronous.
- bjelkeman-again 2y agoFunny thing is that for decades there were predictions about how there was a need for millions of more IT workers. It was assumed one needed local knowledge in companies. Instead what we got was more and more outsourced systems and centralized services. This today is one of the many downsides.
- downrightmike 2y agoTwo weeks ago it was just about all car dealers
- smsm42 2y agoThe problem here would be that there's not enough people who can provide the level of protection a third-party vendor claims to provide, and a person (or persons) with comparable level of expertise would be much more expensive likely. So companies who do their own IT would be routinely outcompeted by ones that outsource, only for the latter to get into trouble when the black swan swoops in. The problem is all other kinds of companies are mostly extinct by then unless their investors had some super-human foresight and discipline to invest for years into something that year after year looks like losing money.
- brightlancer 2y ago> The problem here would be that there's not enough people who can provide the level of protection a third-party vendor claims to provide, and a person (or persons) with comparable level of expertise would be much more expensive likely. Is that because of economies of scale or because the vendor is just cutting costs while hiding their negligence? I don't understand how a single vendor was able to deploy an update to all of these systems virtually simultaneously, and _that_ wasn't identified as a risk. This smells of mindless box checking rather than sincere risk assessment and security auditing.
- downrightmike 2y agothe vendor is just cutting costs while hiding their negligence? That's how it works.
- smsm42 2y agoKinda both I think, with an addition of principal agent problem. If you found a formula that provides the client with an acceptable CYA picture it is very scalable. And the model of "IT person knowledgeable in both security, modern threats and company's business" is not very scalable. The former, as we now know, is prone to catastrophic failures, but those are rare enough for a particular decision-maker to not be bothered by it.
- yeauldfellows 2y agoDepressing thought that this phenomena is some kind of Nash equilibrium. That in the space of competition between firms, the equilibrium is for companies to outsource IT labor, saving on IT costs and passing that cost savings onto whatever service they are providing. -> Firms that outsource, out-compete their competition + expose their services to black swan catastrophic risk. Is regulation that only way out of this, from a game theory perspective?
- slow_typist 2y agoDepressing, but a good way to think about it. The whole market in which crowdstrike can exist is a result of regulation, albeit bad regulation. And since the returns of selling endpoint protection are increasing with volume, the market can, over time, only be an oligopoly or monopoly. It is a screwed market with artificially increased demand. Also the outsourcing is not only about cost and compliance. There is at least a third force. In a situation like this, no CTO who bought crowdstrike products will be blamed. He did what was considered best industry practice (box ticking approach to security). From their perspective it is risk mitigation. In theory, since most of the security incidents (not this one) involve the loss of personal customer data, if end customers would be willing to a pay a premium for proper handling of their data, AND if firms that don’t outsource and instead pay for competent administrators within their hierarchy had a means of signaling that, the equilibrium could be pushed to where you would like it to be. Those are two very questionable ifs. Also how do you recognise a competent administrator (even IT companies have problems with that), and how many are available in your area (you want them to live in the vicinity) even if you are willing to pay them like the most senior devs? If you want to regulate the problem away, a lot of influencing factors have to be considered.
- mym1990 2y agoMany systems are fault tolerant, and many systems can be made fault tolerant. But once you drift into a level of complexity spawned by many levels of dependencies, it definitely becomes more difficult for system A to understand the threats from system B and so on.
- tbatchelli 2y agoDo you know of any fault tolerant system? Asking because in all the cases I know, when we make a system "fault tolerant" we increase the complexity and we introduce new systemic failure modes related to our fault-tolerant-making-system, making them effectively non fault tolerant. In all the cases I know, we traded frequent and localized failure for infrequent but globalized catastrophic failures. Like in this case.
- lucianbr 2y agoYou can make a system tolerant to certain faults. Other faults are left "untolerated". A system that can tolerate anything, so have perfect availability, seems clearly impossible. So yeah, totally right, it's always a tradeoff. That's reasonable, as long as you trade smart. I wonder if the people deciding to install Crowdstrike are aware of this. If they traded intentionally, and this is something they accepted, I guess it's fine. If not... I further wonder if they will change anything in the aftermath.
- mym1990 2y agoThere will be lawsuits, there will be negotiations for better contracts, and likely there will be processes put in place to make it look like something was done at a deeper level. And yet this will happen again next year or the year after, at another company. I would be surprised if there was a risk assessment for the software that is supposed to be the answer to the risk assessment in the first place. Will be interesting to see what happens once the dust settles.
- slt2021 2y ago- This is system has a single point of failure, it is not fault tolerant. Lets introduce these three things to make it fault-tolerant - Now you have three single points of failure...
- notNNT 2y agoAlso a major point in the Black Swan. In the Black Swan, Taleb describes that it is better for banks to fail more often than for them to be protected from any adversity. Eventually they will become "too big to fail". If something is too big to fail, you are fragile to a catastrophic failure.
- _yb2s 2y agoI was wondering when someone would bring up Taleb RE: this incident. I know you aren't saying it is, but I think Taleb would argue that this incident, as he did with the coronavirus pandemic for example, isn't even a Black Swan event. It was extremely easy to predict, and you had a large number of experts warning people about it for years but being ignored. A Black Swan is unpredictable and unexpected, not something totally predictable that you decided not to prepare for anyways.
- realNNT 2y agoThat is interesting, where does he talk about this? I'm curious to hear his reasoning. What I remember from the Black Swan is that Black Swan events are (1) rare, (2) have a non-linear/massive impact, (3) and easy to predict retrospectively. That is, a lot of people will say "of course that happened" after the fact but were never too concerned about it beforehand. Apart from a few doomsdayers I am not aware of anybody was warning us about a crowd strike type of event. I do not know much about public health but it was my understanding that there were playbooks for an epidemic. Even if we had a proper playbook (and we likely do), the failure is so distributed that one would need a lot of books and a lot of incident commanders to fix the problem. We are dead in the water.
- yeauldfellows 2y agoI think Grey Rhino is the term to use. Risks that we can see and acknowledge yet do nothing about.
- nurbl 2y ago"Antifragile" is even more focused around this.
- jessriedel 2y agoIt's also in line with arguments made by Ted Kaczynski (the Unabomber) > Why must everything collapse? Because, [Kaczynski] says, natural-selection-like competition only works when competing entities have scales of transport and talk that are much less than the scale of the entire system within which they compete. That is, things can work fine when bacteria who each move and talk across only meters compete across an entire planet. The failure of one bacteria doesn’t then threaten the planet. But when competing systems become complex and coupled on global scales, then there are always only a few such systems that matter, and breakdowns often have global scopes. https://www.overcomingbias.com/p/kaczynskis-collapse-theoryhtml https://www.overcomingbias.com/p/kaczynskis-collapse-theoryh... https://en.wikipedia.org/wiki/Anti-Tech_Revolution https://en.wikipedia.org/wiki/Anti-Tech_Revolution
- localfirst 2y agocrazy how much he was right. if he hadn't gone down the path of violence out of self-loathing and anger he might have lived to see a huge audience and following.
- washadjeffmad 2y agoI suppose we wouldn't know whether an audience for those ideas exists today because they would be blacklisted, deplatformed, or deamplified by consolidated authorities. There was a quote last year during the "Twitter files" hearing, something like, "it is axiomatic that the government cannot do indirectly what it is prohibited from doing directly". Perhaps ironically, I had a difficult time using Google to find the exact wording of the quote or its source. The only verbatim result was from a NYPost article about the hearing.
- smsm42 2y ago> it is axiomatic that the government cannot do indirectly what it is prohibited from doing directly Turns out SCOTUS decided it isn't, and the government is free to do exactly that as long as they are using the services of an intermediary.
- 2y ago
- ricardo81 2y agoI haven't read it, but I'd take a leap to presume it's somewhere between the people that say "C is unsafe" and "some other language takes care of all of things". Basically delegation.
- smsm42 2y agoJust yesterday listened to a lecture by Moshe Vardi which covers adjacent topics: https://simons.berkeley.edu/events/lessons-texas-covid-19-737-max-efficiency-vs-resilience-richard-m-karp-distinguished-lecture https://simons.berkeley.edu/events/lessons-texas-covid-19-73...
- joe_the_user 2y agoI think it was "predicted" by Sunburst, the Solarwinds hack. I don't think centrally distributed anti-virus software is the only way to maintain reliability. Instead, I'd say companies to centralize anything like administration since it's cost effective and because they actually aren't concerned about global outage like this. JM Keynes said "A ‘sound’ banker, alas! is not one who foresees danger and avoids it, but one who, when he is ruined, is ruined in a conventional and orthodox way along with his fellows, so that no one can really blame him." and the same goes for corporate IT.
- akira2501 2y ago> we setup failure prevention systems You can't prevent failure. You can only mitigate the impact. Biology has pretty good answers as to how to achieve this without having to increase complexity as a result, in fact, it often shows that simpler systems increase resilliency. Something we used to understand until OS vendors became publicly traded companies and "important to national security" somehow.
- jjav 2y ago> if you ever thought we could make systems fault tolerant The only possible way to fault tolerancy is simplicity and then more simplicity. Things like crowsdtrike have the opposite approach. Add a lot of fragile complexity attempting to catch problems, but introducing more attack surfaces than they can remove. This will never succeed.
- ryukoposting 2y agoAs an architect of secure, real-time systems, the hardest lesson I had to learn is there's no such thing as a secure, real-time system in the absolute sense. Don't tell my boss.
- deleted 2y ago[deleted]
- oneepic 2y agoI think a surprising amount of people already share this view, even if they don't go into extensive treatment with references like Dekker presumably does (I haven't read it). I suspect most people in power just don't subscribe to that. which is precisely why it's systemic to see the engineer shouting "no!" when John CEO says "we're doing it anyway." I'm not sure this is something you can just teach, because the audience definitely has reservations about adopting it.