6 ms·
The 'experts' also made similar criticisms with the Fastly outage in 2021 and did anything obvious change as a result? In a week's time no national newspapers w
by sunrunner 1y ago
The 'experts' also made similar criticisms with the Fastly outage in 2021 and did anything obvious change as a result? In a week's time no national newspapers will be talking about this.
Meanwhile, everyone that spends actual time in these areas:
- Knows that running an operation at AWS scale is difficult and any armchair critism from 'experts' is exactly that. Actions speak louder than words.
- Understands that the cost of actually accounting for this kind of scenarios is incredibly high for the benefit in most cases
- Knows that genuinely 'critical' services (i.e. health) should be designed to account for this, and every other 'serious' issue such as 'I can't log in to Fortnite' just shows what the price and effort of actually making that work is versus how much it costs affected companies when it happens
- Knows how much time national newspapers spend actually talking about the importance of multi-region/multi-cloud redundancy, that is, it's zero until the one day where it happens and then it's old news
- Is just curious as to just what exactly happened from a technical perspective
This isn't to say that good blameless post-mortem shouldn't happen to figure out process and technical issues, but the armchair criticism with no actual followup? All noise, no signal.
- free_bip 1y agoBecause the experts have no say in policy. The only people who have a say are the people bribing (sorry I mean "lobbying") Congress. And even they have very little say because Congress is currently on a hot streak of doing absolutely nothing.
- gnerd00 1y agomaybe your VC overlords need a reality check?
- BrenBarn 1y agoI think all of that is mostly irrelevant. You don't need to pay a huge cost to avoid the small benefit, you don't need every service to be resilient to this, or any of that. You just need multiple different providers so that not everyone gets screwed at once.
- bamboozled 1y agoYes but we live in a highly anti-competitive monopolized world now. With more to come under the new admin.
- jdminhbg 1y agoIt’s hard to think of anything less monopolized than cloud hosting. There are hundreds of providers.
- estimator7292 1y agoFor any business that matters, your choices are amazon, google, Microsoft, and that's about it. I couldn't even name another provider except maybe Hetzner
- bamboozled 1y agoThe three you mentioned have over 60% market share which is why this article exists at all. Knowing what I know about cloud ifnra, anyone who is actually anyone is hosting on the big three. So it's not just a market share, it's market share + impact / importance. You could also argue that YT is on GCP (to some level) and that would probably bump that number up much higher. The vast majority of people hosting things on the internet are on these providers. But you get downvoted for pointing that out now.
- bamboozled 1y agoYeah right, and how many of them have any substantial customer base compared to AWS and Azure?
- hdgvhicv 1y agoThere’s two or three gartner approved ways of doing things for fortune 500 ctos, and f500 wannabes. It’s not a monopoly but it’s close.
- sunrunner 1y ago
- deleted 1y ago[deleted]
- alecco 1y ago> - Knows that running an operation at AWS scale is difficult and any armchair critism from 'experts' is exactly that. Actions speak louder than words. NO. From their own reports, clearly AWS is too centralized and dependent on a specific region (us-east-1) and a specific service (DynamoDB). This has been observed for well over 10 years. Why do they stay in this centralized architecture? Cloud services need much higher standards than the average corporation. Just look how they took down 2000+ services for many hours. [1] https://health.aws.amazon.com/health/status https://health.aws.amazon.com/health/status
- sunrunner 1y agoI don't deny that an incident of this scope should prompt a serious technical and process review (and as you describe it, it sounds like this is long overdue), however how often does this kind of thing not affect 2000+ services? Companies should be tracking the time they don't have issues as much as the time they do in order to actually understand if they'd be better off elsewhere. And to be clear, I'm not at all arguing for the monopolisation of cloud providers, only stating that it's easy to point from far away and say 'This is bad' while simultaneously not doing anything to understand the cost and make that change that you say is important, because it's actually costly (in many dimensions) to do.
- inopinatus 1y agoEven wearing my ex-AWS hat and understanding to some degree the internal complexity of these services, I too am boggled that foundational stuff is still out of Virginia and not a separately operated global region for the subset of control-plane dependencies that can’t be refactored into tolerating eventual consistency (such as parts of IAM). We always used to talk a lot about minimising blast radius and there’s been enough time, and enough scale, to fix it. Nevertheless the Guardian’s choice to label self-promoting policy wonks as “experts” is a cringe-inducing reminder that journalists don’t know anything about anything.
- imgabe 1y agoThe "experts" in this case are > Dr Corinne Cath-Speth, the head of digital at human rights organisation Article 19 Dr. Cath-Speth has a PhD in cultural anthropology > Cori Crider, the executive director of the Future of Technology Institute A lawyer > Madeline Carr, professor of global politics and cybersecurity at University College London A professor. Her bio doesn't say what her degree is in, but she mostly seems to publish in political science and international relations So, not a single technical expert. Not anyone who has ever run a hosting service before or even worked for one. Just people who write papers and sit around waiting for journalists to call them for quotes.
- wagwang 1y agoOpinions are valid but also worthless. Just give me a funny tweet to digest the situation.
- kopecs 1y agoDo you not think it a bit too hyperbolic to throw scare quotes around experts and imply the only people who can have opinions on systemic risk are software engineers? I don't think it is unreasonable for people who haven't run or worked for a hosting service to have opinions on the policy aspect or economic impact of hyperscalers.
- imgabe 1y agoAnyone can have an opinion, I never said or implied otherwise. Having an opinion does not make one an expert, hence the scare quotes. The headline is misleading because when there is news about experts saying something about technology, one would naturally think that they are at least somewhat technical experts. Instead the "expert" is the director of the "Big Tech is Bad Institute" who says that "Big Tech is Bad". And their qualification of being an expert is solely that they are director of the "Big Tech is Bad Institute".
- ghaff 1y agoAnd one would hope that the stats being quoted about desktop share were from someone who has been at that research firm in the last 20 years or so. I'm not sure how active he is at all at this point. I have a feeling someone looking for some stats found something old that may or may not have actually had a date on it. (If I'm wrong mea culpa but I'm pretty sure.)
- zenoprax 1y agoI think your third point is what I've had to attune to when criticizing cloud dependence. I think if your entire source of revenue is dependent on AWS then you should be prepared for 16+ hours of downtime per year. Individuals notice it more when something is down for hours but with good observability I am guessing the business notices it more when performance drags for the other 8742 hours of the year. Bursts of downtime per day can still be attributed to the device, wifi, ISP, or some other intermediary's DNS/BGP. If your margins are so tight that 16 hours of downtime will bankrupt you then I think either: a) I have no idea how to run a business; or b) you have no idea how to run a business. I'm also biased because I love highly fault-tolerant, geo-redundant, durable systems much more than "good enough for this KPI".
- sunrunner 1y ago> but with good observability I am guessing the business notices it more when performance drags for the other 8742 hours of the year This is really good point that aligns with my experience. Today's event was LOUD and (compared to other incidents) long, but perhaps not really that long compared to the situation you describe that for most businesses is going to be more pernicious. Business intelligence and analytics-type folks at $DAYJOB are _very_ watchful for the year-on-year deviations and even periods where the prediction lines didn't match up for even just a few hours.
- hippo77 1y agoThese are Guardian 'experts' so can be safely ignored.
- sysguest 1y ago> - Knows that genuinely 'critical' services (i.e. health) should be designed to account for this yeah but aws advertises as "trust me bro I won't go down for 99.99999%" I've seen a lot of gov proposals using aws to 'get away with downtime management'
- A4ET8a8uTh0_v2 1y agoUm.. you don't need to be an expert in security, comp.science or economics to know that putting all eggs in one basket may not be a great idea as introduces one giant systemic target. If anything, regular people here are uniquely qualified to say something along the lines of: Oi, this is ridiculous. Maybe more things should be ran locally.. FWIW, it was instructive to me as to which companies were not able to function today.
- benterix 1y ago> Knows how much time national newspapers spend actually talking about the importance of multi-region/multi-cloud redundancy For the record, multi-region redundancy is moot, and I can't stress it enough. It is not the first time that on the surface it looks like a single region but in fact services in multiple regions are affected. And multi-cloud hot standby can be terribly expensive, unless your infra is very simple. And it's not easy to get it right either until you planned for it from day one.