17 ms·
AWS outage shows internet users 'at mercy' of too few providers, experts say
- saltysalt 1y agoAWS is this generation's mainframe. /joking
- labrador 1y agoThis new post is interesting: https://news.ycombinator.com/item?id=45646777 https://news.ycombinator.com/item?id=45646777 "October 17, 2025, was my last day at Amazon Web Services... CloudFront is a CDN, a content delivery network, or, simply put, a large distributed cache for your cat photos. And a very successful one. Something like 30% of all internet traffic goes through CloudFront in one way or another. Pretty cool, huh? In practice, this means that with any change, you have a chance of crashing 30% of the internet."
- dynamite-ready 1y agoThe whole industry walked straight into the cloud service lock-in trap. How would we begin to wind back? I also think Docker is as much to blame as the bigger cloud vendors.
- Jcowell 1y agoWhy is docker to blame?
- dynamite-ready 1y agoIt's subjective I guess, but I feel as though containerisation has greatly supported the large Cloud vendor's desire to subvert the more common model of computing... Like, before, your server was a computer, much like your desktop machine, and you programmed it much like your desktop machine. But now, people are quite happy to put their app in a Docker container and outsource all design and architecture decisions pertaining to data storage and performance. And with that, the likes of ECS, Dynamo, RedShift, etc, are a somewhat reasonable answer to that. It's much easier to offer a distinct proposition around that state of affairs, than say a market that was solely based on EC2-esque VMs. What I did not like, but absolutely expected, was this lurch towards near enough standardising one specific vendor's model. We're in quite a strange place atm, where AWS specific knowledge might actually have a slightly higher value than traditional DevOps skills for many organisations. Felt like this all happened both at the speed of light, and in slow motion, at the same time.
- pythonaut_16 1y agoI don't see how Docker makes that worse. Before Docker you had things like Heroku and Amazon Elastic Beanstalk with a much greater degree of lock in than Docker. ECS and its analogues on the other cloud providers have very little lock in. You should be able to deploy your container to any provider or your own VM. I don't see what Dynamo and data storage have to do with that. If we were all on EC2s with no other services you'd still have to figure out how to move your data somewhere else? Like I truly don't understand your argument here.
- throwaway894345 1y agoContainers have nothing to do with storage. They are completely orthogonal to storage (you can use Dynamo or RedShift from EC2), and many people run Docker directly on VMs. Plenty of us still spend lots of time thinking about storage and state even with containers. Containers allow me to outsource host management. I gladly spend far less time troubleshooting cloud-init, SSH, process managers, and logging/metrics agents.
- spjt 1y agoI don't think it wants to. Ask any on-call engineer or support tech how they felt when, after having their phone blow up at 1am because everything is falling apart, they found out that this was an AWS-wide outage.
- bix6 1y agoCan someone educate me on the solution to this? I assume most organizations, both small and large, just host on whatever provider they know or that costs them the least. If you have budget maybe you deploy to multiple providers for redundancy? But that increases cost and complexity. Who’s going to bother with colo given the cost / complexity? Who’s going to run a server from their office given ISP restrictions and downtime fears? What is the realistic antidote here?
- gytisgreitai 1y agoWhat cost? Complexity - yes, to some extent.
- 98codes 1y agoCompanies can architect their backends to be able to fail back to another region in case of outage, and either don't test it or don't bother to have it in place because they can just blame Amazon, and don't otherwise have an SLA for their service. To fix it, test your failback procedures. For everything else, there's nothing to fix, it's working by design.
- maccard 1y ago> Companies can architect their backends to be able to fail back to another region in case of outage, and either don't test it or don't bother to have it in place because they can just blame Amazon, and don't otherwise have an SLA for their service. My CI was down for 2 hours this morning, despite not even being on AWS. We have a set of credentials on that host that we call assumeRole with and push to an S3 bucket, which has a lambda that duplicates to buckets in other regions. All our IAM calls were failing due to this outage, and we have 0 items deployed in us-east-1 (we're european)
- cyberax 1y agoYou likely used a us-east-1 IAM endpoint instead of a regionalized one ( https://aws.amazon.com/blogs/security/how-to-use-regional-aws-sts-endpoints/ https://aws.amazon.com/blogs/security/how-to-use-regional-aw... ). We've been using it, and we're not experiencing any issues whatsoever in us-east-2. One thing that AWS should do is provide an easier way to detect these hidden dependencies. You can do that with CloudTrail if you know how to do it (filter operations by region and check that none are in us-east-1), but a more explicit service would be nice.
- neom 1y agoBeen a while since I worked in cloud but at least when I got out of it, the primitives where all shoring up to be generally very similar. Did multi cloud redundancy end up being too expensive? Tech didn't line up enough? No good business case? The elastic cloud story that never was? https://www.slideshare.net/slideshow/pets-vs-cattle-the-elastic-cloud-story/31735707 https://www.slideshare.net/slideshow/pets-vs-cattle-the-elas... What happened?
- LaurensBER 1y agoThe (cognitive) overhead of managing and deploying to multiple clouds usually isn't worth it for most teams. Hiring experts and maintaining knowledge about the ins and outs of two (or more) clouds is less feasible for small, fast moving teams. Simplicity is linked to uptime and having a single cloud solution is a simpeler solution. For large companies, its mostly cost savings. Easier to negotiate a good discount at N million versus N/2 million. Besides that no-one ever got fired for picking AWS ;)
- rubiquity 1y agoNetworking leaving the cloud provider (or even just to another zone on the same cloud) is $0.02 GB. That adds up real fast.
- tadfisher 1y agoNot a justifiable expense when no one else is resilient against their AWS region going down either. Also cross-cloud orchestration is quite dead because every provider is still 100% proprietary bullshit and the control plane is... kubernetes. We settled for kubernetes.
- morshu9001 1y agoAlso if you can't even do cross region, cross cloud won't happen
- dylan604 1y agoCross region isn't simple when you have terabytes of storage in buckets in a region. Building services in other regions without that data doesn't really do any good. Maintaining instances in various regions is easy, but it's that data that complicates everything. If you need to use the instances in a different region because your main region is down, you still can't do anything because those cross region instances can't access the necessary data.
- physicsguy 1y agoWe don’t use AWS at work but we still experienced disruption because lots of our customers do, and use it to transfer data to us. That means we then saw an uplift in data transfers as their systems came back online. There is no panacea. The reason many people use these is because it’s easy and hard to find people that know other clouds and their quirks.
- halis 1y agoJust need to retire the us-east-1 region, it's becoming a meme at this point.
- racl101 1y agoI find it weird many people are just realizing this. I've had this conversation with regards to talking about what should happen if a couple of bad earth quakes, not even "the big one", were to occur. But on the other hand, maybe I hang around too many tech people to not empathically understand the other point of view.
- bee_rider 1y agoUS east is pretty geologically stable I think.
- morshu9001 1y agoWe've seen big outages already but nothing that lasts too long. If an outage became prolonged enough, people would find solutions. We don't know what this massive outage would even look like, so whatever preparation you do, it might still break. Also there are some outages that affect real life like airlines, but tech news overstates some like Facebook. It turns out that FB and IG can be totally broken for a whole day, the world will keep spinning, and they won't even lose users.
- jraph 1y agoI think many (most?) non tech people don't even know that Amazon is first and foremost a cloud provider (and one of the biggest at that, if not the biggest) and that its market thing is almost a side activity at this point.
- morshu9001 1y agoThe expert opinions are more about geopolitics, like maybe don't have all your country's systems realtime depend on a foreign company. If you are just one company whose goal is to maximize uptime without bringing in the complexity of multi-cloud, relying on AWS is reasonable. You probably won't get better uptime using something else, you'll only be down at different times than most others, which in most cases is actually worse.
- kristianc 1y agoFor the kind of person being quoted, the stock in trade is not actually doing anything to fix it, it's in being the person quoted when something goes wrong.
- 0xbadcafebee 1y agoThis is what I call "fool's availability": reducing single points of failure (one cloud provider) without adding any actual redundancy. If you removed AWS/GCP/Azure/etc and just had 100 small providers scattered all over, the result would be hundreds of outages throughout the year, as opposed to one big outage every other year [in one region]. AWS is already way more reliable than any other provider. The real problem here is that companies that use AWS are morons who don't know how to architect/build infrastructure properly. If it's important, it should be built right, regardless of who the provider is. A software building code would mandate how companies could use infrastructure (AWS or any provider) so that important services would not go down when one service or region goes down. This is the basic concept behind things like the electrical code. It doesn't matter how great a public utility is; if your business is wired up so badly that a stiff breeze sets it on fire, just switching utilities isn't gonna help. And some utilities do occasionally have problems that persist down their lines to the customers, so customers need to set up equipment to protect against those failures. Whole-house surge protectors, lightning arresters, EMP shields, etc are necessary so that a rare event doesn't fry expensive customer equipment.
- dudeinjapan 1y agoIts probably worse—a given stack using multiple of these small providers will probably have more “single points of failure” (providers used in series rather than parallel.) (If most companies liked using cloud providers in parallel, they’d already be doing it today between AWS, Azure, and GCP.)
- deleted 1y ago[deleted]
- morshu9001 1y agoYes but most of those companies aren't morons, they're just taking an acceptable risk. Multi-region or multi-cloud setup is nontrivial.
- 0xbadcafebee 1y agoMost companies I've worked for (and have heard about from others) have either lacked the knowledge, or the will, to evaluate risk. They build things until they "just work", and their thought process ends there. They don't examine the design to identify its reliability and security risks. They don't calculate the losses. They still have issues, but they just happen to be acceptable most of the time. Example1: A company's infra goes down, but it doesn't come back up correctly. People run around trying to get it working again. It takes much longer than they hoped/expected, and they lose a lot more money than they expected. This is because they never really understood the risk they were exposed to. If they understood it, they would have done more ahead of time to mitigate that much risk. (today's outage is this case. A lot of companies are going to lose money after today, because their customers are not happy with these "acceptable risks". Presumably, losing this much money due to one outage will not be an acceptable risk in hindsight. So the company either didn't understand its risk, or it did but was too stupid to prevent it) Example 2: A company gets hacked, and its data is either exposed or wiped. This is a much worse result; they can lose tons of money, chase off customers, damage their brand, open them up to lawsuits and fines, even tank the whole company. It's clear that this risk is pretty unacceptable. But it keeps happening. And the reason usually isn't "some genius hacker"; it was a lack of understanding the risk of not investing in security. (there's tons of examples of these in the news. presumably, not investing in security was not an acceptable risk in hindsight when it ended their business! almost always, the people involved in making these products don't know enough about security to understand the risks. but they also don't invest in security training, mandatory security controls, checklists, processes, quality gates, etc) You don't need multi-region or multi-cloud to mitigate reliability risks. Just like you don't need to hire a big security team or invest tons of cash to mitigate security risks. You can use your existing infra and tools, and mitigate both issues. You just have to use them wisely. It takes some effort and time, but you do it once and it pays dividends indefinitely. Building something without identifying its security/reliability risks, and then not calculating those risks' impact, is not acceptable risk; it's ignored risk. Is tanking your company and shedding customers an acceptable risk? Well, there's one way to find out.
- ChrisArchitect 1y agoMore discussion: https://news.ycombinator.com/item?id=45640838 https://news.ycombinator.com/item?id=45640838
- dijit 1y agoAnd we lean into it by saying "Well, if everyone else is down, I get a free pass". (which, is not true in reality if you have ordinary customers).
- Jzush 1y agoIf only there was a system of computers on the Internet that was distributed across the world where we could host things instead of all in one location. We could call it the "cloud".
- 123sereusername 1y ago[dead]
- impure 1y agoWe already have diversification. You can rent a VPS from hundreds of possible companies. And people are very happy with them, it seems every month or two there’s a post here about how some company slashed their cloud bill by switching to a VPS. What we have here is a lock-in and marketing problem.
- jasode 1y ago>You can rent a VPS from hundreds of possible companies. And people are very happy with them, it seems every month or two there’s a post here about how some company slashed their cloud bill by switching to a VPS. Companies are using higher-level "PaaS" suite of services from AWS such as DynamoDB, RedShift, etc and not just the lower-level "IaaS" such as basic EC2 instances or pure containers. Same "lock-in" situation with using the higher-level services from MS Azure and Google Cloud. For those dependent on high-level services, migrating to a VPS like Hetzner or self-hosting is not possible unless they re-invent the AWS stack by installing/babysitting a bunch of open-source software. It's going to be a lot more involved than just installing a PostgreSQL db instance on a VPS.
- SoftTalker 1y ago> It's going to be a lot more involved Yes, and you can't escape that by outsourcing it. The complexity is still there, and it will still bite you when your outsourcer fails to manage it.
- candiddevmike 1y agoSame thing applies to AWS...
- throwaway894345 1y agoI’m not really making a point here as much as an observation, but if my stack that I manage atop VMs in a data center goes down, my customers are pissed at me. If AWS goes down along with half the Internet, my customers are completely sympathetic.
- binary132 1y agoWow, thanks experts! I never could have figured this out without you :)))
- esafak 1y agoThere has to be an Onion article for this.
- 01HNNWZ0MV43FF 1y ago"No way to prevent this, says only region where this regularly happens"
- patrickmcnamara 1y agoThis article isn't written for you. It's written for my mom, etc.
- hippo77 1y agoSurprised to see an article like that even getting shared here. The Guardian seems to be wrong on almost every tech issue.
- binary132 1y agoDoes your mother frequent hackernews?
- patrickmcnamara 1y agoThis article was written for The Guardian, not Hacker News.
- binary132 1y agoyet here it is, posted on hackernews
- 1y ago
- shadowgovt 1y agoSure. Are the "experts" going to pony up the cash to build in redundancy, or change the market fundamentals that make it make more sense for a startup to rush to product on a shoestring and then keep adding features instead of building against not-yet-happened failure modes? If not, I look forward to the next single-point-of-failure outage. And the next. And the next.
- Aeolun 1y agoIt’s only a single region. If anything it shows how many people just double down on the default without any redundancy.
- arbll 1y agoA single region that is a SPOF for global AWS services*
- deleted 1y ago[deleted]
- starman55 1y agoIs us-east-2 services impacted today? which ones?
- stronglikedan 1y ago> It’s only a single region Which was effectively the only region
- heavyset_go 1y agoIt makes us vulnerable to a centrality attack either foreign or domestic. If someone wants to fuck society up, only a handful of data centers, routers, networking junctions, etc could do it.
- ryandvm 1y agoMan, I did not have "AWS us-east-1 will only have TWO 9s this year" on my bingo card.
- aurumque 1y agoFor those of us who have been using AWS for almost 20 years now, I can't imagine why anyone would willingly choose us-east-1 for anything. It is the oldest, highest traffic, most critical path region and is subject to turbulence.
- morshu9001 1y agoIt can make sense to depend on the thing that will attract massive worldwide attention if/when it goes down. Or, more likely, it's just a default people don't change.
- interroboink 1y agoBy some logic, that would mean it is the most battle-tested and highest-stakes (and therefore most carefully-managed) choice. I.e. reasons in favor. Not that I disagree with you, but maybe not for the reasons you say (:
- swiftcoder 1y ago> By some logic, that would mean it is the most battle-tested and highest-stakes (and therefore most carefully-managed) choice As someone who used to work on the inside, us-east-1 has the biggest pile of legacy workarounds for internal AWS issues, it has a variety of legacy API behaviours that don't exist in other regions, and because everyone picks it as the default, it has significantly more pressure on contested resources (i.e. things like spot instance pools). Plus since it's the default in all the tooling, if you ever decide to go multi-region, you'll find tons of things break right away.
- tlogan 1y agoI think it is a little complicated. For example, your service might be using full failover but you use API from other service which are down. Or you might use BART to come to work and you got stuck: https://www.kqed.org/news/12060687/bart-resumes-service-but-delays-remain-after-another-major-disruption https://www.kqed.org/news/12060687/bart-resumes-service-but-...
- cpncrunch 1y ago"The root cause is an underlying internal subsystem responsible for monitoring the health of our network load balancers." https://health.aws.amazon.com/health/status?path=service-history https://health.aws.amazon.com/health/status?path=service-his...
- boznz 1y agoI've really got to get me one of these 'expert' job gigs!
- labrador 1y agoKieran Healy @kjhealy@mastodon.social Always worth taking sentences that use “the Cloud” or “the Internet” and try replacing those phrases with “A shed in Virginia” to see how they hold up. “Our service is fully based in a shed in Virginia”; “All my files are in a shed in Virginia”; “A shed in Virginia was designed to survive a nuclear war”, etc. https://mastodon.social/@kjhealy/115407725852594322 https://mastodon.social/@kjhealy/115407725852594322
- SpicyLemonZest 1y agoSounds like a pretty good shed! Like a lot of pithy commentary on the cloud, this ignores the fact the practical alternative to a shed in Virginia for most businesses is a shelf in the supply closet. "Oops, Jim Bob tripped over the power cord, guess we won't get any emails until the IT guy shows up" - this used to be a routine experience.
- darkwater 1y agoWith a gazillion of shelves, closets, Jims and cables. So if Fortnite's Jim trips on a wire, Canva's Jim is quitely sipping coffee at his desk.
- noir_lord 1y agoOnce had a site wide outage (biggish manufacturing company) of the internet and backup servers because one of the women wanted to plug her hair straighteners in for the xmas party. In a surprise to literally no one that happening on the last friday before xmas break got my "We need to secure the main comms cabinet" (which had the backup server and main ingress for WAN and was in a separate building on other side of site) item that I'd been asking about for months to the top of the list. Still one of my favourite "outages" because I got to my desk, turned PC on, no network, walked across the landing into the main office, opened comms cabinet, plugged it back in and was "resolved" before the MD got to my desk.
- gspencley 1y ago> "Oops, Jim Bob tripped over the power cord, guess we won't get any emails until the IT guy shows up" - this used to be a routine experience. You're not entirely wrong, but you're being hyperbolic too. I'm actually curious how old you are / how long you've worked in tech, because I started out pre-cloud and things weren't nearly as bad or as limited as you suggest. First, on-prem servers are not the only alternative to "cloud." Many businesses, including the ones I worked for, did co-location. The companies owned their own bare metal servers, but would rent a rack in a data centre, and certain things - like the network admin - was entirely outsourced to the data centre / hosting company. You could also rent managed bare metal servers (you still can). This means that you can pretty much outsource your entire IT department, but you're still not doing cloud services. Meaning you've got bare metal servers, someone you're paying at the hosting company is handling security updates and troubleshooting. You don't get things like auto-scaling or serverless or other cloud features, but you also don't have to worry about Jim tripping over the power cable either. There's also still virtual servers. Which is basically a VM running on a server that hosts multiple clients. All of this is to say that the alternative is not "cloud" or "box in a closet." The alternative is "cloud" and a ton of different server options: owned, rented, co-located, on-prem, dedicated, virtual, managed v un-managed (outsource IT vs admin your own) and the list goes on and on.
- greenavocado 1y agoThe only reason we can't leave AWS is because we have 500 terabytes of data in S3
- jewel 1y agoTalk to the other vendors. I know of a place that had about that same amount and decided to have a redundant copy of all of their data in another vendor's S3-compatible product. That vendor paid for all of their egress fees as long as they signed a 12-month contract and used their tool for the migration.
- coredog64 1y agoAWS will credit your egress fees if you incur them via leaving. https://aws.amazon.com/blogs/aws/free-data-transfer-out-to-internet-when-moving-out-of-aws/ https://aws.amazon.com/blogs/aws/free-data-transfer-out-to-i...
- ovaistariq 1y agoWhat other AWS services do you depend on?
- greenavocado 1y agoMostly EC2 for data mining terabytes of historical data stored in S3. Production usage is fairly lightweight compared to the EC2 and S3 stuff. We did cut our bill a lot by moving to single AZ redundancy.
- sunrunner 1y agoThe 'experts' also made similar criticisms with the Fastly outage in 2021 and did anything obvious change as a result? In a week's time no national newspapers will be talking about this. Meanwhile, everyone that spends actual time in these areas: - Knows that running an operation at AWS scale is difficult and any armchair critism from 'experts' is exactly that. Actions speak louder than words. - Understands that the cost of actually accounting for this kind of scenarios is incredibly high for the benefit in most cases - Knows that genuinely 'critical' services (i.e. health) should be designed to account for this, and every other 'serious' issue such as 'I can't log in to Fortnite' just shows what the price and effort of actually making that work is versus how much it costs affected companies when it happens - Knows how much time national newspapers spend actually talking about the importance of multi-region/multi-cloud redundancy, that is, it's zero until the one day where it happens and then it's old news - Is just curious as to just what exactly happened from a technical perspective This isn't to say that good blameless post-mortem shouldn't happen to figure out process and technical issues, but the armchair criticism with no actual followup? All noise, no signal.
- free_bip 1y agoBecause the experts have no say in policy. The only people who have a say are the people bribing (sorry I mean "lobbying") Congress. And even they have very little say because Congress is currently on a hot streak of doing absolutely nothing.
- gnerd00 1y agomaybe your VC overlords need a reality check?
- BrenBarn 1y agoI think all of that is mostly irrelevant. You don't need to pay a huge cost to avoid the small benefit, you don't need every service to be resilient to this, or any of that. You just need multiple different providers so that not everyone gets screwed at once.
- bamboozled 1y ago
- deleted 1y ago[deleted]
- jimmar 1y agoMy company has been ahead of all of this by causing outages in our own data center without waiting for the cloud to do it for them. On a serious note, resiliency takes effort and investment no matter where you host your content.
- _pvzn 1y agoThe "experts" should lay out a good alternative in that case. Smaller providers also run into outages.
- dexterdog 1y agoAnd they all get to claim that they have better uptime to potential customers because nobody other than their current customers remembers their outages.
- JadoJodo 1y ago> "Also in the UK, Ring users complained on social media that their doorbells were not working." I sincerely hope that the base functionality of these doorbells (i.e., triggering the ringing of the bell within the home) is preserved in the event of an internet outage.
- dabinat 1y agoThis is coming right after we switched back to AWS after trying to switch storage to Cloudflare R2. Even with this outage, I still consider AWS more reliable than Cloudflare.
- KronisLV 1y agoSo, how many people will actually switch their setups to multi-cloud as a consequence of this? How many will move over to self-hosting? Or will they just do a post-incident report, wave hands around and do nothing? Because I think it's very much the same way as it is with Cloudflare - while the large vendors aren't always openly hostile, we can just smile and hope that they don't get too keen on reminding us that they're holding us hostage. I don't see that changing anytime soon. I've personally also used Hetzner, Contabo, Scaleway, Vultr, DigitalOcean, Time4VPS and some other platforms, but when people couple their setups to CF/AWS/GCP/Azure, typically that coupling is hard to get rid of and doing so is hard to justify.
- 1970-01-01 1y agoGCP and Azure should be running a 10% sale/discount (Coupon code: RAINYDAY) for new accounts during the week of an AWS outage. The bean counters would take note.
- SkyPuncher 1y agoFor most companies, I suspect this will actually re-affirm _not_ switching to multi-cloud. Lots of businesses who will be completely forgotten as having an outage today because all of their customers were dealing with their own outages and outages in dozens of other providers. Obviously, that doesn't fly for everyone.
- jimbokun 1y agoNobody ever got fired for buying IBM… …no, Microsoft… …no, AWS.
- xp84 1y agoIn 2011 there was some kind of big outage at some major AWS US-east pop. I started a job at a company (very boring B2C startup) which had taken the lesson from that, that "cloud anything is dangerous." They went and bought a bunch of literal servers and installed them in a datacenter, 90 miles away from our offices, and this is where all our applications ran for the remainder of that company's existence (about 6 more years). For the whole time I was at that company, we had somewhat more, and usually more lengthy, outages than the average startup. The only difference is that when some piece of networking gear took a crap, or a disk failed, or whatever, our guys had to diagnose and resolve it (Their karma, I guess, since this was their idea). Anyway, I do think it would be good if at least so-calld 'tech companies' had a little less obsession to outsource everything -- even easy things -- to AWS, GCP, and Azure. I feel that way mainly for cost reasons as many of these services are wildly overpriced. But also we shouldn't kid ourselves by ignoring the advantages of operating at the scale those guys do. They can afford to have multiple absolute wizards available around the clock who make sure that when a problem happens, it's not the kind of "S-show" we had at my old company where we're all on a slack room or zoom or whatever and just guessing at to try for half an hour before we can figure out what the actual issue is.
- 999900000999 1y agoI largely agree with you. When AWS goes down, for most situations I can just go outside and smoke a cigarette and not worry about it. It's someone else's problem.
- robomc 1y agoThis. And when a service goes down it's a lot easier to explain to your client/boss that "half the internet is down" than "our boutique solution is broken so it's just us actually".
- midtake 1y agoThere are many public clouds and VPS providers out there. Who the fuck are these experts? The real issue is that business pricks will cut costs and single-homing in a single availability zone will be the only workable solution. On top of that, infrastructure ops are seen as a nuisance who get in the way of the sexy stuff like shipping your latest code changes now. If you complicate the ops pipeline that gets in the way of sexy dev work. So fuck that just ship lol!
- judahmeek 1y agoI recall reading that when the costs of distribution (but not the costs of discoverability) are low, generally you end up with a power law sort of distribution of consumers to providers, where provider #1 has exponentially more market share than provider #2 and provider #2 has exponentially more market share than provider #3, #4, etc. Examples of this are Windows/Mac, McDonalds/Burger King, Playstation/Xbox, Nvidia/?, AWS/Azure?, Android/iPhone, etc... Basically, the majority of users all using the same dependency/platform/product is basic economics.
- spullara 1y agoproviders should stop using just us-east-1 like idiots.
- dang 1y agoRelated ongoing thread: AWS Multiple Services Down in us-east-1 - https://news.ycombinator.com/item?id=45640838 https://news.ycombinator.com/item?id=45640838 -(1650 comments so far)
- jesterson 1y agoThis is not a provider scarcity problem - there are numerous providers out there, but user's problem - they voluntarily choose crappy service at large scale, believing sales managers "it's reliable".
- 0xbadcafebee 1y agoIt is reliable. Even considering the inflated availability numbers, it's stupidly reliable.
- jesterson 1y agoRecent (and not so recent) events prove it isn't, or is it?
- 0xbadcafebee 1y agoTerms like reliability have specific definitions in computer systems: Term | Definition | Measurement -------------- ----------------------------------- ------------------------------------------- Availability | Basically, system uptime | A percentage over time Durability | Basically, persistence of data | A percentage over time Resiliency | Basically, self-healing | A probability within a time period (usually) Reliability | Basically, operational probability | A probability within a time period (usually) Fault tolerant | Basically, it cannot fail | Binary (it has faults or it doesn't) Unlike more mathy fields, reliability is more of a "quality" that is qualified by one or more measurements (like Mean Time Between Failure). You define your metric, you give an estimate of what that value should be, and if you come in under it, you're reliable. AWS has always stretched the truth when it comes to these numbers, but they do come pretty close to them most of the time. If you can find a different provider who'll even offer a number, it is usually not as close, and there's usually no contract that has any teeth to enforce it. Or they'll give very vague claims that don't get into specifics. At least, not for "cloud providers" (other than the hyperscalers). You can find a datacenter who'll give you a number, but that's for like, their power reliability. That's a very different thing than saying "there is X probability over Y time that a server I run for you will not go down". Partly because it's pretty freakin' hard to wrangle all the different things that can go wrong with so much certainty that you can put a number on it. So most people give things like reliability, durability, availability, etc numbers for specific components of a system. AWS S3 offers 99.999999999% durability and 99.99% availability. Now, did AWS S3 go down completely during the outage? Not as far as I'm aware. Maybe the control plane did, or a management portal, or billing, or something? But I'll bet you the PUT, GET, DELETE operations kept on flowing within 99.99% availability. Some other components in AWS may have been failing like crazy (which may have no guarantees...), but that one component probably stayed up within its guaranteed amount. Design your apps to run on AWS using the components with specific guarantees, and you can estimate how reliable your end product will be. As far as I know, nobody has a better track record for meeting the guarantees. Even considering events like this.