7 ms·
US-East-1 is more than just a normal region. It also provides the backbone for other services, including those in other regions. Thus simply being in another re
by JCM9 1y ago
US-East-1 is more than just a normal region. It also provides the backbone for other services, including those in other regions. Thus simply being in another region doesn’t protect you from the consistent us-east-1 shenanigans.
AWS doesn’t talk about that much publicly, but if you press them they will admit in private that there are some pretty nasty single points of failure in the design of AWS that can materialize if us-east-1 has an issue. Most people would say that means AWS isn’t truly multi-region in some areas.
Not entirely clear yet if those single points of failure were at play here, but risk mitigation isn’t as simple as just “don’t use us-east-1” or “deploy in multiple regions with load balancing failover.”
- helsinkiandrew 1y ago>US-East-1 is more than just a normal region. It also provides the backbone for other services, including those in other regions I thought that if us-east-1 goes down you might not be able to administer (or bring up new services) in other zones, but if you have services running that can take over from us-east-1, you can maintain your app/website etc. I haven’t had to do this for several years but that was my experience a few years ago on an outage - obviously it depends on the services you’re using. You can’t start cloning things to other zones after us-east-1 is down - you’ve left it too late
- cmiles8 1y agoWell that sounds like exactly the sort of thing that shouldn’t happen when there’s an issue given the usual response is to spin things up elsewhere, especially on lower priority services where instant failover isn’t needed.
- sgarland 1y agoIt depends on the outage. There was one a year or two ago (I think? They run together) that impacted EC2 such that as long as you weren’t trying to scale, or issue any commands, your service would continue to operate. The EKS clusters at my job at the time kept chugging along, but had Karptenter tried to schedule more nodes, we’d have had a bad time.
- bpicolo 1y agoStatic stability is a very valuable infra attribute. You should definitely consider how statically stable your services are in architecting them
- yencabulator 1y agoMeanwhile, AWS has always marketed itself as "elastic". Not being able to start new VMs in the morning to handle the daytime load will wreck many sites.
- Yeul 1y agoInternet was supposed to be a communication network if the East Coast was nuked. What it turned into was Daedalus from Deus Ex lol.
- t_sawyer 1y agoYeah because Amazon engineers are hypocrites. They want you to spend extra money for region failover and multi-az deploys but they don't do it themselves.
- ajkjk 1y agoThey absolutely do do it themselves..
- falcor84 1y agoWhat do you mean? Obviously, as TFA shows and as others here pointed-out, AWS relies globally on services that are fully-dependent on us-east-1, so they aren't fully multi-region.
- ajkjk 1y agoThe claim was that that they're total hypocrites aren't multi region at all. That's totally false, the amount of redundancy in aws is staggering. But there are foundational parts which, I guess, have been too difficult to do that for (or perhaps they are redundant but the redundancy failed in this case? I dunno)
- t_sawyer 1y agoThere's multiple single points of failure for their entire cloud in us-east-1. I think it's hypocritical for them to push customers to double or triple their spend in AWS when they themselves have single points of failure on a single region.
- ajkjk 1y agoThat's absurd. It's hypocritical to describe best practices as best practices because you haven't perfectly implemented them? Either they're best practice or they aren't. The customers have the option of risking non-redundancy also, you know.
- 1y ago
- bravetraveler 1y agoCall me crazy, because this is, perhaps it's their "Room 641a". The purpose of a system is what it does, no point arguing 'should' against reality, etc. They've been charging a premium for, and marketing, "Availability" for decades at this point. I worked for a competitor and made a better product: it could endure any of the zones failing.
- voxadam 1y ago> perhaps it's their "Room 641a". For the uninitiated: https://en.wikipedia.org/wiki/Room_641A https://en.wikipedia.org/wiki/Room_641A
- jf 1y agoInteresting. Langley isn’t that far away
- deleted 1y ago[deleted]
- nevir 1y agoIt's really not that nefarious. IAD datacenters have forever been the place where Amazon software developers implement services first (well before AWS was a thing). Multi-AZ support often comes second (more than you think; Amazon is a pragmatic company), and not every service is easy to make TRULY multi-AZ. And then other services depend on those services, and may also fall into the same trap. ...and so much of the tech/architectural debt gets concentrated into a single region.
- bravetraveler 1y agoRight, like I said: crazy. Anything production with certain other clouds must be multi-AZ. Both reinforced by culture and technical constraints. Sometimes BCDR/contract audits [zones chosen by a third party at random].
- nevir 1y ago
- gchamonlive 1y agoBeen a while since I last suffered from AWS arbitrary complexity, but afaik you can only associate certificates to cloudfront if they are generated in us-east-1, so it's undoubtedly a single point of failure for all CDN if this is still the case.
- kokanee 1y agoI worked at AMZN for a bit and the complexity is not exactly arbitrary; it's political. Engineers and managers are highly incentivized to make technical decisions based on how they affect inter-team dependencies and the related corporate dynamics. It's all about review time.
- sharpy 1y agoI have seen one promo docket get rejected for doing work that is not complex enough... I thought the problem was challenging, and the simple solution brilliant, but the tech assessor disagreed. I mean once you see there is a simple solution to a problem, it looks like the problem is simple...
- bdbdkdksk 1y agoI had a job interview like this recently: "what's the most technically complex problem you've ever worked on?" The stuff I'm proudest of solved a problem and made money but it wasn't complicated for the sake of being complicated. It's like asking a mechanical engineer "what's the thing you've designed with the most parts"
- arethuza 1y agoI was once very unpopular with a team of developers when I pointed out a complete solution to what they had decided was an "interesting" problem - my solution didn't involve any code being written.
- SoftTalker 1y agoI suppose it depends on what you are interviewing for but questions like that I assume are asked more to see how you answer than the specifics of what you say. Most web jobs are not technically complex. They use standard software stacks in standard ways. If they didn't, average developers (or LLMs) would not be able to write code for them.
- xbar 1y agoThis set of facts comes to light every 3-5 years when US-East-1 has another failure. Clearly they could have architected their way out of this blast radius problem by now, but they do not. Why would they keep a large set of centralized, core traffic services in Virginia for decades despite it being a bad design?
- firesteelrain 1y agoIt’s probably because there is a lot of tech debt plus look at where it is - Virgina. It shouldn’t take much of imagination to figure out why that is strategic
- dsr_ 1y agoThey could put a failover site in Colorado or Seattle or Atlanta, handling just their infrastructure. It's not like the NSA wouldn't be able to backhaul from those places.
- deleted 1y ago[deleted]
- knotimpressed 1y agoYou mean the surveillance angle as reason for it being in Virginia?
- deleted 1y ago[deleted]
- AtlasBarfed 1y agoWhat is the motivation of an effective Monopoly to do anything? I mean look at their console. Their console application is pretty subpar.
- cyberax 1y agoAWS _had_ architected away from single-region failure modes. There are only a few services that are us-east-1 only in AWS (IAM and Route53, mostly), and even they are designed with static stability so that their control plane failure doesn't take down systems. It's the rest of the world that has not. For a long time companies just ran everything in us-east-1 (e.g. Heroku), without even having an option to switch to another region.
- api 1y agoMy contention for a long time has been that cloud is full of single points of failure (and nightmarish security hazards) that are just hidden from the customer. "We can't run things on just a box! That's a single point of failure. We're moving to cloud!" The difference is that when the cloud goes down you can shift the blame to them, not you, and fixing it is their problem. The corporate world is full of stuff like this. A huge role of consultants like McKinsey is to provide complicated reports and presentations backing the ideas that the CEO or other board members want to pursue. That way if things don't work out they can blame McKinsey.
- raw_anon_1111 1y agoYou act as if that is a bug not a feature. As hypothetically someone who is responsible for my site staying up, I would much rather blame AWS than myself. Besides none of your customers are going to blame you if every other major site is down.
- unethical_ban 1y agoAs someone who hypothetically runs a critical service, I would rather my service be up than down.
- raw_anon_1111 1y agoAnd you have never had downtime? If your data center went down - then what?
- unethical_ban 1y agoI'm saying the importance is on uptime, not on who to blame, when services are critical. You don't have one data center with critical services. You know lots of companies are still not in the cloud, and they manage their own datacenters, and they have 2-3 of them. There are cost, support, availability and regulatory reasons not to be in the cloud for many parties.
- 1y ago
- qaq 1y agoeven if us-east-1 was a normal region there is not enough spare capacity to take up all the workloads from us-east-1 in other regions so t's a moot point
- nevir 1y agoIt also doesn't help that most companies using AWS aren't remotely close to multi-region support, and that us-east-1 is likely the most populated region.
- einrealist 1y agoIt sounds like they want to avoid split-brain scenarios as much as possible while sacrificing resilience. For things like DNS, this is probably unavoidable. So, not all the responsibility can be placed on AWS. If my application relies on receipts (such as an airline ticket), I should make sure I have an offline version stored on my phone so that I can still check in for my flight. But I can accept not to be able to access Reddit or order at McDonalds with my phone. And always having cash at hand is a given, although I almost always pay with my phone nowadays. I hope they release a good root cause analysis report.
- immibis 1y agoIt's not unavoidable for DNS. DNS is inherently eventually consistent anyway, due to time-based caching.
- einrealist 1y agoSure, but you want to make sure that changes propagate as soon as possible from the central authority. And for AWS, the control plane for that authority happens to be placed in US-EAST-1. Maybe Blockchain technology can decentralize the control plane?
- immibis 1y agoOr Paxos or Raft...
- masfuerte 1y agoAmazon are planning to launch the EU Sovereign Cloud by the end of the year. They claim it will be completely independent. It may be possible then to have genuine resiliency on AWS. We'll see.
- louthy 1y agoThen it will be eu-east-1 taking down the EU
- samcat116 1y agoThis is the difference between “partitions” and “regions”. Partitions have fully separate IAM, DNS names, etc. This is how there are things like US Gov Cloud, the Chinese AWS cloud, and now the EU sovereign cloud
- JCM9 1y agoYes, although unfortunately it’s not how AWS sold regions to customers. AWS folks consistently told customers that regions were independent and customers architected on that belief. It was only when stuff started breaking that all this crap about “well actually stuff still relies on us-east-1” starts coming out.
- Agingcoder 1y agoYes - they told me quite specifically that until they launch their their sovereign cloud, the mothership will be around.
- immibis 1y agoWhich are lies btw - Amazon has admitted the "EU sovereign cloud" is still susceptible to US government whims.
- seany 1y agogov, iso*, cn are also already separate (unless you need to mess with your bill, or certain kinds of support tickets)
- belter 1y ago> Thus simply being in another region doesn’t protect you from the consistent us-east-1 shenanigans. Well it did for me today...Dont use us-east-1 explicitly just other regions and I had no outage today...( I get the point about the skeletons in the closet of us-east-1 ...maybe the power plug goes via Bezos wood desk? )
- thayne 1y agoThere are hints at in their documentation. For example ACM certs for cloudfront and KMS keys for route53 DNSSEC have to be in the us-east1 region.
- everfrustrated 1y agoHowever these services don't need high write uptime.
- cyberax 1y agoFWIW, I tried creating a DNSSEC entry for one of my domains during the outage, and it worked just fine.