5 ms·
>The fun thing about these types of outages are seeing all of the people that depend upon these services with no graceful fallback. Whats a graceful fallback?
by s_dev 5y ago
>The fun thing about these types of outages are seeing all of the people that depend upon these services with no graceful fallback.
Whats a graceful fallback? Switching to another hosting service when AWS goes down? Wouldn't that present another set of complications for a very small edge case at huge cost?
- rehevkor5 5y agoUsually this refers to falling back to a different region in AWS. It's typical for systems to be deployed in multiple regions due to latency concerns, but it's also important for resiliency. What you call "a very small edge case" is occurring as we speak, and if you're vulnerable to it you could be losing millions of dollars.
- simmanian 5y agoAWS itself has a huge single point of failure on us-east-1 region. Usually, if us-east-1 goes down, others soon follow. At that point, it doesn't matter how many regions you're deploying to.
- scoopertrooper 5y agoMy workloads on Sydney and London are unaffected. I can't speak for anywhere else.
- lowbloodsugar 5y ago"Usually"? When has that ever happened?
- dylan604 5y agohttps://awsmaniac.com/aws-outages/ https://awsmaniac.com/aws-outages/
- lowbloodsugar 5y agoThanks. That shows that OPs claim that "Usually, if us-east-1 goes down, others soon follow." is false.
- dijit 5y agoCognito (and r53) have hard dependencies in us-east-1. It’s mentioned up and down this thread.
- lowbloodsugar 5y agoI don't want to minimize the impact from cognito and r53, but that's quite a different scale of failure than the OP implies. It's been a while since I used AWS, but we had multiple regions and saw no impact to our services in other regions the one time that us-east-1 had a major failure. And we used r53.
- simmanian 5y agoPerhaps I could've been more precise with my words. What I meant to say is IAM and r53 are two of many critical services that all depend on us-east-1. It goes without saying that if those services go down in us-east-1, the whole AWS is affected. This doesn't just happen "usually." When IAM goes down, AWS experiences major issues across all regions. If you were okay, perhaps you got lucky? Our team has 7 different prod regions and we see multiple regions go down every time a problem of this scale occurs. If your product requires 100% uptime, you may need to look at backup options or design your product in such a way that can handle temporary cloud failures.
- stevehawk 5y agoprobably not possible for a lot mroe services than you'd think because AWS Cognito has no decent failover method
- cowmoo728 5y agoI heard from someone at <big media company> that they couldn't switch to their fallback region because they couldn't update DNS on Route53. All the console commands and web interface were failing.
- winrid 5y agoIn this case, just connect over LAN.
- politician 5y agoOr BlueTooth.
- winrid 5y agoWell, BT is a hell of a protocol. I wouldn't wish that on anyone.
- s_dev 5y agoRight -- I think I've misread OP as graceful fallback e.g. working offline. Rather than implement a dynamically switching backup in the event of AWS going down which is not trivial.
- deleted 5y ago[deleted]
- itisit 5y ago> Wouldn't that present another set of complications for a very small edge case at huge cost? One has to crunch the numbers. What does a service outage cost your business every minute/hour/day/etc in terms of lost revenue, reputational damage, violated SLAs, and other factors? For some enterprises, it's well worth the added expense and trouble of having multi-site active-active setups that span clouds and on-prem.
- twistedpair 5y agoThere is a company that delivers broadcast video ads to hundreds of TV stations on demand. The ad has to run and run now, so they cannot tolerate failure. They write the videos to GCS storage in Google Cloud, and to S3 in AWS. Every point of their workflows are checkpointed and cross referenced across GCP and AWS. If either side drops the ball, the other picks it up. So yes, you can design a super fault tolerant system. This company did it because failing to deliver a few ads would mean lose of major contracts.