5 ms·
A lot more companies will go to a multi-cloud active active architecture with maybe even bare metal redundancies.
by 1cvmask 5y ago
A lot more companies will go to a multi-cloud active active architecture with maybe even bare metal redundancies.
- psanford 5y agoI would stay away from any company that thinks "we need to go multicloud" as a response to this outage. This affected a single az in a single region. If it caused you downtime or a partial outage, it means you are not fully resilient to single az failures. The correct thing is to fix your application to handle that. If you can't handle a single az failure there is no way you are going to handle failing over across different cloud providers correctly.
- CubsFan1060 5y agoThere will, however, be a lot of executives _talking_ about going multi cloud.
- luhn 5y agoTo be fair though, from what I heard AZ outage caused an EC2 API brownout, so people couldn't launch new instances in the other AZs. That put a wrench in a lot of multi-AZ architectures. Not advocating for multi-cloud though...
- dragonwriter 5y ago> If it caused you downtime or a partial outage, it means you are not fully resilient to single az failures. Given the number of AWS global services that have dependencies on infra in US-EAST-1 (and, from the impacts of this and other past outages, seen vulnerable to single-AZ failures in US-EAST-1) that's...less avoidable for certain regions/AZs than one might naively expect. Most clouds seem to have at least some degree of this kind of vulnerability.
- nonane 5y ago> This affected a single az in a single region. If it caused you downtime or a partial outage, it means you are not fully resilient to single az failures. This is not true. Amazon is not being upfront about what happened here. It was simply not a single AZ failure. Our us-east-1 ELB load balancers were hosed and were unable to direct traffic to other AZs - they simply stopped working an were dropping traffic. We tried creating load balancers in different AZs and that didn't work either. How can you be resilient to single AZ failures if load balancers stop working region wide during a single AZ outage?
- acdha 5y agoDid your TAM go into any details on that? Over a couple hundred load-balancers, the only issue we had was taking longer to register new instances and that affected only a couple of them. Running services weren't interrupted, latency remained, etc. which is what I'd expect for a single AZ failure.
- deleted 5y ago[deleted]
- tyingq 5y agoMulti-cloud is odd to me, unless you're a company selling a service to cloud customers. By definition, you would have to either go lowest-common-denominator, or build complicated facades in front of like services. If you're going lowest-common-denominator, then multi-old-school-hosting would be far cheaper.
- isbvhodnvemrwvn 5y agoYou'd have to go worse than lowest common denominator. To have active-active replication you need to take the cost of latency between the different DCs of the various providers, and this thing will kill performance.
- ransom1538 5y agome: "Yeah! AWS went down a few times, I think we should pay double!" vp: no.
- emodendroket 5y agoMore than double realistically. But you could achieve most of the benefit at much lower cost by going multi-AZ.
- karmasimida 5y agoInfra is pure cost center. A few hours downtime is not going to justify double cost, and worse whose benefits can only be demonstrated during those few hours. It will get shutdown immediately. And what is more important, corporate world really don't care that much about downtime, not as much as they care about who to assign the blame, in this case AWS is a perfect irresistible externality much like a natural disaster. Disclaimer: Ex-AWS employee
- profmonocle 5y agoEgress fees can make replicating databases, storage buckets, etc. between clouds very expensive. Multi-region is a much more affordable option. Multi-region outages aren't unheard of among the major cloud operators, but they're less common than single-AZ or single-region outages. IMO, most companies just aren't sensitive enough to downtime that multi-AZ + multi-region deployment within a single cloud provider isn't good enough.
- beermonster 5y agoAs large cloud outages become more frequent and the impact greater each time, I feel it’s more likely people will reconsider moving workloads off-prem.