4 ms·
I think this is a good example of how the "cloud" is not a silver bullet to making your site always up. AWS provides a way to keep it up, but it is up to each d
by mathrawka 15y ago
I think this is a good example of how the "cloud" is not a silver bullet to making your site always up. AWS provides a way to keep it up, but it is up to each developer to ensure that they are using AWS in a way to make sure their site can handle problems in one availability zone.
I think we will see more of a focus from big users of AWS about focusing on how to create a redundant service using AWS. Or at least I hope we will!
- tybris 15y agoNetFlix should sell their chaos monkey as a commercial product.
- justinsb 15y agoDon't let AWS hear that, or they'll charge us for their failures by rebranding it as a feature.
- jedberg 15y agoThis outage is affecting all AZ's in the East. So even a multizone setup wouldn't help for this one. Only a multiregion setup. This outage is a lot like having your entire datacenter lose power.
- mathrawka 15y agoI thought AZs were supposed to be different physical data centers. If that is not the case, then having a multi-region setup would be a necessity for any major sites on AWS. Perhaps there will be a time where to truly be redundant, one would need to use multiple cloud providers. Which would be a _huge_ pain to do now I imagine, with all the provider lock-ins we have.
- jedberg 15y ago> I thought AZs were supposed to be different physical data centers. They are. Which means this is probably a software issue or some other systemic issue.
- mathrawka 15y agoWell that isn't supposed to happen :( I think I'll go home and wait it out there, but it appears that they are having some progress in recovering it. But our site is still affected.
- 6ren 15y agoIf "multi-region" means North America, Europe and Asia Pacific, doing so would also improve world-wide latency (e.g. here in Australia...). Could you use this outage to justify switching to multi-region?
- Maxious 15y agoA blog post last month touched on this: "Q: Why is reddit tied so tightly to the affected availability zone? A: When we started with Amazon, our code was written with the assumption that there would be one data center. We have been working towards fixing this since we moved two years ago. Unfortunately, progress has been slow in this area. Luckily, we are currently in a hiring round which will increase the technical staff by 200% :) These new programmers will help us address this issue." Not sure if the costs of data transfer between regions (charged at full internet price) would justify the added reliability/lower latency though.
- alecco 15y agoI don't know reddit, but concurrency is very hard with high latency.
- justinsb 15y agoAll well and good, but the elephant in the room is that multiple availability zones have failed at the same time. It looks like AWS have a single point of failure they weren't previously aware of.