4 ms·
AWS US-East-1 Outage: Postmortem and Lessons Learned
- thread_id 5y agoJeremy Daly, 'Off-by-none' "For the other 99.99% of us building workloads in the cloud, a multi-hour outage may sting, and perhaps even result in significant revenue losses, but compared to the cost of implementing and maintaining solutions to mitigate these outages (that only happen every few years), it’s a drop in the bucket. I’ve got more important things to focus on." This is an interesting perspective. All of our servers were up and running, our applications were accessible and running. We manage our DNS through an external provider. However, by design, our applications integrate tightly with multiple AWS services for various architectural components. The features that depend on those services were faililng intermitently throughout the event window. As a remediation, we could replace all of those services with various platforms hosted on servers or provided by SaaS - therefore eliminating the dependency on AWS. Our cost model, support model, and speed of delivery would change significantly. This would obviate our reasons for adopting cloud based architecture in the first place. Instead our approach is to examine how we have also tightly coupled our integration to us-east-1. And look for 'High Availability' design patterns that will loosely couple to a region with the ability to fail over to warn/hot region when needed. Let me know how you are defining your response. I would like to consider additional points of view as we work through this problem.