6 ms·
Agreed with the expertise thing, but to add to your summary of the post-mortem, it looks like human error compounded by: 1) An architecture bug (the EBS "contr
by ianso 15y ago
Agreed with the expertise thing, but to add to your summary of the post-mortem, it looks like human error compounded by:
1) An architecture bug (the EBS "control plane" cuts across Availability Zones and EBS clusters, leading to a single point of failure: this is what broke the "service contract"),
2) a spec/programming bug: No aggressive back-off on retry attempts of EBS ops, and
3) two separate logical bugs: the race condition in the EBS nodes & problems with MySQL replication.
I think that's everything. It just goes to show, most disasters in very well-engineered systems are generally the result of a series of things all going wrong at once, not individual failures...
- OllieJones 15y agoAgree completely. We are still in the early days of AWS-like technology. Providers and users both will experience some serious issues in the years to come. Heck, users of municipal water systems still experience outages and that technology is arguably mature. The really encouraging thing is that the Amazon post-mortem writeup indicates they're taking it very seriously.