4 ms·
This affected on region, us-east-1, which is the oldest and largest aws region, and also hosts many core aws services. Each region has several availability zon
by matteotom 5y ago
This affected on region, us-east-1, which is the oldest and largest aws region, and also hosts many core aws services. Each region has several availability zones (AZs) that are basically whole datacenters. In theory AZs should be mostly isolated, but as we saw bugs happen. Reading between the lines of the status updates, it sounds like this either affected some core infra shared between AZs, or was from a change rolled out safely previously (maybe last night) and failed later (eg due to increased load in the morning).
At the end of the day, building redundant systems is expensive (and may introduce whole new bugs), so they probably did the math and figured the risk of a whole region outage was less than the cost of building redundancy in some systems that are hard to make redundant.
- smsm42 5y agoThe whole concept that there are "core AWS services" living in one single zone sounds like anti-thesis to everything AWS should be about. What's the use of having all this nice distributed setup - and paying for it! - if a single failure in a single zone takes everything down anyway? I mean sure, maybe they built it on a shoe-string budget years ago - but since then, they had years and billions, and still didn't bother to fix it?