4 ms·
“Based on our investigation, the issue appears to be related to DNS resolution of the DynamoDB API endpoint in US-EAST-1. We are working on multiple parallel pa
by stepri 1y ago
“Based on our investigation, the issue appears to be related to DNS resolution of the DynamoDB API endpoint in US-EAST-1. We are working on multiple parallel paths to accelerate recovery.”
It’s always DNS.
- commandersaki 1y agoSomeone probably failed to lint the zone file.
- huflungdung 1y ago[dead]
- DrewADesign 1y agoDNS strikes me as the kind of solution someone designed thinking “eh, this is good enough for now. We can work out some of the clunkiness when more organizations start using the Internet.” But it just ended up being pretty much the best approach indefinitely.
- cindyllm 1y ago[dead]
- movpasd 1y agoSeems like an example of "worse is better". The worse solution has better survival characteristics (on account of getting actually made).
- DrewADesign 1y agoI wouldn’t say it’s the worst… a largely decentralized worldwide namespace is not an easy thing to tackle and for the most part it totally works.
- ifwinterco 1y agoI actually think the design of DNS is really cool. I'm sure we could do better designing from a clean slate today, especially around security (designing with the assumption of an adversarial environment). But DNS was designed in the 80s! It's actually a minor miracle it works as well as it does
- bayindirh 1y agoEven when it's not DNS, it's DNS.
- JoBrad 1y agoSometimes it’s BGP. /s
- shamil0xff 1y agoMight just be BGP dressed as DNS
- Nextgrid 1y agoI wonder how much of this is "DNS resolution" vs "underlying config/datastore of the DNS server is broken". I'd expect the latter.
- huflungdung 1y agoI don’t think it is DNS. The DNS A records were 2h before they announced it was DNS but _after_ reporting it was a DNS issue.
- wdfx 1y ago... wonders if the dns config store is in fact dynamodb ...
- CaptainOfCoit 1y agoI feel like even Amazon/AWS wouldn't be that dim, they surely have professionals who know how to build somewhat resilient distributed systems when DNS is involved :)
- grogers 1y agoI doubt a circular dependency is the cause here (probably something even more basic). That being said, I could absolutely see how a circular dependency could accidentally creep in, especially as systems evolve over time. Systems often start with minimal dependencies, and then over time you add a dependency on X for a limited use case as a convenience. Then over time, since it's already being used it gets added to other use cases until you eventually find out that it's a critical dependency.
- kjsingh 1y agoDNS is managed by Route53 which has no dependency on Dynamodb for data plane
- ej_campbell 1y agoBackground on the service: https://aws.amazon.com/builders-library/reliability-and-constant-work/ https://aws.amazon.com/builders-library/reliability-and-cons...
- koliber 1y agoIt's always US-EAST-1 :)
- us0r 1y agoOr expired domains which I suppose is related?
- oneeyedpigeon 1y agoDowntime Never Stops!
- indycliff 1y agothe answer is always DNS
- dexterdog 1y agoThat's why they wrote the haiku
- nijave 1y agoI don't think that's necessarily true. The outage updates later identified failing network load balancers as the cause--I think DNS was just a symptom of the root cause I suppose it's possible DNS broke health checks but it seems more likely to be the other way around imo
- lkjdsklf 1y agoI don’t work for aws, but a different cloud provider so this is not a description of this incident, but an example of the kind of thing that can happen One particular “dns” issue that caused an outage was actually a bug in software that monitors healthchecks. It would actively monitor all servers for a particular service (by updating itself based on what was deployed) and update dns based on those checks. So when the health check monitors failed, servers would get removed from dns within a few milliseconds. Bug gets deployed to health check service. All of a sudden users can’t resolve dns names because everything is marked as unhealthy and removed from dns. So not really a “dns” issue, but it looks like one to users