3 ms·
I gather, the root cause was a latent race condition in the DynamoDB DNS management system that allowed an outdated DNS plan to overwrite the current one, resul
by WaitWaitWha 1y ago
I gather, the root cause was a latent race condition in the DynamoDB DNS management system that allowed an outdated DNS plan to overwrite the current one, resulting in an empty DNS record for the regional endpoint.
Correct?
- tptacek 1y agoI think you have to be careful with ideas like "the root cause". They underwent a metastable congestive collapse. A large component of the outage was them not having a runbook to safely recover an adequately performing state for their droplet manager service. The precipitating event was a race condition with the DynamoDB planner/enactor system. https://how.complexsystems.fail/ https://how.complexsystems.fail/
- 1970-01-01 1y agoWhy can't a race condition bug be seen as the single root cause? Yes, there were other factors that accelerated collapse, but those are inherent to DNS, which is outside the scope of a summary.
- tptacek 1y agoBecause the DNS race condition is just one flaw in the system. The more important latent flaw† is probably the metastable failure mode for the droplet manager, which, when it loses connectivity to Dynamo, gradually itself loses connectivity with the Droplets, until a critical mass is hit where the Droplet manager has to be throttled and manually recovered. Importantly: the DNS problem was resolved (to degraded state) in 1hr15, and fully resolved in 2hr30. The Droplet Manager problem took much longer! This is the point of complex failure analysis, and why that school of thought says "root causing" is counterproductive. There will always be other precipitating events! † which itself could very well be a second-order effect of some even deeper and more latent issue that would be more useful to address!
- 1970-01-01 1y agoTwo different questions here. 1. How did it break? 2. Why did it collapse? A1: Race condition A2: What you said.
- tptacek 1y agoWhat is the purpose of identifying "root causes" in this model? Is the root cause of a memory corruption vulnerability holding a stale pointer to a freed value, or is it the lack of memory safety? Where does AWS gain more advantage: in identifying and mitigating metastable failure modes in EC2, or in trying to identify every possible way DNS might take down DynamoDB? (The latter is actually not an easy question, but that's the point!)
- 1970-01-01 1y agoTwo things can be important for an audience. For most, it's the race condition lesson. Locks are there for a reason. For AWS, it's the stability lesson. DNS can and did take down the empire for several hours.
- tptacek 1y agoDid DNS take it down, or did a pattern of latent failures take it down? DNS was restored fairly quickly! Nobody is saying that locks aren't interesting or important.