4 ms·
And apparently they had never tried rebooting some of the most important parts of that system. Just when you start to think that someone's really gotten it righ
by tlack 10y ago
And apparently they had never tried rebooting some of the most important parts of that system. Just when you start to think that someone's really gotten it right you come to learn they're just fumbling around in the dark like everyone else.
- jonhohle 10y agoMy interpretation of this is that the indexing system was resilient to lost of a certain amount of capacity (probably around ⅓ + 1 host). As a guess, the indexing system probably used some form of consensus (e.g. paxos) which has had an active leader for years. Deployments stay within that capacity constraint, so while hosts have been restarted and replaced (data center migrations, hardware lease expiration, failures, upgrades, etc.), they may have not recently run into a situation where quorum wasn't available for a partition, especially at the scale of restarting the entire fleet. Since restarting the entire fleet would incur downtime of all relevant S3 operations, it's unlikely that it was something ever intentionally done in production (and they may or may not have run that scenario in other environments). Source: I used to run several large scale services at Amazon.
- pyre 10y agoTo put what @jonhohle said another way, Amazon had probably never brought up the entirety of S3 from Zero to Production-Ready on in a production environment before. I wouldn't necessarily classify this as "fumbling around in the dark." Perhaps they should have tested this in a simulated environment, but (to be fair) on a distributed fault-tolerant system, it probably wasn't a top-priority situation to test.
- joelthelion 10y ago> Perhaps they should have tested this in a simulated environment What makes you think they didn't?
- idlewords 10y agoMoreover, beyond a certain scale, it becomes really hard to simulate.