6 ms·
This part is also interesting: > While this is an operation that we have relied on to maintain our systems since the launch of S3, we have not completely resta
by conorh 10y ago
This part is also interesting:
> While this is an operation that we have relied on to maintain our systems since the launch of S3, we have not completely restarted the index subsystem or the placement subsystem in our larger regions for many years.
These sorts of things make me understand why the Netflix "Chaos Gorilla" style of operating is so important. As they say in this post:
> We build our systems with the assumption that things will occasionally fail
Failure at every level has to be simulated pretty often to understand how to handle it, and it is a really difficult problem to solve well.
- urda 10y agoI think you meant "Chaos Monkey" [1]. [1] https://github.com/Netflix/chaosmonkey https://github.com/Netflix/chaosmonkey
- deleted 10y ago[deleted]
- tantalor 10y agoNo, Chaos Gorilla is similar to Chaos Monkey, but simulates an outage of an entire Amazon availability zone http://techblog.netflix.com/2011/07/netflix-simian-army.html http://techblog.netflix.com/2011/07/netflix-simian-army.html
- ceejayoz 10y agohttp://techblog.netflix.com/2011/07/netflix-simian-army.html http://techblog.netflix.com/2011/07/netflix-simian-army.html > Chaos Gorilla is similar to Chaos Monkey, but simulates an outage of an entire Amazon availability zone. We want to verify that our services automatically re-balance to the functional availability zones without user-visible impact or manual intervention.
- deleted 10y ago[deleted]
- deleted 10y ago[deleted]
- saisundar 10y agoChaos gorilla is a thing as well, simulates outage of an entire AZ. http://techblog.netflix.com/2011/07/netflix-simian-army.html http://techblog.netflix.com/2011/07/netflix-simian-army.html
- adzicg 10y agoI recently saw a talk where they referred to Chaos Monkey (kills instances), Chaos Gorilla (kills many instances for a single service in a single region) and Chaos Kong (takes an entire region offline)
- Pxtl 10y agoThey have an entire simian army of chaos for the purposes of simulated destruction of their network.
- snewman 10y ago> Failure at every level has to be simulated pretty often to understand how to handle it, and it is a really difficult problem to solve well. Exactly. It seems likely that Amazon tests the restart operation, but it would be hard to test it at full us-east-1 scale. Running a full S3 test cluster at that scale would likely be a prohibitive expense. Perhaps the "index subsystem" and "placement subsystem" are small enough for full-scale tests to be tractable, but certainly not cheap, and how often do you run it? Also, hindsight is 20/20, but before this incident it might have been hard to identify "full-scale restart of the index subsystem" as rising to the top of the list of things to test. One approach is to try to extrapolate from smaller-scale tests. It would be interesting to know what kinds of disaster testing Amazon does do, and at what scale, and whether a careful reading could have predicted this outcome.
- pbkhrv 10y ago> Perhaps the "index subsystem" and "placement subsystem" are small enough for full-scale tests to be tractable, but certainly not cheap, and how often do you run it? Rough guide: CT = cost of 1 full scale test with necessary infrastructure and labor costs added up CF = amount of money paid out in SLA claims + subjective estimate of business lost due to reputation damage etc PF = estimate of probability of this event happening in a given year if PF * CF > CT, then you run such a test at least once a year. Think of such an expense as an insurance premium. What Netflix does with their simian army is amortize the cost of doing the test across millions of tests per year and the extra design complications arising from having to deal with failures that often.
- hollander 10y agoTesting a full zone test is only possible when they have a new zone available, unused. I bet they do these test, and they now have a new scenario to test.
- snewman 10y agoThey also probably have one or more test regions where they could perform a test like this. But it's presumably not at nearly the same scale as us-east-1, the region affected by this incident. And to a considerable extent the problem was one of scale. The writeup makes the recovery sound fairly straightforward; but due to the sheer size of S3 in this region, it took hours for the system to come back up, which was apparently unexpected. (Nit: this incident affected a region, not a zone. us-east-1 is a region, which is divided into zones us-east-1a, us-east-1b, etc. S3 operates on regions.)
- spydum 10y agoTotally agree. Would also point out that if you have systems up for many years, they like haven't been updated in the same... shouldn't people find that alarming?
- InclinedPlane 10y agoYep. There's a transition period where you can't rely on redundancy any longer because there are so many components that it's basically inevitable that at any given time somewhere something will be in a degraded state. So you design for that case, the degraded normalcy case. You make something failing somewhere a non-emergency. It takes a lot of work to do but when you have things working in that way then you can guarantee that you're in that state by testing it routinely in production.