5 ms·
Netflix actually added the additional AZs because of a prior outage that did take them down. "After a 2012 storm-related power outage at Amazon during which Ne
by angstrom 7y ago
Netflix actually added the additional AZs because of a prior outage that did take them down.
"After a 2012 storm-related power outage at Amazon during which Netflix suffered through three hours of downtime, a Netflix engineer noted that the company had begun to work with Amazon to eliminate “single points of failure that cause region-wide outages.” They understood it was the company’s responsibility to ensure Netflix was available to entertain their customers no matter what. It would not suffice to blame their cloud provider when someone could not relax and watch a movie at the end of a long day."
https://www.networkworld.com/article/3178076/why-netflix-didnt-sink-when-amazon-s3-went-down.html https://www.networkworld.com/article/3178076/why-netflix-did...
- aaronblohowiak 7y agoWe went multi-region as a result of the 2012 inc. source: I now manage the team responsible for performing regional evacuations (shifting traffic and scaling the savior regions).
- mkl 7y agoThat sounds fascinating! How often does your team have to leap into action?
- aaronblohowiak 7y agoWe don’t usually discuss the frequency of unplanned failovers, but I will tell you that we do a planned failover at least every two weeks. The team also uses traffic shaping to perform whole system load tests with production traffic, which happens quarterly.
- justinator 7y agoDo you do any chaos testing? Seems like it would slot right in, there.
- Zobat 7y agoI'd say yes. I heard about this tool just a week ago at a developer conference. https://github.com/Netflix/chaosmonkey https://github.com/Netflix/chaosmonkey
- arainwater 7y agothey have invented the term, so probably yes :)
- a_t48 7y agoNetflix was a pioneer of chaos testing, right? https://en.m.wikipedia.org/wiki/Chaos_engineering https://en.m.wikipedia.org/wiki/Chaos_engineering
- aaronblohowiak 7y agohttps://www.oreilly.com/library/view/chaos-engineering/9781491988459/ https://www.oreilly.com/library/view/chaos-engineering/97814... ;)
- azimuth11 7y agoI think some Google engineers published a free Meap book on service relatability and uptime guarantees. Seemingly counterintuitive, scheduling downtime, without other teams’ prior knowledge, encourages teams to handle outages properly and reduce single points of failure, among other things.
- fnord123 7y agoService Reliability Engineering is on OReilly press. It's a good book. Up there with ZeroMQ and Data Intensive Applications as maybe the best three books from OReilly in the past ten years.
- fnord123 7y agoDerp, Site Reliability Engineering. https://landing.google.com/sre/books/ https://landing.google.com/sre/books/