3 ms·
Thanks for the nice and informative post. Can you add a bit more color to extreme chaos failure with maybe a more concrete action and its outcome? I'm assuming
by gshx 11y ago
Thanks for the nice and informative post. Can you add a bit more color to extreme chaos failure with maybe a more concrete action and its outcome? I'm assuming you're pre-classifying availability/sla's for the serving parts of the search system and its sub-systems. Also, a real world example of a silent failure observed in your systems will be nice to learn from.
- henakama 11y agoThanks for the feedback! I do have an example of a silent failure that would be classified as extreme chaos. Before we went into production, we were having problems with leader election that could cause multiple masters per cluster. This is an extreme chaos state not only because it could result in corrupted or missing data (as a cloud service, we take data integrity very, very seriously), but also because it didn't raise any immediate notifications, only downstream errors. This problem has since been fixed, but even before we did that, the highest priority was to ensure our system detected this state as quickly as possible. We added a service that monitored cluster leader status and that quickly alerted us whenever anything was unexpected. Even after fixing the problem, I've set the Search Chaos Monkey to injecting failure into leader election at various points to ensure that the cluster can always recover in an expected way that won't put customer data at risk.