4 ms·
Thundering herd / cascading outages. You take out a large enough portion of your fleet, and the remaining load overloads your remaining nodes one by one as they
by arccy 2mo ago
Thundering herd / cascading outages. You take out a large enough portion of your fleet, and the remaining load overloads your remaining nodes one by one as they restart, so you can never have enough healthy nodes.
- erulabs 2mo agoSRE team debates correctness versus availability for the 540th time this year You're both correct, of course!
- atmosx 2mo agoWhat this guy said :point_up: My personal take-away is this: whatever you choose, make sure it's consistent across services (not serviceA behaves like X and serviceB like Y) and make sure eng teams know _how_ these are configured and what can go wrong. They'll figure out the rest.
- solatic 2mo agoYou'd be surprised how many engineering leaders don't understand the CAP theorem and will fail engineers on interviews for picking the one they don't agree with instead of communicating their expectations clearly (dodged a bullet on that one ...)
- jaggederest 2mo agoThat's a problem for circuitbreakers on these kinds of actions, not lying on health checks. Something like healthcheck fails -> restart -> healthcheck fails -> restart -> healthcheck fails -> circuit breaker trip, alarm raised, give up until manual intervention or X minutes have passed
- deathanatos 2mo agoThat circuitbreaker exists, by default. It is "CrashloopBackoff", here, and TFA covers it. (& it's an "until X minutes have passed" kind, by default.)
- dilyevsky 2mo agobackoff is only applied to individual pods/containers not across pods. the point is at scale it's easy to get into a situation where it's not possible to recover without (usually manual) full service drain
- jaggederest 2mo agoYeah that's why I have a manual intervention breaker that goes across all the pods/nodes etc. CrashLoopBackoff is great for selfhealing but when things go really pear shaped you want something that catches the global state. Saw it activate during an AWS outage one time where new nodes were unhealthy on start, for example.