4 ms·
The right answer is for the RPC framework to accurately communicate "try again on another node" vs "don't try again, just hard fail". When one end user request
by jdm2212 2mo ago
The right answer is for the RPC framework to accurately communicate "try again on another node" vs "don't try again, just hard fail".
When one end user request fans out to hundreds of backend requests (typical for microservices), you can't have each of those backend requests do its own exponential backoff. If they do it in parallel, they're a thundering herd, and if they do it in serial, the end user request will time out before you finish all the work, at which point you're doing a bunch of slow expensive work for no gain (and the enqueued slow expensive work will make your outage worse).
- llama052 2mo agoThis is why you have circuit breakers upstream. Not on every individual instance.
- jdm2212 2mo agoDoesn't do you any good if the outage is in the circuit breaking layer, which it was for GitHub (this started as a load balancer outage).
- llama052 2mo agoIdeally you have levers further up from your local load balancers as well. Even at the edge. Granted you never want those to trigger but it’s better than fighting a storm while you fix things.
- otterley 2mo agoI wondered that myself. Curious as to why they couldn’t shed load at the edge to help protect goodput.
- Maxion 2mo agoIsn't that what they did though? Start returning more-or-less hardcoded 403s for the Copilot endpoint that was causing the issues?
- otterley 2mo agoThat’s more surgical than load shedding. With load shedding you intentionally return 503s to a proportion of all legitimate requests. It turns a hard blackout (total outage) into a flakiness issue.
- unscaled 2mo agoIn highly distributed microservice architecture, there's almost never a single upstream. In some cases you may have a couple of customer-facing entry-points (a global API gateway, and a couple of BFFs), but these are not the only paths that need to be protected. There are client-side retries (which have broken GitHub in this case) and server-side initiated API calls between microservices that don't pass through any of your ingresses (e.g. triggered by an ETL pipeline, or a scheduled job). With a complex architecture you can't just slap a circuit breaker on a couple of ingresses and call it a day. Don't get me wrong, putting them there does go a long way, but you won't be covering all your bases.
- llama052 2mo agoI completely agree, I currently manage a fleet of microservices that handles a few trillion requests a month. It’s about defense in layers to these sorts of things. All the way through the stack if possible starting at the edge. At least it should be required for critical level services in production.
- deleted 2mo ago[deleted]
- otterley 2mo agoClose but not quite: https://builder.aws.com/content/3EumjoZascWd1oZiEgL8ORlv3qE/timeouts-retries-and-backoff-with-jitter https://builder.aws.com/content/3EumjoZascWd1oZiEgL8ORlv3qE/... “Circuit breakers, where calls to a downstream service are stopped entirely when an error threshold is exceeded, are widely promoted to solve this problem. Unfortunately, circuit breakers introduce modal behavior into systems that can be difficult to test, and can introduce significant additional time to recovery. We have found that we can mitigate this risk by limiting retries locally using a token bucket. This allows all calls to retry as long as there are tokens, and then retry at a fixed rate when the tokens are exhausted.”
- otterley 2mo agoThis is probably the best and most thorough explanation I’ve seen: https://builder.aws.com/content/3EumjoZascWd1oZiEgL8ORlv3qE/timeouts-retries-and-backoff-with-jitter https://builder.aws.com/content/3EumjoZascWd1oZiEgL8ORlv3qE/... I’m particularly fond of the token-bucket mechanism for pacing recovery.
- throwawayqqq11 2mo agoThats why you only back off at the edges, (not internally, where you can spin up on demand) and why you should add jitter to the timing.