4 ms·
This is why you have circuit breakers upstream. Not on every individual instance.
by llama052 1mo ago
This is why you have circuit breakers upstream. Not on every individual instance.
- jdm2212 1mo agoDoesn't do you any good if the outage is in the circuit breaking layer, which it was for GitHub (this started as a load balancer outage).
- llama052 1mo agoIdeally you have levers further up from your local load balancers as well. Even at the edge. Granted you never want those to trigger but it’s better than fighting a storm while you fix things.
- otterley 1mo agoI wondered that myself. Curious as to why they couldn’t shed load at the edge to help protect goodput.
- Maxion 1mo agoIsn't that what they did though? Start returning more-or-less hardcoded 403s for the Copilot endpoint that was causing the issues?
- otterley 1mo agoThat’s more surgical than load shedding. With load shedding you intentionally return 503s to a proportion of all legitimate requests. It turns a hard blackout (total outage) into a flakiness issue.
- unscaled 1mo agoIn highly distributed microservice architecture, there's almost never a single upstream. In some cases you may have a couple of customer-facing entry-points (a global API gateway, and a couple of BFFs), but these are not the only paths that need to be protected. There are client-side retries (which have broken GitHub in this case) and server-side initiated API calls between microservices that don't pass through any of your ingresses (e.g. triggered by an ETL pipeline, or a scheduled job). With a complex architecture you can't just slap a circuit breaker on a couple of ingresses and call it a day. Don't get me wrong, putting them there does go a long way, but you won't be covering all your bases.
- llama052 1mo agoI completely agree, I currently manage a fleet of microservices that handles a few trillion requests a month. It’s about defense in layers to these sorts of things. All the way through the stack if possible starting at the edge. At least it should be required for critical level services in production.
- deleted 1mo ago[deleted]
- otterley 1mo agoClose but not quite: https://builder.aws.com/content/3EumjoZascWd1oZiEgL8ORlv3qE/timeouts-retries-and-backoff-with-jitter https://builder.aws.com/content/3EumjoZascWd1oZiEgL8ORlv3qE/... “Circuit breakers, where calls to a downstream service are stopped entirely when an error threshold is exceeded, are widely promoted to solve this problem. Unfortunately, circuit breakers introduce modal behavior into systems that can be difficult to test, and can introduce significant additional time to recovery. We have found that we can mitigate this risk by limiting retries locally using a token bucket. This allows all calls to retry as long as there are tokens, and then retry at a fixed rate when the tokens are exhausted.”