4 ms·
I'd be interested to hear other strategies in this space. I've done the naive thing of allowing retries everywhere, and gotten into retry storms. When I was nex
by aftbit 18d ago
I'd be interested to hear other strategies in this space. I've done the naive thing of allowing retries everywhere, and gotten into retry storms. When I was next presented with the problem, I tried the other naive thing of only allowing retries from the very top level service, which led me to redoing absolutely tons of work for each failure. What's a nice middle path that doesn't add too much complexity?
- tregoning 17d agohttps://en.wikipedia.org/wiki/Exponential_backoff https://en.wikipedia.org/wiki/Exponential_backoff
- anonymars 17d agohttps://devblogs.microsoft.com/oldnewthing/20051107-20/?p=33433 https://devblogs.microsoft.com/oldnewthing/20051107-20/?p=33...
- applfanboysbgon 17d agoThis is trading a good developer experience for a bad user experience. There are situations where it makes sense to force manual retry, but there's no reason to apply one universal rule to all possible situations. Lack of considering nuance for your situation is just intellectual laziness.
- anonymars 17d agoMy point was that just throwing exponential backoffs at the retry problem is not a magic solution I don't follow how being cautious about avoiding multiplicative layers of backoffs is trading a good developer experience for a bad user experience. The described situation is an awful user experience. Simply adding a retry and calling it a day sounds like the easy developer experience at the expense of the user experience
- applfanboysbgon 17d ago> The described situation is an awful user experience Sure, the worst case scenario is. 99.9999% of the time, a transient error will actually just work on the first or second auto-retry and save your users the effort of paying attention and manually retrying things. This is especially prudent for background tasks where the failure may not be noticed right away; coming back to something fire-and-forget 30m later to see it never tried to finish is not a good user experience. > My point was that just throwing exponential backoffs at the retry problem is not a magic solution Nobody said it was. In fact, I suggested the exact opposite - a proper solution takes dev effort. Adhering to an iron rule of "just make them manually retry" is throwing your hands up and not even trying to solve the problem because laziness is convenient.
- anonymars 17d ago> Nobody said it was I responded to a post that merely linked to the Wikipedia article for exponential backoff (in response to "I'd be interested to hear other strategies in [protecting against retry storms]") The submitted article is precisely about the degenerate case and the difficult work of dealing with it > This approach works for transient or low-rate failures. However, during moderate or severe degradation, it becomes counterproductive. Aggressively retrying against an already struggling service increases load, accelerates failure, and amplifies retry traffic across upstream dependencies. What begins as a localized outage can quickly escalate into a stack-wide incident—ultimately degrading, or in the worst case, completely breaking, the end user experience.
- applfanboysbgon 17d agoYes, and then responds to that degenerate case by suggesting that you never automatically retry. It's like saying "you should never drive a car/take a flight/ride public transportation because it's gone wrong so many times". Things go wrong. You should absolutely consider the impact and what will happen when they go wrong, but the end takeaway to just never engage with them because they can go wrong is, frankly speaking, lazy and bad advice.
- mitxela 17d agoTitle is "Take it easy on the automatic retries" Microsoft breaks all Old New Thing links every few years so it's necessary to post the title so the right post can still be found.
- sroussey 17d agoYes, exponential back off and jitter are the first things to work on, and good if you don’t have a better signal (like loss of network). Also, a simple signal status server or queue system helps to keep global state such that everyone doesn’t retry all at once. If you have a central error rate server you can skip your retry based on the error rate (100% error rate, don’t retry, etc).
- sroussey 17d agoSo many variables, but the simple thing is to set things up like normal rate limiting (which you would want to do anyways). The one generating the errors passes back a retry time. You can add jitter here, tell low priority requests to wait longer, etc. BTW: do keep track of priority. It’s like having a database that gets flooded with connections and won’t allow new ones in—but will for admin users (btw, it did not used to be that way in the early days of MySQL).
- otterley 17d agohttps://aws.amazon.com/blogs/developer/introducing-retry-throttling/ https://aws.amazon.com/blogs/developer/introducing-retry-thr... (2016 -- 10 years ago!) https://builder.aws.com/content/3EumjoZascWd1oZiEgL8ORlv3qE/timeouts-retries-and-backoff-with-jitter https://builder.aws.com/content/3EumjoZascWd1oZiEgL8ORlv3qE/... (originally published 2020, republished 2026) https://docs.aws.amazon.com/sdkref/latest/guide/feature-retry-behavior.html#retry-quota-token-bucket https://docs.aws.amazon.com/sdkref/latest/guide/feature-retr... https://aws.amazon.com/blogs/developer/announcing-updated-retry-behavior-for-aws-sdks-and-tools/ https://aws.amazon.com/blogs/developer/announcing-updated-re... (2026)
- sfraxo 17d ago[flagged]
- CBLT 17d agoThere's a good amount of literature about this (check the other comments), but you can vastly simplify this into two things you need to do: 1. Your service that retries should have some retry budget. This is a good place to be "smart", because you can reason entirely locally instead of turning it into a distributed systems problem. The best library I've seen for this was doing Exponential Moving Average of requests per second sent down that pipe (not counting retries) and only allowing 20% more requests per second as retries, total. Each individual request could be retried 3 times. This was critical as it bounds the additional load from retries. 2. Whenever a service retries but has to give up, the error it sends to its callers should never be retried. There has to be some agreement that that HTTP code will never be retried. This prevents the multiplicative factor of retry on top of retry, which is why those storms can generate so much load. Everything else is nice-to-have, but those two alone should bound the total requests you get in a retry storm.
- otterley 17d ago> 2. Whenever a service retries but has to give up, the error it sends to its callers should never be retried. There has to be some agreement that that HTTP code will never be retried. This prevents the multiplicative factor of retry on top of retry, which is why those storms can generate so much load. Ooh, I like the idea of propagating "no retries" hints in the responses back upstream. Have you seen it implemented in the wild, or in public discussions about the practice?
- CBLT 17d agoI've only seen it in bigcorp cross-service typedefs, or in startup's code that re-implements the checks in every service.
- hxtk 17d agoIn gRPC, the statuses it returns in trailers can include arbitrary details, and Google has a well-known proto for common ones in `google/rpc/error_details.proto`. One such detail is RetryInfo [1]. When we implement retries where I work, the general rule is that if a request is suitable for retry, it should include the RetryInfo in the error status and use it as the base delay for the exponential backoff. The absence of that detail means don’t retry, and we have a client interceptor that parses the response status and retries according to that logic. 1: https://github.com/googleapis/googleapis/blob/bba4c646b1f85a4cd499207926865205caad0384/google/rpc/error_details.proto#L92 https://github.com/googleapis/googleapis/blob/bba4c646b1f85a...
- mandevil 17d agoAt $previousJob we implemented circuit breakers: centralize all requests to the foreign service (every call to service theta went through the service theta client which had some shared state so everything so we could keep track of requests) and then monitor, when error % got above a certain limit start to dump requests to a text file for sending in the future instead of now. And the centralized caller will send one message every time gap (we started at 30s) and as long as that errors out we keep writing. We did that because otherwise we would get 2x30 second timeouts to a dead service on every user interaction and it made for a terrible user experience. Keeping track and handling it smartly made the average user experience a lot better.
- cyberax 17d agoOne good option that is not (yet?) mentioned here is a deadline for retries. You can cap the request duration by, say, 500ms and pass the remaining time budget to downstream services. This can be done via an HTTP header and enforced by the middleware.
- konaraddi 17d agoDepending on the context, circuit breakers
- nirmeetimthebes 17d ago[flagged]
- jeffbee 17d agoJust limiting your retry budget to 1% of normal rates using a client-local token bucket with no distributed coordination will eliminate the possibility of long-lived retry storms.
- sisuo30390 17d ago[flagged]