3 ms·
A common pattern in highly available services is that sometimes you should retry immediately (because the node you hit is rolling/broken/overloaded, but the oth
by jdm2212 2mo ago
A common pattern in highly available services is that sometimes you should retry immediately (because the node you hit is rolling/broken/overloaded, but the others aren't) and other times you should back off aggressively (because the service is degraded).
If your server indicates with 100% accuracy when to retry immediately vs backoff, AND if all your clients consume that information with 100% accuracy, things go great. But there are lots of situations where one or both of those breaks down.
- dapperdrake 2mo agoCAP theorem. Pick one of those.
- deleted 2mo ago[deleted]
- MrWiffles 2mo agoWe used to be able to afford two! Acronym letters cost as much as houses nowadays!
- jordanb 2mo agoAnyone who designs such a system should know to use an exponential backoff to avoid the thundering herd. Maybe copilot missed that while it was reviewing its own PR
- jdm2212 2mo agoExponential backoff is the wrong answer in a highly available system in the typical case where (a) failure is expected and (b) you have nodes you are supposed to fail over to.
- grim_io 2mo agoWhy? You can retry, but there is nothing wrong with increasingly waiting slightly longer if we fail many times.
- jdm2212 2mo agoTry asking Opus or Fable that question. It'll give you a good answer on why microservice architectures work the way they do in order to keep user-facing latency acceptable and minimize downtime. It's a complicated enough topic that I don't feel like explaining it for free to you in a HN comment.
- ctvo 2mo agoWhy are you pretending this is a harder topic than it is and you have some secret insight. And you're a dick on top of it: Yes, exponential backoffs alone are insufficient. Yes, adding jitter helps randomize the calls across a fleet and should be the default with exponential backoffs. Yes, both of these may be sufficient for most systems. Yes, you can dive more into circuit breakers and adaptive retries to limit thundering herd. https://aws.amazon.com/blogs/architecture/exponential-backoff-and-jitter/ https://aws.amazon.com/blogs/architecture/exponential-backof... https://brooker.co.za/blog/2022/02/28/retries.html https://brooker.co.za/blog/2022/02/28/retries.html
- Dylan16807 2mo ago
- vlovich123 2mo ago> because the node you hit is rolling/broken/overloaded, but the others aren't Retries in such a situation should be handled internally with the client at most responsible for failing over with a circuit breaker to another zone. Having the client auto retry right away is not something that behaves well as shown here, even if in the happy path it happens to stimulate increased availability without actually investing in the proper architecture for it