6 ms·
Exponential backoff is the wrong answer in a highly available system in the typical case where (a) failure is expected and (b) you have nodes you are supposed t
by jdm2212 2mo ago
Exponential backoff is the wrong answer in a highly available system in the typical case where (a) failure is expected and (b) you have nodes you are supposed to fail over to.
- grim_io 2mo agoWhy? You can retry, but there is nothing wrong with increasingly waiting slightly longer if we fail many times.
- jdm2212 2mo agoTry asking Opus or Fable that question. It'll give you a good answer on why microservice architectures work the way they do in order to keep user-facing latency acceptable and minimize downtime. It's a complicated enough topic that I don't feel like explaining it for free to you in a HN comment.
- ctvo 2mo agoWhy are you pretending this is a harder topic than it is and you have some secret insight. And you're a dick on top of it: Yes, exponential backoffs alone are insufficient. Yes, adding jitter helps randomize the calls across a fleet and should be the default with exponential backoffs. Yes, both of these may be sufficient for most systems. Yes, you can dive more into circuit breakers and adaptive retries to limit thundering herd. https://aws.amazon.com/blogs/architecture/exponential-backoff-and-jitter/ https://aws.amazon.com/blogs/architecture/exponential-backof... https://brooker.co.za/blog/2022/02/28/retries.html https://brooker.co.za/blog/2022/02/28/retries.html
- Dylan16807 2mo ago"exponential" and "slightly longer" are very different backoff patterns.
- grim_io 2mo agoThe typical ceiling for these is around a minute.
- 27183 2mo agoThis can either be a relatively short time or an eternity, depending on the upstream consequences. If that means holding a connection open for 60000 milliseconds while waiting for some downstream rpc to go through its backoff ritual, that's an eternity, and under load that connection pool will get exhausted quickly. So now a problem which should only affect maybe 1% of users has completely hosed everyone. I've seen this happen multiple times. Someone designs some clever backoff strategy without considering how it fits in the context of the rest of the system. Hilarity ensues. If you have something taking an entire minute on a computer, please ensure you implement it in such a way that no connections are actually held open for that entire minute.
- grim_io 2mo agoWe are probably talking past each other. I'm not proposing doing any exponential backoff retries, or even retries at all for internal services. In my mind, the retries with exponential backoff and jitter belong only on the end-client(VSCode in this case). Everything else -> fail fast.
- 27183 2mo agoWith http you can do it elegantly with a 429 and retry-after. 100% doing it at the edge is the way to go. The machinery which determines how long to delay a client's retry can benefit from knowledge of internal services' state, but I agree the ultimate decision must lie with the serving layer. That's the only way to efficiently deal with misbehaving clients. On that note, one of the more memorable incidents of my career was when a 10M+ node client decided to retry as hard as possible on 4xx. That was fun x_x. [edit] that is to say, for this mechanism to be robust your retry-after enforcement mechanism needs to be capable of withstanding almost every single one of your users attempting to illegally retry as fast as they physically can without negatively impacting that one user requesting legitimate traffic. https://media.tenor.com/p3mss3YI6TcAAAAM/wat.gif https://media.tenor.com/p3mss3YI6TcAAAAM/wat.gif
- wat10000 2mo agoThe whole point of exponential backoff is that the first retry can be quick.
- jdm2212 2mo agoThe right answer is for the RPC framework to accurately communicate "try again on another node" vs "don't try again, just hard fail". When one end user request fans out to hundreds of backend requests (typical for microservices), you can't have each of those backend requests do its own exponential backoff. If they do it in parallel, they're a thundering herd, and if they do it in serial, the end user request will time out before you finish all the work, at which point you're doing a bunch of slow expensive work for no gain (and the enqueued slow expensive work will make your outage worse).
- llama052 2mo agoThis is why you have circuit breakers upstream. Not on every individual instance.
- jdm2212 2mo agoDoesn't do you any good if the outage is in the circuit breaking layer, which it was for GitHub (this started as a load balancer outage).
- llama052 2mo agoIdeally you have levers further up from your local load balancers as well. Even at the edge. Granted you never want those to trigger but it’s better than fighting a storm while you fix things.
- otterley 2mo agoI wondered that myself. Curious as to why they couldn’t shed load at the edge to help protect goodput.
- dannyw 2mo agoYour highly available system is probably somewhat important, otherwise you won’t have invested in making it HA. While your premise holds for happy cases, when you do have a cascading series of outages, not using exponential backoff is just adding a self-inflicted DoS to when you do go down. I don’t really follow your premise and can’t really articulate many cases for when you shouldn’t use exponential backoff. Maybe if you’re working at Jane St or something; or other circumstances where you can deploy immediate changes to the client; and you’re willing to trade ‘better p50 for worse outages’. But in the case of shipped code that’s run on clients, I’ll continue exponentially backing off all the way, all the time, for everything.
- jdm2212 2mo agoWhen you have an outage, you should not retry at all. Exponential backoff is exactly how you get cascading outages. If service A fails a request to service B and decides to exponentially back off, now service A is holding open an end user request that will claim resources on service A. Fast forward ten minutes and the service B degradation has metastasized into a service A degradation. And even after service B has recovered, service A might still be dead. To handle this correctly you need your RPC framework to accurately communicate retryable vs non-retryable failures to clients. Then service A knows service B is dead, does not retry, and proapgates the failure to clients. This is hard to do perfectly, but there's no alternative that works.
- andrekandre 2mo ago> To handle this correctly you need your RPC framework to accurately communicate retryable vs non-retryable failures to clients. basically enumerate your errors, and depending on the type, retry or just return/forward that same "dont retry this" error?
- jdm2212 2mo agoPretty much, but ideally it should be transparent to your app developers. App developer writes `rpc.doThing(...)` on one side, and an implementation on the other, and the infrastructure -- the RPC framework or the service mesh or whatever -- transparently handles when/whether to retry and where to route retries.
- inigyou 2mo agoSomeone linked the Google SRE book. It explains that client services watch the global distribution (for that client) of number or retries. It will normally retry immediately on another node, because as you say, some failures are expected. But if it notices that more than about 1% of requests are having to retry twice, that indicates the server service needs reduced load and it starts refusing to retry. If it gets really bad it even starts preemptively failing first attempts so that the service doesn't get contacted at all.