3 ms·
Why? You can retry, but there is nothing wrong with increasingly waiting slightly longer if we fail many times.
by grim_io 2mo ago
Why? You can retry, but there is nothing wrong with increasingly waiting slightly longer if we fail many times.
- jdm2212 2mo agoTry asking Opus or Fable that question. It'll give you a good answer on why microservice architectures work the way they do in order to keep user-facing latency acceptable and minimize downtime. It's a complicated enough topic that I don't feel like explaining it for free to you in a HN comment.
- ctvo 2mo agoWhy are you pretending this is a harder topic than it is and you have some secret insight. And you're a dick on top of it: Yes, exponential backoffs alone are insufficient. Yes, adding jitter helps randomize the calls across a fleet and should be the default with exponential backoffs. Yes, both of these may be sufficient for most systems. Yes, you can dive more into circuit breakers and adaptive retries to limit thundering herd. https://aws.amazon.com/blogs/architecture/exponential-backoff-and-jitter/ https://aws.amazon.com/blogs/architecture/exponential-backof... https://brooker.co.za/blog/2022/02/28/retries.html https://brooker.co.za/blog/2022/02/28/retries.html
- Dylan16807 2mo ago"exponential" and "slightly longer" are very different backoff patterns.
- grim_io 2mo agoThe typical ceiling for these is around a minute.
- 27183 2mo agoThis can either be a relatively short time or an eternity, depending on the upstream consequences. If that means holding a connection open for 60000 milliseconds while waiting for some downstream rpc to go through its backoff ritual, that's an eternity, and under load that connection pool will get exhausted quickly. So now a problem which should only affect maybe 1% of users has completely hosed everyone. I've seen this happen multiple times. Someone designs some clever backoff strategy without considering how it fits in the context of the rest of the system. Hilarity ensues. If you have something taking an entire minute on a computer, please ensure you implement it in such a way that no connections are actually held open for that entire minute.
- grim_io 2mo agoWe are probably talking past each other. I'm not proposing doing any exponential backoff retries, or even retries at all for internal services. In my mind, the retries with exponential backoff and jitter belong only on the end-client(VSCode in this case). Everything else -> fail fast.
- 27183 2mo agoWith http you can do it elegantly with a 429 and retry-after. 100% doing it at the edge is the way to go. The machinery which determines how long to delay a client's retry can benefit from knowledge of internal services' state, but I agree the ultimate decision must lie with the serving layer. That's the only way to efficiently deal with misbehaving clients. On that note, one of the more memorable incidents of my career was when a 10M+ node client decided to retry as hard as possible on 4xx. That was fun x_x. [edit] that is to say, for this mechanism to be robust your retry-after enforcement mechanism needs to be capable of withstanding almost every single one of your users attempting to illegally retry as fast as they physically can without negatively impacting that one user requesting legitimate traffic. https://media.tenor.com/p3mss3YI6TcAAAAM/wat.gif https://media.tenor.com/p3mss3YI6TcAAAAM/wat.gif
- Dylan16807 2mo agoWhich is probably hundreds of times your normal wait.