5 ms·
Are retries bad? These are the sort of reason they make me generally uncomfortable. I appreciate they might be useful in scenarios where connectivity is inher
by Quarrelsome 1mo ago
Are retries bad?
These are the sort of reason they make me generally uncomfortable. I appreciate they might be useful in scenarios where connectivity is inherently problematic (e.g. mobile connectivity), but for a super connected and very desktoppy service I'd rather not retry much, if at all. As it obscures it when stuff has genuinely gone wrong, and this worst case scenario is tragic.
I feel like I'm mildly stupid in trying to out retries as heresy but I'm not sure.
- dapperdrake 1mo agoIt seems like retries are sometimes best left to the human being in front of the screen. Works well enough.
- madeofpalk 1mo agoI was using claude tethered via my phone, and would lose signal every now and then as we went through a tunnel. I was glad for how resilient it was its its eventual retries.
- frollogaston 1mo agoI don't like blind retries. It's different if the server or LB knows it's overloaded and asks clients to retry in X seconds.
- frollogaston 1mo agoOh and this is already assuming the blind retries are randomized exponential backoff. Thought it went without saying but maybe not.
- GeorgeDewar 1mo agoI totally agree with you, I think retries are overused, with the exception of operations that are known to be unreliable and can't be improved. In my experience, errors which go away within a few seconds are quite rare, and are mainly due to flaws which are usually caught in testing. I think a very careful cost/risk/benefit analysis should be done when adding automatic retries to things. As well as potentially causing cascading failures, it is a degraded user experience when it doesn't succeed. As a user I would rather see an error straight away than see many seconds of spinning while something silently retries, and THEN an error.
- wat10000 1mo agoI have the exact opposite view. Way too often, I’ll be presented with an error to the effect of, “something went wrong, please try again” and often the retry works. And I’m left wondering why this machine whose sole purpose is to automate things can’t do that for me automatically. In particular, networks tend to be a LOT less reliable than the typical developer accounts for. And the failures are very often transient. A case I run into often is doing something with my phone while leaving the house. There’s a window where it still thinks it’s on the WiFi but it’s too far away for it to work anymore. Initiating an action in that window often produces an alert telling me to try again, and trying again a few seconds later almost always works.
- frollogaston 1mo agoThe Github outage was about internal clients. Phone apps are a reasonable place to say things are known to be unreliable and can't be fixed. Your IP address changes when you leave the house. Btw, PWAs added offline capabilities to websites. I hate how the only thing that got used for was these stupid pages that look like you were able to reach the site but it's actually just saying you have no internet, like YouTube.
- wat10000 1mo agoI’m sure some retries are helpful there too (TCP is doing them, at the very least) but yeah, different approaches for different situations. Maybe you retry but you don’t spend many seconds hoping for it to work.
- frollogaston 1mo agoI'm ok with TCP retries generally. But the assumption at L4 is that L3 can and will drop/reorder packets at random, and retrying is cheap. Also a lottt of tuning has gone into TCP already.
- inigyou 1mo agoIf you want a retry to prevent that, it must only be done at the outermost layer. Note that what you think is the outermost layer might not be, and that accidentally deploying a retry at a non-outermost layer is much worse than no retry at all.
- taylor-s 1mo agoRetries are good, conditional on having a client-side circuit breaker that stops retries quickly when nothing is working. Otherwise, they are good in good times and bad in bad times.
- Quarrelsome 1mo agoisn't that a bomb with a pair of scissors to cut the fuse that could break down under certain conditions? I feel like they could also hide an issue that might get fixed if there were no retries. Is it slow or is our resource sporadically offline?
- stingraycharles 1mo agoGoogle “thundering herd” and you’ll understand why uncontrolled retries can be / are bad.
- inigyou 1mo agoQuarrelsome is suggesting you should have no retries, not unlimited.
- stingraycharles 1mo agoYes I’m aware, and they’re probably suggesting that something on a higher level should instead retry, and I’m arguing that you need to coordinate the retry mechanisms and behavior between differently layers of your stack.
- MeetingsBrowser 1mo agoExponential back off retries could hide a real issue, but I would estimate something like 99.99999999% of network retries are resolved within the first 2 attempts. Not using retries is optimizing for the astronomically rare case, which is better mitigated by other means
- stingraycharles 1mo ago
- Olreich 1mo agohttps://brooker.co.za/blog/2022/02/28/retries.html https://brooker.co.za/blog/2022/02/28/retries.html Seems that retries are good when the error is rare, and bad when the error is common. Typically outages have you transitioning from "everything is fine" to "nothing works", so being able detect that transition early is helpful
- chrisjj 1mo ago> Seems that retries are good when the error is rare, and bad when the error is common. Retries are a great way to turn errors rare into common.
- deleted 1mo ago[deleted]
- anonymars 1mo agoYou're not wrong - https://devblogs.microsoft.com/oldnewthing/20051107-20/?p=33433 https://devblogs.microsoft.com/oldnewthing/20051107-20/?p=33...
- chrisjj 1mo ago> Are retries bad? For sh*ty providers they are great. Best of all when backsourced to the user by "Try again later." > As it obscures it when stuff has genuinely gone wrong Works as designed - at every level.