5 ms·
> Errors in those services triggered a client-side retry loop that increased traffic during recovery Symptomic of a wider trend to avoid showing the user any e
by cube00 2mo ago
> Errors in those services triggered a client-side retry loop that increased traffic during recovery
Symptomic of a wider trend to avoid showing the user any error at all costs, even if that means they sit watching a spinner for 7 hours.
> Delayed replies to a single internal endpoint triggered a latent retry bug in VS Code that amplified traffic by approximately 10x and caused delayed recovery for the Copilot Token Service.
The detailed root analysis tries to pass this off as a "bug". You can't seriously tell me client retry doesn't have a unit test which ensures the retry back off behaviour is functioning exactly as designed. In this case aggressively to try and hide problems if token service responses become flakey.
- XorNot 2mo agoExcept the other side of this is interrupting a service which would otherwise have succeeded: there's a lot of unattended or minimally attended processes where an interruption is just asking the user to do the only thing they were going to do anyway - retry it. In GitHub's case this is especially relevant - the only reason to throw an error message at the user is the hope they - the human - give up and walk away (or you break all the CI/CD builds and the time it takes humans to hit "retry" gives you some breathing room).
- hizyyo 2mo ago[flagged]
- sqquima 2mo agoMaybe the retry logic was vibecoded instead of using an existing hardened library. After all, according to Twitter, nobody is looking at the code anymore.
- skissane 2mo ago> instead of using an existing hardened library A lot of retry libraries I’ve seen require the user to configure them. You can use a library with all the right settings, but if you configure it wrong, you are really no better off than if you hadn’t
- jdm2212 2mo agoA common pattern in highly available services is that sometimes you should retry immediately (because the node you hit is rolling/broken/overloaded, but the others aren't) and other times you should back off aggressively (because the service is degraded). If your server indicates with 100% accuracy when to retry immediately vs backoff, AND if all your clients consume that information with 100% accuracy, things go great. But there are lots of situations where one or both of those breaks down.
- dapperdrake 2mo agoCAP theorem. Pick one of those.
- deleted 2mo ago[deleted]
- MrWiffles 2mo agoWe used to be able to afford two! Acronym letters cost as much as houses nowadays!
- 27183 2mo agoOK I'll pick C and A! https://web.archive.org/web/20250128235041/https://codahale.com/you-cant-sacrifice-partition-tolerance/ https://web.archive.org/web/20250128235041/https://codahale....
- ACCount37 2mo ago"You can't seriously tell me that the unhappy leg of the code path has no test coverage." Sometimes I forget how ignorant HN can be of real world software development and the bar of corporate code quality, and then bangers like this remind me of it.
- bbarn 2mo agoExactly this. I've seen production level trading systems grind to a halt over a simple bug and no matter what tests you have in place, it happens.
- lbrandy 2mo agoI cannot even begin to express how many times I've seen engineers working super hard to optimize happy-paths so that we turn 3 nines of availability into 4 nines but introduce unintended emergent behaviors in unhappy-paths that turn 1 nine into zero nines via thundering herds, retry storms, etc.
- Breakthrough 2mo agoYou have my empathy for this kind of sentiment. Personally this seems somewhat rare in practice. That being said I'm curious if anyone has anecdotes they can share about these kinds of things?
- aprdm 2mo agoConfiguring postgres to automatically failover instead of doing it manually. The automated system caused more downtime in a few months than manually doing it did for years before. All in the name of more automations and less downtime
- eudamoniac 2mo agoCould you elaborate on why that would cause more downtime?
- giancarlostoro 2mo agoBackend API rate limiting has to surely kick in and force you to wait x amount of time before you try again… Discords bot API actually sends you how long before you retry.
- chasd00 2mo agoAs a client you can hit a server as much and as often as you like. The only thing the server can do is return an error code or try to hold the socket open (which the client can then close on their own).
- inigyou 2mo agoIt can block your IP address at the firewall level. It can't stop you DDoSing it with raw packets, but that's very unlikely to happen unintentionally, because TCP will wait for a SYN-ACK response for at least several seconds and possibly up to several minutes.
- bluerooibos 2mo ago> You can't seriously tell me client retry doesn't have a unit test which ensures the retry back off behaviour That wouldn't be a unit test - that's more like an end-to-end or integration test. Have you ever worked anywhere that had perfect test coverage? It just doesn't happen, nor is it possible unless you're building a calculator app or todo list.
- hirvi74 2mo agoMy employer has 0% test coverage lol. I've begged and pleaded, but the claim is that "risk is low" and "that's what QA is for." Hell, I've complained to senior management about how there are senior devs that forgo backend validation. It's truly Hell in the trenches sometimes. Some days, I would seriously rather work at Wendy's.
- nnx 2mo ago> Some days, I would seriously rather work at Wendy's. Narrator: He would not.
- hirvi74 2mo agoIf such a job paid even half as much, I would quit today. I can continue to program computers as a hobby. Occupational programming has essentially eroded my passion over the years. Money is not important to me insofar as I have enough to live an average life. I don't need anymore than that.
- inigyou 2mo agoJust write the tests. It's your job, not your employer's.
- hirvi74 2mo agoIt's their code, not mine. If they don't want us writing tests, then so be it. I get paid either way.
- Lammy 2mo ago> a wider trend to avoid showing the user any error at all costs And in fact you can see the degradation of software over the previous decade-plus via Google Trends search for ‘something went wrong’ lol: https://trends.google.com/trends/explore?date=all&q=%22Something%20went%20wrong%22&hl=en https://trends.google.com/trends/explore?date=all&q=%22Somet...
- inigyou 2mo agoIronically I get an error page when I click on this.
- tclancy 2mo ago> You can't seriously tell me client retry doesn't have a unit test which ensures the retry back off behaviour Not to join the parade, but what would a unit test that confirms a cycling behavior across all the instances in-flight even look like? I mean, besides "Not a unit test".
- cube00 2mo agoMock out the network call (which you'd be doing anyway because unit tests never connect to the network) so it always returns a failure. Mock out the timer. Call the function multiple times and ensure it'd passing the expected wait durations in for each time it's called followed by a fatal error after say 30 seconds.
- tclancy 2mo ago> Mock out the network call Man, the day this hits, just know I didn't judge you at all.
- tverbeure 2mo agoThey added 3 million CPUs. You reduce the complexity of their systems to a unit test…
- _kidlike 2mo agoSo, in all of your software, you have introduced randomness in your retries so that the billions of your clients avoid retry synchronization dances?