3 ms·
> because of some flakey third-party service I've started building in auto retries to the CI scripts for this type of thing, at least once it annoys me to a ce
by gregmac 5y ago
> because of some flakey third-party service
I've started building in auto retries to the CI scripts for this type of thing, at least once it annoys me to a certain extent. The most glaring and unavoidable ones that come to mind across several projects are external certificate time stamping services.
I generally hate this type of thing: retry mechanisms are a lazy band-aid that can mask real problems. But at the same time, it just isn't possible to get to 100% and retry is a better band-aid than constant human intervention.
- runeks 5y ago> I generally hate this type of thing: retry mechanisms are a lazy band-aid that can mask real problems. I disagree. Retry mechanisms is mandatory for any network request. You cannot expect a message to travel across the globe on a wire without ever failing. I would say the opposite: you must have a retry mechanism for all requests performed by your application; unless it’s acceptable to pass on this failure to the client (who will retry the request).
- gregmac 5y ago> I disagree. Retry mechanisms is mandatory for any network request. You cannot expect a message to travel across the globe on a wire without ever failing. Actually good point, and I agree. I didn't state that well. What I was thinking of was some other things in builds I've seen retries be the band-aid for: flaky, timing-based unit tests; locked files/directories; or applying database migrations. My usual reaction to adding a retry to a build step is "no, let's fix the actual problem". Network requests to external services are a special case though, and absolutely should always have retries.
- fivea 5y ago> I generally hate this type of thing: retry mechanisms are a lazy band-aid that can mask real problems. They can, but beyond the old "it broke once it deployed" obvious pattern, 9 out of 10 times retries only sidestep basic nuisances from transient errors such as a request timing out.