4 ms·
IDK why mobile dev folks do blind retries thinking server is some kind of mage. I’ve had debates on how much retry makes sense for a login service I built. Peop
by maxpert 6y ago
IDK why mobile dev folks do blind retries thinking server is some kind of mage. I’ve had debates on how much retry makes sense for a login service I built. People have the tendency of hey this endpoint used to work with these parameters didn’t work? Fine I will keep retrying! The only difference here I think is in company paid job you step up immediately put in a server side fix to fix outage and then make these changes to roll them out. All of this is achieved in matter of hour or so rather than 12+ hours.
Circuit breaking, exponential backoff, bulkheading should be key pillars and no dev should be allowed to write a client without these key skills.
Edit: Please add Jitter to the list of items. Any good library like https://resilience4j.readme.io/ https://resilience4j.readme.io/ will give you all of this for free out of box!
- tyingq 6y agoHas the "batteries included" aspect of things like exponential backoff improved? Last I was in the details, most of the popular client libraries offered little, and you had to implement it yourself.
- maxpert 6y agoResilience4J has improved quite a lot. I have been using for over two years now, never missed out on anything.
- abhi_kr 6y agoAnother scenario to keep in mind is the Thundering Herd Problem[0]. Exponential backoff without added jitter could still DDoS the servers. [0]: https://en.wikipedia.org/wiki/Thundering_herd_problem#Mitigation https://en.wikipedia.org/wiki/Thundering_herd_problem#Mitiga...
- VWWHFSfQ 6y agoexponential backoff will still kill the servers because the first thing people do is kill/restart the app or reload the webpage and it will just restart the backoff again
- anigbrowl 6y agoSo just set 'backoff until [timestamp]' rather than 'backoff for [time interval]'. Generally users restart because they don't know what's going an assume the client is stuck in loop or something. Think of a client that can say 'internet is up but my.server is having [local, regional, global] problems. I will try again at 11:37am EST.' Downforeveryoneorjustme is exploring an API for service monitoring and it seems to me like every cloud client should have a standardized approach of trying to reach its own server, then checking internet access, then checking a service monitor, then checking social media status updates. https://downforeveryoneorjustme.com/services/api https://downforeveryoneorjustme.com/services/api Also think apps and devices should have limited peer-to-peer information sharing instead of only talking to the operating system. In many ways our devices are like a roomful of people whispering status updates to the operating system and/or user but never talking to each other because 'security'.
- edoceo 6y agoI make a mobile app and we do this. When the app wakes up it tries to GET a file from /.well-known and if that works, proceed. If it fails we have a notice with "will retry at T". Timing is random between 10 and 50 seconds
- anonymousab 6y agoYes, but even they much back pressure provides a lot of reprieve. Restarting an app or even manually refreshing a web page takes a lot longer than an in-memory automatic retry() function call.
- orojackson 6y agoWhich form of jitter is better: adding a random wait time to a predetermined wait time that grows exponentially with each retry attempt, or following something like [0] where every retry attempt increases the possible wait time choices and the "jitter" is to randomly pick one of them? To illustrate the latter option, suppose the smallest retry time unit is 1 second. The first attempt gives you a random choice in {0, 1}. The second attempt gives you a random choice in {0, 1, 2}. The third attempt gives you a random choice in {0, 1, 2, 4}. The fourth attempt gives you a random choice in {0, 1, 2, 4, 8}. This goes on until a ceiling in the number of attempts or a set wall clock time is reached. [0]: https://en.wikipedia.org/wiki/Exponential_backoff#Example_exponential_backoff_algorithm https://en.wikipedia.org/wiki/Exponential_backoff#Example_ex...
- toast0 6y agoIt depends on what you want. The first option will tend towards longer waits, and the second towards shorter waits. I would tend towards increasing the minimum wait at each iteration (until you get to some maximim wait), because it it failed 10 times in the last minute, it's likely going to fail many times in the next minute, so we don't need to try more than once or twice. Also: in case the client random is broken, you don't want to accidentaly end up with everyone retrying after zero seconds forever.
- yyhhsj0521 6y agoEthernet uses the second one to avoid collision.
- saati 6y agoIs anyone still using shared medium Ethernet?
- baby 6y agoOr back pressure. How does bulkhead helps here?
- treeman79 6y agoDozen Years ago I built a backend Ruby server, (not rails), I needed max performance. Our front end flex guy didn’t understand backing off on retries. Some system had an issue, that caused the flex app for 1000 people to go insane. Each calling ruby app endlessly with no delay. Ruby app remained healthy, server crashed in under 5 minutes from running out of disk space. Our alarm interval on free space ran every 5 minutes. So no time for an alarm to sound. A mix, of I was very proud of my little ruby app for scaling. And a a wtf, I needed to be more involved with front end people.
- zrm 6y agoI've seen it regularly happen where a user's company account is used for email that whenever the user changes their password, the email app on their phone causes their account to get locked out by doing repeated retries with the old password.
- maxpert 6y agoMind sharing the app name?
- AareyBaba 6y agoMicrosoft Outlook does this.
- tatersolid 6y agoNative Mail client for Android still retries aggressively on an incorrect password, locking out accounts. The native iOS mail client stopped the endless retry behavior a few releases ago (so sad it only took a decade to fix),
- vbezhenar 6y agoThat kind of design is wrong because anyone can lock out any account (or all of them).
- adkadskhj 6y agoI'm unfamiliar with "circuit breaking" and "bulkheading" in this context. Could you provide a bit more description so i can research them and make sure i know what i should know? :)
- cnasc 6y agoCircuit breaking: also called “kill switch”. It’s having a way to shut down some feature if it becomes problematic. “Bulkheading” is making sure that failures in one area don’t cascade into causing failures elsewhere analogous to how bulkheads on a ship prevent a single breach from sinking the whole ship.
- ignoramous 6y agoAs far as I know, the original source for these patterns is Michael Nygard's book Release It: https://www.amazon.com/gp/product/0978739213 https://www.amazon.com/gp/product/0978739213
- sciurus 6y agoA fantastic book! The first edition is a bit dated in places, but there's a second edition now: https://www.amazon.com/Release-Design-Deploy-Production-Ready-Software-dp-1680502395/dp/1680502395/ https://www.amazon.com/Release-Design-Deploy-Production-Read...
- dtech 6y agoThe difference between a circuit breaker and a kill switch is that a circuit breaker, like a real electrical one, automatically trips and stops requests going through for some time after a certain error threshold to the remote server has passes.
- hayst4ck 6y agoThe context is somewhat important. While 508s were being sent, there was also a significant number of 503s (IIRC). A 503 is an absolute hallmark of something reaching max utilization. Sometimes it's a bad code push that results in memory bloat and then swap or significantly increased request handling time on a poorly threaded server (think a code push that makes a blocking linear request in a loop), but the vast majority of time an upstream dependency (specifically a data store) has been overloaded. So for whatever reason a data store gets slower. What happens upstream when this happens? The number of incoming requests is constant, but the time each individual thread spends attempting to talk to the data store is constant (or worsening). This means each request blocks longer (resulting in potential thread starvation) or there are more concurrent requests (load) to the data store. This creates a feedback loop of doom: as a data store slows down, its load (the number of requests it's handling at once) increases until complete failure. The only way to stop this behavior is by "failing fast." This is how a circuit breaker works. When the data store starts responding slowly, it’s important not to hammer it with even more load, so your client watches the number of load related failures (or response times) and automatically fails requests immediately without sending it to the data store (circuit breaker). This allows the data store to become unloaded and get out of the doomed feedback loop.
- wtmt 6y ago> Circuit breaking, exponential backoff, bulkheading should be key pillars and no dev should be allowed to write a client without these key skills. This is all the more critical for client applications that depend on app stores to review and approve updates and end users who may not immediately update to a newer version because they don’t have auto update turned on. Edit: I can’t seem to find it, but there was a comment recently about servers using flags to tell clients to shut up for a specific duration and avoid exacerbating problems with retries. I think it was about some specific design of Dropbox.
- eli 6y agoThere’s a whole HTTP status for “you are connecting too fast” but that assumes the server is kinda working. https://developer.mozilla.org/en-US/docs/Web/HTTP/Status/429 https://developer.mozilla.org/en-US/docs/Web/HTTP/Status/429
- mey 6y agoMost client code I have ever seen unfortunately only ever looks for status code != 200
- ju-st 6y agoSignal started to handle 404 7 months ago and the 508 code 13 hours ago. (https://github.com/signalapp/Signal-Android/blame/master/libsignal/service/src/main/java/org/whispersystems/signalservice/api/SignalServiceMessagePipe.java https://github.com/signalapp/Signal-Android/blame/master/lib... strg+f 508)
- maxpert 6y agoAgreed server should have a rate-limiting on each account to prevent DDoSing in future.
- amluto 6y agoIn an emergency, one could rig up a server to drop TCP connections without sending RST. The result would be a client that thinks it’s connected without using server resources. Bonus points for some code on the server that responds to some fraction of incoming segments with ACKs to keep stringing clients along.
- vaduz 6y agoFirst distrbuted computing fallacy is "Network is reliable". Detecting if you can deliver a packet in advance is not exactly a solved problem - network detection APIs at best inform you if the PHY/MAC layer for a given connection is up, not if any given packet you generate can be routed out - and fail to detect a number of rather common edge cases (all traffic being routed over VPN on WiFi, without mobile fallback, for instance, or mobile connection appearing to be up, but not transferring data) - therefore ultimately the solution is to try, and try again, with proper backoff to prevent overwhelming your own infrastructure. In this case the latter appears to have been missing.
- maxpert 6y agoI've seen "Network is unreliable" being used as an argument to hide laziness. I do agree you can have flaky network, but the biggest give of the laziness behind the argument is not inspecting error codes at all. In this particular case 508 for example is clear indicator of some error status. Why would you even keep blind retries in place? I hate the argument of devs arguing for "dumb clients" and server "taking care of everything". Even for network failures IDK how having a retry within millisecond will actually help. It will be more detrimental to your infra, bandwidth, battery than a once or twice snappy experience argument.
- theossuary 6y agoSure but that just falls under proper backoff. If they decide to use different backoff strategies for different classes if errors, that'd be great. But it's an easy mistake to make, and a very hard one to catch in testing. Sure if they wrote the code better (or been less "lazy") this wouldn't have happened, but that type of feedback isn't constructive imo. This wasn't a conscious decision that was made, it was written this way by default because of some combination of their team/framework/api client/testing/documentation/etc. Talking about why those led to this is much more interesting I think.
- aftbit 6y agoSmart server dumb client makes a lot of sense in the web-dev world, where the core client code is not under your control. As soon as you need to write your own app (Android/iOS/desktop), you should revisit this decision. The server should be able to serve a status code & header that requests the client to back off for a certain time period or with certain exponential parameters. This is a very hard thing to deploy after the fact and can really save your bacon, especially if you have a number of very old clients still out in the field refusing auto-updates. Sure, you can refuse to serve them, but if they get stuck in a tight loop retrying until you do, you're pretty hosed.
- husainalshehhi 6y agoAlso token bucket: https://en.wikipedia.org/wiki/Token_bucket https://en.wikipedia.org/wiki/Token_bucket
- champtar 6y agoI remember some years ago iPhone Activesync client was retrying 5 time the same login password on authentication failure (http 401), and our AD was locking people out after 5 attempts ... We ended up writing a small proxy to fix it server side.
- wh33zle 6y agoYou'd think the semantics of "4xx is the client's fault, don't bother retrying without changing the request" is pretty easy to understand.
- rglullis 6y agoTake the flip side of it. If you are working on the backend, you should always assume/provision for the case where there will be hostile clients. If your backend can only work if the client plays nice, I'd be weary of pointing fingers too quickly.
- als0 6y agoTrouble is...you can't tell who is malicious in a DDoS scenario.
- rglullis 6y agoYeah, that's my point actually. If your service is failing because of a DDOS, don't blame those trying to use your infra. Before pointing to issues of bad clients, Signal should be ready for this influx of users - especially if your leader believes that centralized services are the only ones able to create a viable alternative for the masses.
- wadkar 6y agoit's also quite possible to design and implement this one the server side. I mean one can implement Circuit breaking, exponential backoff etc. at network level (e.g. istio) based on the "internal service" (could be VMs/nginx rever proxies) outage error.
- amerine 6y agoOn man. Your list of pillars is spot on.