7 ms·
> Extremely strict RPC settings. I’m talking zero retries (or MAYBE one) [...] I disagree. If we're talking about distributed systems, then one thing is guaran
by rytis 4y ago
> Extremely strict RPC settings. I’m talking zero retries (or MAYBE one) [...]
I disagree. If we're talking about distributed systems, then one thing is guaranteed - network is not going to be reliable. And if we have 10's or 100's of services, this policy means that at the smallest blip the whole thing collapses like a house of cards. If the concern is "it's hard to troubleshoot", well then perhaps implement better logging ("connectivity to service X has been unreliable with X% failure rate over the last Xhrs" instead of "connection terminated", or no logging at all).
- outworlder 4y ago> Extremely strict RPC settings. I’m talking zero retries (or MAYBE one) [...] > I disagree I too disagree with that too, pretty strongly. The network is unreliable. Technology has become pretty good, to the point that we have become spoiled and take the reliability for granted. But it is at its very core an unreliable medium. Applications should expect that and survive accordingly. Now, before packets go out to the network, or after they have made their way inside your network, they are in an environment that you have a degree of control. Errors and retries should be monitored. If they increase and remain elevated, they should be investigated. But guess what, if your services are resilient, you should have some time to investigate before things start breaking down - as they would if you treated everything as if it was _localhost_ Case in point: our app started breaking in horrendous ways once we deployed it in China and it tried to cross over the Great Firewall. Most issues would have been survivable if they just backed off and retried. The GFW doesn't usually like the first attempt to/from a new address, but it will usually allow the traffic after a while. We had to pay a company a not insignificant chunk of cash to get better connectivity (that would have been fine as an improvement, but not to make the system work at all). Retries are at the core of everything we do. TCP has retries (if the author followed their own advice, they should switch to UDP!). Kubernetes has a whole bunch of retries before reaching error states: CrashLoopBackOff, ImagePullBackOff. Your ethernet card has retries. But it also tracks errors. Track errors. And have retries. For as long as it makes sense to retry for your use-case. > well then perhaps implement better logging Not logging. Metrics! Logs are good for troubleshooting. For signaling issues, metrics are better. Errors can be a simple counter that you increment. Have Prometheus or similar scrape that and throw alerts as necessary.
- vinay_ys 4y agoIf you do structured binary logging, then it is better than doing unstructured/text logging and metrics separately. With structured binary logs you can extract whatever metrics you want to your hearts content and you have the freedom to turn metrics processing on/off as needed.
- nijave 4y agoYup, better advice (imo) is put metrics on all incoming requests and any external services. Add request count, request time buckets, and error counters. A lot of APM agents will instrument this automatically but it's still easy to do yourself. Some frameworks have this functionality built in
- smrtinsert 4y agoAgreed. That bullet point hit me as a very wrong. Returning failures downstream is costly and most likely a terrible user experience.
- mattpallissard 4y agoI'm with you. I recognize that "fail hard and restart" can be a valid design, but if you aren't careful you're going to hit some form of thundering herd problem at some point. > If your service can’t load the config on startup for any reason, it should just crash Not always the case. You don't want to carry on as usual as if everything is normal, but running in a degraded state or dedicated failure mode that you can coax information out of is often quite useful. It almost seems like the author err's on the side of KISS to a fault. Writing off or failing to recognize things like retry-logic, metrics, logging, failure modes, etc. It strikes me that the author may have dealt with a lot of projects that are far from feature complete and may not have seen one taken one all the way to rock-solid. And by rock-solid I mean a system that auto-remediates to a certain extent and lets you ask questions of when the sky is starting to fall. That said most of the advice is alright, especially for a relatively young project. Rome wasn't built in a day and much of the failure behavior I tend to plumb into projects gets added only after a problem is encountered. Edit: grammar
- joshka 4y agoIn practice, due to the way network / system failures tend to work at scale, failure of a first retry is generally strongly correlated with failure of a second retry. Thus a second retry can be more problematic than a first (especially if each retry causes load). From that you can infer that a single retry at the highest level is the right approach (most of the time as always YMMV). It's worth measuring this for your own services in production with real workloads by including a metric that captures how often a first and a second retry succeed. When you don't choose zero / one as your multiplier, there's a strong risk of implementing a retry strategy that is multiplicative. E.g. given 3 layers with a try and then 3 retries at each layer you cause a potential 64x (4x4x4) amplification of any failure at the lowest level. Retries are an easy way to overload a service that would otherwise recover from a problematic situation. Adaptive retry using a token bucket / circuit breaker approach are reasonable second alternative to zero/one. In practice, for resilient systems, you can actually go even further than zero retries when you have shared knowledge of an outage to the downstream service (due to concurrent calls from the same source). You can choose not make the call altogether, and look at only making a small amount of calls to the service to let it recover sanely. Obviously this is only useful for calls that are optional parts of a call chain. An example implementation skip a percentage of calls based on the percentage of failing calls (e.g. if 50% of the calls are failing due to an overloaded downstream service, backing off to only make around 50% of the call volume is directionally appropriate). Better logging is always appreciated regardless of situation ;)
- cwingrav 4y agoThe stage of testing is important here. Early on, I agree with the author. Catch code/config bugs early with no wiggle room for retries. But later, testing in a live system that lives with unreliability, you need those configs that allow for it. This enters the chaos testing phase, where you can assume with some degree that the code works deterministically, but now you have to test how it works in non-deterministic settings. Or more likely, why it failed in retrospect and how it recovered previously. This is much harder.
- throwaway012282 4y agoYou are 100% correct. It's crucial to understand and manage the failure modes of your system around transient network failures and permanent bottlenecks. Retries are a must in most systems, but need to be planned, otherwise you DoS your own network or services.
- pgo 4y agoAgreed, If you have retries then you also need circuit breakers
- saiya-jin 4y agoThis particular concern costed some 10 folks 2 mandays per head last week across the globe and due to consequences of the issue got escalated to higher management. Different topic a bit - messaging & routing ecosystem working 10 years without flaw suddenly started exhibiting slowness and randomly would just stop, needing restart of client. We debugged like crazy, java messaging system by me on one end, Tibco ems system on the other. Tibco refused to help due to server being out of support (note for us/US team). We had network guys on 13h call too, but they didn't have as much experience with WANs. After 2 days, they discovered some internal backbone network system between US and Switzerland just failed out of blue in the worst way possible - dropped some +-20% of the packets, so things kept chugging along somehow, till they didn't (some Acks on transaction commits rarely didn't happen and then all got blocked without a hint why). 2 lessons - don't always doubt yourself and your skills when SHTF. And don't take things like servers, OS, network for granted. Don't expect some clever monitoring already in place will figure out issues for you.
- quadrifoliate 4y ago> If we're talking about distributed systems, then one thing is guaranteed - network is not going to be reliable. And if we have 10's or 100's of services, this policy means that at the smallest blip the whole thing collapses like a house of cards. With RPC, I believe the author is talking about retries at the application level. There are already enough retries in the TCP layer below it that happen with exponential backoff. Tuning that and also your HTTP library's timeout settings is possible if you happen to have a unique enough network that the defaults don't work. But very likely, your slowness or problems will exist in the application layer – either on "your" side (your service is tied up doing something too long) or on the other side (their service is tied up doing something too long). The correct fix is to "Fix the flaky service!" as the author recommends, and this can take many forms – spin up more copies of the service, or fix any CPU or I/O resource problems. Slapping on another layer of "just retry" on top of all the other retries at the application layer is what the author is recommending against – this is because you will end up inventing a newer, complicated model of a distributed system.
- nine_k 4y agoNo, TCP retries are not enough. TCP retries won't help if your backend is restarting, if failover switching is happening, if your overloaded cluster has just been scaled up to add nodes, etc. A reasonable application-level retry policy (exponential randomized delays, limited attempts) would turn these from a service disruption for the client into a mere delay, often pretty short.
- quadrifoliate 4y ago> TCP retries won't help if your backend is restarting, if failover switching is happening, if your overloaded cluster has just been scaled up to add nodes, etc. Yeah, this is where the nuance begins. You are correct that TCP is not always sufficient. Perhaps where we differ is that in my experience, it still helps for this to be a feature of the framework or infrastructure that the applications are running on (e.g. a retry budget in the service mesh, or a load balancer) rather than scattered around in the application itself. At some point it becomes a word game – you could say that the service mesh is also kind of an "application" itself, but the core principle is that the retries should be kept in a few simple, common places that are rarely tuned. Otherwise, you will find that N developers who are tasked with figuring out something like this will scatter N different version of your exponential randomized delay policy all across your codebase. It is always possible to avoid this with enough code review discipline, but once the trend starts, it's much harder to say "No, you need to fix this the right way".
- zaphar 4y agoI was just about to post the exact same thing. 9 times out of 10 the issue isn't a broken service. It's going to be something environmental. You absolutely must have retries. Set alerts for elevated numbers of retries so you can respond when it can't recover on it's own but don't leave these out. You'll just generate alert fatigue for whoever is on call in your teams.
- jasonhansel 4y agoThe answer is to "kick retries up the stack"; when you fail to reach a service, you return a 503 and have your clients retry, to avoid a case where every service in the stack starts retrying all at once and causes a massive increase in traffic. IMHO you should only add retries if it proves necessary in practice to reach your SLA, if you're building something that isn't itself triggered by an RPC, or if you're performing an operation that can never be made idempotent.
- dmux 4y agoI like the idea of pushing the retry up to the caller, but in a lot of apps I've seen they've built up an in-memory object-graph from the results of calling out to other services. Wouldn't failing due to one bad service and asking the caller to retry result in every other service being unnecessarily hit?
- CBLT 4y agoI had a tech lead who vehemently agreed with the parent commenter (retry from the top), but I ended up learning different lessons. * Differentiate retryable and non-retryable errors. If the service can't return success because the DB it queries is borked, it should send a non-retryable error. Then it won't get overwhelmed by retries from upstream. * Retry configuration should have sane defaults. Even "retry once" is too often; many services aren't overprovisioned for 100% increase in traffic. What ended up working here for us was having the retry module collect req/sec statistics, then only allowing 20% of that number in retries/second. Individual requests can be retried twice. That was small enough to not push over any services, but enough retries to compensate for garden-variety unavailability. * Services shouldn't serve requests first-come-first-served under load. When SHTF, fairness means everyone has to suffer long delays, often much longer than rpc timeout. Instead, serve them in an unfair order - the most recent requests to come in are the most likely to have a caller still interested in them. Answer those! * Use headless services in kubernetes. By exposing the replicaset to the client, the client can load balance itself more intelligently. Retries should go to different replicas than the failing request. Furthermore, you can perform request hedging to different replicas than the lagging request. * Define a degraded form of your in-memory object-graph. If a feature is optional to the core business flow, it shouldn't take down your whole product. This one is a lot more involved. We needed custom monitoring for degraded responses, in-memory collection and storage of "guesses" to substitute for degraded portions of the object graph, as well as some other work I can't think of right now. This does enable an organization to compartmentalize better, having faster, less fearful deploys of newer initiatives.
- knicholes 4y agoI prefer to keep statistics like that in my metrics system instead of my logging system.
- Thaxll 4y agoRetries are complicated to implement properly, you have some many timers in an HTTP request ( tcp hanshake, tls, headers, send bytes etc ... ) Let say you have 15sec to do an API call you have to take all the timers in consideration for the retry to be ffective and also to cancel the request if you go over 15sec.