10 ms·
Tail latency might matter more than you think
- timf 5y agoThis reminded me of the similar cascading effect you get with availability across the aggregate of many services as discussed in "The Calculus of Service Availability": https://queue.acm.org/detail.cfm?id=3096459 https://queue.acm.org/detail.cfm?id=3096459
- PaulHoule 5y ago"Tail Latency Matters More Than You Think" is more like it. When you see a beach ball or other indicator of delay on your computer you are quite likely to be experiencing tail latency.
- cratermoon 5y agoA lovely thing about tail latency in the chains that the post talks about is how one service being slow can cascade. Especially in serial chains, when on component is slow, the rest are waiting, in the meantime using resources like memory, sockets, cpu cycles, that could be used to service other requests. In the worst cases, those other services can start responding slowly to other requests, resulting in further degradation. Having circuit breakers and carefully tuning timeouts can help.
- mjb 5y ago> Having circuit breakers and carefully tuning timeouts can help. Can help, for sure. Can also turn a 'slow' into 'unavailable', which could be what you want, or could be a disaster.
- cratermoon 5y agoI'd take an 'unavailable' lasting for a short time while connections exponentially back off with jitter over minutes or hours of degraded service for everyone.
- klodolph 5y ago> Choosing Summary Statistics This is the #1 thing that people get wrong. It's something that otherwise smart software engineers get wrong because they don't have enough of a background in data analytics. The problem has two parts. One part of the problem is that once you reduce an observation to summary statistics, you can't go back. The other part of a problem is that web services usually generate too much data if you don't summarize.
- nvilcins 5y agoWhat are good guidelines to tackle this?
- klodolph 5y agoCan’t think of a way to reduce this to guidelines. I’ve seen a bunch of different errors. Mostly it’s a bad understanding of the math, or a bad understanding of what you’d want to do with the collected data. Just one example that comes to mind: suppose you want average (mean) latency. So you record, for every minute, the average latency for requests during that minute. Now you want to calculate the average latency for an entire day… but it’s impossible. You can calculate the daily average of the per-minute averages, but you’re averaging over minutes, rather than averaging over requests. This is just a simple example, but you can see how a seemingly innocuous decision sabotaged your ability to run the query you want. Maybe my advice is “learn calculus”.
- noodlenotes 5y agoAverage is actually one of the "nicer" summary statistics to work with because you can recalculate it over a different aggregation level if you kept the volume. It's statistics like median, percentile that you have to worry about. In my experience, people over-rely on averages just because they're easy to re-aggregate when they really should be using something else.
- klodolph 5y agoYeah, median and percentiles are awful. You usually want them only at the very top-end, as the final summary of some system. Like, “page renders within 500ms for 90% of users”. So, well meaning engineers will collect percentiles on individual components in the system, with the idea that if you want percentiles at the end, just collect percentiles everywhere. You’re left with garbage data that can’t be aggregated in a way that makes any sense.
- cassianoleal 5y agoThe talk "How NOT to Measure Latency" [0] taught me all I needed to know in order to start worrying about tail latency in a very well presented way. [0] https://www.infoq.com/presentations/latency-response-time/ https://www.infoq.com/presentations/latency-response-time/
- atombender 5y agoIt's a fantastic talk that taught me several important principles about measuring performance. Gil Tene also developed HdrHistogram, which has been ported to a bunch of languages, and is a great, lightweight way to collect accurate performance histograms: http://hdrhistogram.org/ http://hdrhistogram.org/.
- cassianoleal 5y agoThanks, I'll check it out. Already a winner on my book for this bit of humour: Support or Contact Don't call me, I won't call you.
- FriedrichN 5y agoThis is why I don't like require.js, one script requires this script which requires that script. If there is one hiccup somewhere down the line it causes the whole page to have to wait. One of my clients had their website made and wondered why it was so slow (the designers said it needed a faster server) but I found out it was requesting hundreds of .js files in roughly 10 waves. Causing the whole page to take up to 10 seconds to load completely.
- wging 5y agoWith requirejs, that work can and should be done at build time. There shouldn't be a dependency waterfall. https://requirejs.org/docs/optimization.html https://requirejs.org/docs/optimization.html
- FriedrichN 5y agoExcept nobody does that. They don't know how to. For the record, I bundle all my JS/CSS assets in one file.
- earthboundkid 5y agoIt sounds like you and your clients need to learn more about modern JS bundling techniques. Making a waterfall of requests can be okay in development but should never ship to production.
- kqr 5y agoAnother point often missed is the diagnostic value of tail measurements. One of the first things I do at any job is replace the 90th percentile with the maximum in all plots. Sure, it gets messier, and definitely less visually appealing, but the reaction by others has uniformly been "Did we have this data available all along and just never showed it?!" It's also worth mentioning that even in a system where technically tail latencies aren't a big problem, psychologically they are. If you visit a site 20 times and just one of those are slow, you're likely to associate it mentally with "slow site" rather than "fast site".
- viraptor 5y agomin(max_time, const_reasonable_max) is also a good graph if your software supports it. It stops the outliers from polluting the view. After all, your user will leave/refresh after a few seconds, so only matters your response took longer than a minute - it's irrelevant it was 20minutes.
- kqr 5y agoDue to the systems in question themselves timing out relatively soon, this is what I in practise end up looking at anyway. It's a good point, even though it cuts both ways: given some assumptions about the tail behaviour of latencies, the 20 minute extreme event is a treasure trove for estimating the probabilities of smaller tail events.
- jberryman 5y agoFor most of the services I've looked this closely at, maximum would be a proxy for load. Or in the case of a benchmark suite you'd expect max to increase with the number of iterations. What do you find the maximum useful for?
- kqr 5y agoAren't all performance numbers a proxy for load, almost by definition? I tend to look at them per request, per user, per iteration, and so on, to control for that effect.
- deleted 5y ago[deleted]
- nostrebored 5y agoOne complaint here is that serial/parallel is not the way to think about most modern architectures. In modern architectures you are typically working with decoupled event buses which invert the relationship with dependencies. In this case, you become resilient to many negative impacts of tail latency as you're inherently eventually consistent.
- tybit 5y agoI don’t know that I’d call this modern so much as just a different school of thought. Using streaming/event buses and getting rid of the chain of calls in front of the initial request definitely has upsides. Syncing data all over the place, with consumers being forced to interpret that data from outside of their domain has large downsides too though.
- jkire 5y ago> A common pattern in these systems is that there's some frontend, which could be a service or some Javascript or an app, which calls a number of backend services to do what it needs to do. I think an important idea here is that you should be trying to measure the experience of a user (or as close as possible). If there is a slow service somewhere in your stack, but has no impact on user experience, then who cares? Conversely, if users are complaining that the app feels sluggish, then it doesn't matter if all your graphs say that everything is OK. I find it helpful to split up graphs/monitoring into two categories: 1) if these graphs look fine then the service is probably fine, and 2) if problems are being reported then these graphs might give an insight into why things are going wibbly. In general, we alert on the former and diagnose with the latter. Of course, its nigh on impossible to get perfect metrics that track actual user experience, but we've definitely found it worthwhile to try and get as close as possible to it. --- Another fun problem with using summary statistics is they can easily "lie" if the API can do a variable amount of work. For example, if you have a "get updates API" that is called regularly to see updates since the last call, then you end up with two "modes": 1) small amount of time between calls and so super fast and 2) a large amount of time between calls and so is slow. Now, in any given time period the vast majority of the calls are going to be super quick, but every user will hit the slow case the first time they open the app for the first time that day. This results in summary statistics that all but ignore those slow API calls when opening the app.
- jeffbee 5y agoThe flip side is that errors often pollute service latency statistics. If your service is capable of serving a fast failure, for example by returning 503 instantly for all requests when it is overloaded, you need another dimension in your statistics to handle that.
- oavdeev 5y agoI have once built an interactive calculator[1] for this exact problem, maybe someone else will find it useful too [1] https://observablehq.com/@oavdeev/parallel-task-latency-calculator https://observablehq.com/@oavdeev/parallel-task-latency-calc...