3 ms·
> but that Percentiles can be misleading as well. I'm not sure I agree. If they're computed wrong, sure, but this is what your system is actually doing. And ho
by Xorlev 8y ago
> but that Percentiles can be misleading as well.
I'm not sure I agree. If they're computed wrong, sure, but this is what your system is actually doing. And honestly, the tail has a way of dictating your system's performance as a whole.
> the 99% at 867 ms latency makes you have a moment of panic, but when you see that 95% is 60 ms
It's easy to write off 1 in 100 users, but the reality is a little more dim. If your P(slow request) is normally distributed (it isn't always -- some requests are more expensive, some data is on worse disks, etc.), then you can compute the (extremely rough) probability a user will run into a slow request in a session:
P(slow request for user) = 1 - (0.99)^N
N = number of requests.
For example, lets say a user visits 15 pages in a session with that call in each. They have a ~13.9% chance of running into that 99th percentile. :(
Now if you're fanning out lookups (as one often does), you could easy have 50 lookups for a single request. Now you're at 39.5%! What happens in 1% of requests can become extremely important and essentially dictate your user's experience.
The Tail at Scale [1] talks a lot about this. I'd recommend it as reading.
[1] https://blog.acolyer.org/2015/01/15/the-tail-at-scale/ https://blog.acolyer.org/2015/01/15/the-tail-at-scale/