8 ms·
Averages Can Be Misleading: Try a Percentile (2014)
- SketchySeaBeast 8y agoMore knowledge is always better, but percentiles are a little misleading as well - the 99% at 867 ms latency makes you have a moment of panic, but when you see that 95% is 60 ms, then you really realize how few of your visitors are experiencing the slow response. Might it be a problem? Possibly, and I has brought awareness to that potential, but it also has the possibility to blow it out of proportion if you don't look at the rest of the data. Edit: I'm not saying Averages are better, but that Percentiles can be misleading as well.
- Pfhreak 8y agoHaving multiple percentiles is key, but I can't think of a time when I've ever found average to be useful.
- deleted 8y ago[deleted]
- dragontamer 8y agoThe average (or more specifically: arithmetic mean) has a number of key properties that allows for advanced analysis. If you have two pieces of data: say the average roll of a D6 (6-sided dice) + average roll of D20 (20-sided dice), you wouldn't be able to do anything with percentiles. The 90% percentile roll of D20 is 18. The 90% percentile roll of D6 is 5(ish). But the sum of these numbers tells us nothing (18 + 5 == 23, which tells us... nothing). In contrast, the mean / average roll of D20 is 10.5, while the average roll of D6 is 3.5. The average roll of D20 + average of D6 is 14. You can add and subtract averages just fine, combining separate pieces of data to make a larger conclusion. You cannot do the same with percentages. ---------- There's nothing "misleading" about averages. Its just that most people are awful at statistics. The real lesson is... learn statistics, so that you can learn to use these tools correctly. Averages, and standard deviation / variance, are excellent for combining data. Its a blurry picture, but its still mathematically correct. Percentiles / quartiles / graphs are more precise and allow for deeper conclusions... but its not always possible to create a percentile graph, especially if you cannot directly measure some attribute. "Indirect measurement", by way of arithmetic, averages, and standard deviation, is quite useful. ---------- EDIT: While I'm talking statistics, don't forget about the three kinds of averages: Arithmetic Mean is the most common, but is often meaningless. You have to also understand geometric mean, and harmonic mean. Multiplicative problems should use a geometric mean. Harmonic Mean is used when comparing speeds (roughly). Going 60MPH for 10 miles, and then going 120MPH for 10 miles is averaged using Harmonic Mean. Benchmarks should typically use harmonic mean: the arithmetic mean is meaningless in the scope of benchmarks. Etc. etc. But this is all statistics, that for some reason is rarely taught well in the US High School level. IMO, its an issue with our school system... we really should be teaching more statistics, especially given how common the analysis of data is in today's world.
- coldtea 8y ago>There's nothing "misleading" about averages Well, the fact that the average of my wealth and a Bill Gates' wealth is dozens of billions of dollars shows why averages are misleading -- and frequently used to give a false image with statistics. Misleading doesn't necessarily mean mathematically or factually wrong. Something can be totally true, and still give the false impression, and averages do that. And saying "people should learn statistics" wont change this. Even knowing statistics, the average doesn't tell me much.
- dragontamer 8y ago> Well, the fact that the average of my wealth and a Bill Gates' wealth is dozens of billions of dollars shows why averages are misleading -- and frequently used to give a false image with statistics. I'm not sure I follow your example. But let me explain my point of view first. Lets say we have a game, where we flip a coin. Heads, you win your amount of money (lets say $100,000). Tails, you win Bill Gate's money (lets say $50 Billion). If we play the game 50 times, how much money will you make on the average? Well, that's just 50 times the expected value of the game. Even if we have grossly separate results for heads vs tails, the mathematical properties of the arithmetic mean / common average remains the same. ---------- "Proper use" of averages depends entirely upon the use of the number afterwards. How you're interpreting the data and why. That's solidly within the realm of statistics: understanding exactly what "Average" means, and using the math correctly. I fully agree with you that very few people out there seem to understand the definition of "Average". In fact, I personally prefer the more precise term "arithmetic mean", specifically so that I don't mislead those who say "Average" (also, "average" is ambiguous in the statistical world: mean, median mode? If mean, then arithmetic, geometric, or harmonic??). But that's more of a writing issue as opposed to an interpretation / mathematical issue. I've used ALL average calculations (median, mode, arithmetic mean, geometric mean, and harmonic mean) at some point in my life. They're useful, and I'm not even a statistician by trade. (Actually, I use those calculations mostly in my video-game analysis...)
- ComputerGuru 8y agoYou have a knack for explaining statistics in an approachable way.
- chewbacha 8y agoPercentages without volume are still useless though. 99% percentile of 1 billion requests could still mean it's impacting a huge number of users. moral: No one number is a panacea
- pmart123 8y agoPercentiles can be misleading if the data follows a Poisson distribution, especially if the lambda coefficient is closer to 1. Latencies I would imagine, would typically be Poisson due to arrival time.
- the8472 8y ago> the 99% at 867 ms latency makes you have a moment of panic, but when you see that 95% is 60 ms, then you really realize how few of your visitors are experiencing the slow response. And then you realize that visitors are hitting your application with hundreds of requests per page load and the 99th percentile suddenly becomes your average. And then you realize that you didn't plot the windowed maximum and have some crazy hangs every now and then that block entire page loads for a whole minute.
- Xorlev 8y ago> but that Percentiles can be misleading as well. I'm not sure I agree. If they're computed wrong, sure, but this is what your system is actually doing. And honestly, the tail has a way of dictating your system's performance as a whole. > the 99% at 867 ms latency makes you have a moment of panic, but when you see that 95% is 60 ms It's easy to write off 1 in 100 users, but the reality is a little more dim. If your P(slow request) is normally distributed (it isn't always -- some requests are more expensive, some data is on worse disks, etc.), then you can compute the (extremely rough) probability a user will run into a slow request in a session: P(slow request for user) = 1 - (0.99)^N N = number of requests. For example, lets say a user visits 15 pages in a session with that call in each. They have a ~13.9% chance of running into that 99th percentile. :( Now if you're fanning out lookups (as one often does), you could easy have 50 lookups for a single request. Now you're at 39.5%! What happens in 1% of requests can become extremely important and essentially dictate your user's experience. The Tail at Scale [1] talks a lot about this. I'd recommend it as reading. [1] https://blog.acolyer.org/2015/01/15/the-tail-at-scale/ https://blog.acolyer.org/2015/01/15/the-tail-at-scale/
- cromulent 8y agoThere's a great story on 99% Invisible about averages, particularly when used to design cockpits for the average pilot. https://99percentinvisible.org/episode/on-average/ https://99percentinvisible.org/episode/on-average/
- baq 8y agoIMHO plotting the distribution should be the first step before trying to compute its statistics. If you know the shape, you can understand the values - otherwise it's guesswork.
- shittyadmin 8y agoWe've switched to box and whisker plots for most of our statistical reports, they give you a good idea for various important percentiles and adding average indicators is fairly simple if desired. Even for things like latency this can be quite useful.
- the8472 8y agoBox plots can still obscure the nature of a distribution, e.g. it might be multi-modal. Violin plots + outlier dots + additional markers are more helpful. Sometimes the CDF is also more useful than the PDF, e.g. for latencies.
- wyldfire 8y agoI wholeheartedly agree! Violin plots are a great way to get a dense appreciation for a distribution. Multi-modal distributions are completely masked by mean, even median+std.
- pytyper2 8y agoThis entire thread is great, 10 data scientists all want their own special chart to be the best. You are all wrong, you should have a view of all these charts!!!! hahaha
- shittyadmin 8y agoSeems like a good idea, I wish I had more of a statistical background for this kind of thing. Seems like it'd have proved more useful for most software purposes than the calculus I had to take instead as a prerequisite. It's basically just my boss making suggestions and me implementing, so the results are probably less than optimal for this kind of thing.
- LiamPa 8y agoSite Reliabilty Engineering goes over this in a lot more detail. https://landing.google.com/sre/books/ https://landing.google.com/sre/books/
- spenthil 8y agoSpecifically Chapter 4, under "Aggregation" https://landing.google.com/sre/sre-book/chapters/service-level-objectives/ https://landing.google.com/sre/sre-book/chapters/service-lev...
- phosfox 8y agoReminds me of “Don’t cross a river if it is four feet deep on average.” — Nassim Nicholas Taleb
- mitchtbaum 8y agothx.. good summary: http://greatesthitsblog.com/the-black-swan-nassim-nicholas-taleb/ http://greatesthitsblog.com/the-black-swan-nassim-nicholas-t...
- Lightbody 8y agoOne of my favorite (short) talks on this topic. Well worth a few minutes of your time: https://www.youtube.com/watch?v=coNDCIMH8bk https://www.youtube.com/watch?v=coNDCIMH8bk
- Aengeuad 8y agoI know it's in the spirit of the talk, but the histogram at 10:45 and the related discussion about how the latency improved for most users but the average latency increasing meaning a worse experience for other users reminds me of the anecdote a Google engineer had when Youtube started rolling out their HTML5 player, the responsiveness of the page had increased but the average latency graphs went up. This wasn't down to it being a bad update, or some users getting a worse experience - not really anyway, but the switch to the HTML5 player allowed a wider audience to start using Youtube where they wouldn't have been able to do this previously. A change increasing average latency, even on a histogram, doesn't necessarily mean it's a bad change. Look at your data indeed.
- Rafuino 8y agoThis topic always leads me to think about this great talk from Gil Tene on how NOT to measure latencies (basically, don't use averages!). https://www.youtube.com/watch?v=lJ8ydIuPFeU https://www.youtube.com/watch?v=lJ8ydIuPFeU I'm also a huge fan of how Dormando showed latency distributions in one of his recent Memcached Extstore posts. The default is 95th percentile but you can change the percentile to what matters to you (i.e. 99th percentile if you ask me!). Scroll down to see what he did and play with it. https://memcached.org/blog/nvm-multidisk/ https://memcached.org/blog/nvm-multidisk/
- camel_gopher 8y agoPercentiles can be misleading, try a histogram - https://www.circonus.com/2018/11/the-problem-with-percentiles-aggregation-brings-aggravation/ https://www.circonus.com/2018/11/the-problem-with-percentile...
- mikorym 8y agoI've used Elasticsearch + Kibana for agricultural data and similarly "expanded" the view out from averages to time series. People in agriculture love averages and it makes a lot of sense in financial data since averages preserve totals e.g.: 50 ton / ha average over 100 ha = 5 000 tons At the same time summing each individual ha gives you 5 000 tons total. But once you realise that you can expand on this, things get really interesting. I don't know of other people working on the same problems that I am working on, but they are relevant both economically (in the sense of making money) and environmentally (in the sense of improving efficiency and managing climate).
- sohkamyung 8y agoCheck out this comic on "Why Not to Trust Statistics" [1]. His book, "Math With Bad Drawings" [2] has a chapter on statistics and why not to trust a single statistical measure only. [1] https://mathwithbaddrawings.com/2016/07/13/why-not-to-trust-statistics/ https://mathwithbaddrawings.com/2016/07/13/why-not-to-trust-... [2] https://mathwithbaddrawings.com/2018/05/23/math-with-bad-drawings-the-book/ https://mathwithbaddrawings.com/2018/05/23/math-with-bad-dra...
- novaleaf 8y agoMy own solution, which might be useful to those using javascript (nodejs or browser): I use mathjs.quantileSeq() and log 0%, 25%, 50%, 75%, and 100%. This seems to be good for "casual metric logs". I've found that this gives a good shape of the data, as well as the absolute min/max values. If you use 1% or 99% you'll miss the absolute worst performers, and I want to be at least aware of what the worst performance numbers are. https://mathjs.org/ https://mathjs.org/ https://mathjs.org/docs/reference/functions/quantileSeq.html https://mathjs.org/docs/reference/functions/quantileSeq.html