5 ms·
How Not to Measure Latency (2015)
- leetrout 4y agoNeeds (2015). I loved the talks from Gil Tene. I always reach for his fork of wrk whenever I need to test throughput: https://github.com/giltene/wrk2 https://github.com/giltene/wrk2
- rootw0rm 4y agoThis is why I love HN. I'm actually working on a lightweight metrics system right now. This accurate little tool (and the author's article) is exactly what I needed right now.
- kqr 4y agoMy poison is vegeta, but one of the genius ideas behind both are HDR histograms, both as a design, as an implementation detail, and a data transport mechanism. There are few production projects where I don't have a reason to depend on HDR histograms for various under the hood functionality.
- efitz 4y agoI found this article hard to read; it felt like bulleted noted after watching a presentation and many of the statements lacked the context to back them up. The bottom line is that measuring performance is hard, and that you have to be aware that your measurement code and your metrics system have a high likelihood of misleading you if you don’t deeply understand how they work. As a side note, the author didn’t mention my favorite metrics pet peeve, that is the right hand cliff. Often systems will graph a 0 while waiting for a metrics bucket to fill, as when you are aggregating counts over time intervals. This can result in your metric appearing to drop off a cliff in the very recent past.
- dikei 4y ago> I found this article hard to read; it felt like bulleted noted after watching a presentation and many of the statements lacked the context to back them up. Because it is. Here's the presentation. https://www.youtube.com/watch?v=lJ8ydIuPFeU https://www.youtube.com/watch?v=lJ8ydIuPFeU
- neonate 4y agohttps://archive.ph/UiyIl https://archive.ph/UiyIl http://web.archive.org/web/20221227140440/http://highscalability.com/blog/2015/10/5/your-load-generator-is-probably-lying-to-you-take-the-red-pi.html http://web.archive.org/web/20221227140440/http://highscalabi...
- deleted 4y ago[deleted]
- dang 4y agoDiscussed at the time: How Not to Measure Latency - https://news.ycombinator.com/item?id=10334335 https://news.ycombinator.com/item?id=10334335 - Oct 2015 (23 comments) Maybe we'll pinch that non-baity title...
- chis 4y ago95% latency seems way more useful as a metric than max() would be. It’s better to spend most of your time improving the experience of most users rather than debugging ultra-rare issues that rarely affect users.
- Gh0stRAT 4y agoIt depends how many requests (both frontend and backend) each user triggers and the odds of a big latency spike on at least one of them. For single requests, targeting 95% would probably be just fine back in the days of static sites served by a single Apache instance, but that doesn't describe many modern systems. Nowadays, a single HTTP request to a load balancer can trigger a flurry of blocking queries to fulfil it. (eg a distributed ElasticSearch query that can't combine the results from each node and return until EVERY node has responded) then your worst-case performance will quickly begin to dominate.
- jcelerier 4y agoUp to the day one of these ultra-rare issues falls on the journalist writing a review in a specialist magazine that will make or break your product
- winrid 4y agoUntil max() is your largest, top paying customers.
- Akronymus 4y agoOr you use microservices where a lot of requests are sent between the services. Which results in max being more and more likely for each user made request. https://youtu.be/_Zoa3xkzgFk https://youtu.be/_Zoa3xkzgFk
- dikei 4y ago95% is really not a lot, it's 5 slow requests out of 100. It takes ~300 requests to load the homepage of Amazon. You need a lot more "9"
- andreareina 4y agoI don't remember where I saw it but there's also the idea of setting a threshold and asking what percentage of users exceeded it, e.g. how many users are seeing page load times > 1 second.
- throwghkgjn 4y agohttps://linuxczar.net/blog/2019/10/28/prometheus-histograms-part-3-using-something-else/ https://linuxczar.net/blog/2019/10/28/prometheus-histograms-... and maybe https://en.m.wikipedia.org/wiki/Apdex https://en.m.wikipedia.org/wiki/Apdex ?
- philbo 4y agoI think there's times when using a max makes sense and times when using a percentile makes sense. For measurements that are entirely within my own infrastructure, I'll always use the max. Outliers there are my responsibility and I want to fix them asap. For measurements that originate from clients, I'll always use percentiles. There you're measuring the internet, which means people on patchy wifi connections or someone who's train just went through a tunnel. The max will always be spiky and there's nothing you can do about it.
- kqr 4y agoYou may (or may not!) be conflating summary statistics with signal filtering here. For latencies as experienced by interactive users, the tail of the distribution (very high percentiles including maximum) are the summary statistics that matter, period. It sounds like in your case the data is a mixed distribution of both non-actionable garbage and actionable signals, and if you could filter out the actionable, you would benefit from looking at the tail of that distribution. However, the best way you have of filtering out the garbage currently is to filter out the central portion of the mixed distribution. This comes with some false negatives and false positives, and may (or may not!) be the ideal filter in your case. It would be interesting to try to find out!
- denton-scratch 4y ago> Service time is how long it takes to do the work. > Response time is the amount of time spent waiting before the work starts. Is that right? I thought response time was the time from user input to completed response. I'm not sure (I found parts of the article hard to parse), but I got the impression that author's generator won't generate a new request until the previous one is complete. That's not how I'd write a generator. Also, it seems odd that the test-rig would discard responses that arrive outside of the test window. Surely it should record all responses to requests that were emitted during the window, even if they arrive outside the window?
- cassianoleal 4y ago> Is that right? I thought response time was the time from user input to completed response. You're right. In fact, that's how the Gil Tene defines it in the linked talk. It's essentially service time + wait time.
- denton-scratch 4y agoPlus network queueing time: oh, I guess that's rolled into wait time.
- cassianoleal 4y agoYes, wait time from the point-of-view of the user, so everything that's not actual service time between time of request and fulfilled response.
- nequo 4y agotl;dr: The 95th percentile response time is a bad metric. A typical user executes more than a single request, so it is close to 100% probable that they get a response time above the 95th percentile. Better to tune the maximum response time instead. Edit: When comparing two systems, the Fisher–Tippett–Gnedenko theorem gives guidance about how to do large-sample statistical inference on estimates of maxima: https://en.wikipedia.org/wiki/Fisher%E2%80%93Tippett%E2%80%93Gnedenko_theorem https://en.wikipedia.org/wiki/Fisher%E2%80%93Tippett%E2%80%9...