3 ms·
Thank you for the feedback. I agree with most of your points. > Most -- nearly all -- benchmarking tools like this work from a normality assumption I don't th
by sharkdp 7y ago
Thank you for the feedback. I agree with most of your points.
> Most -- nearly all -- benchmarking tools like this work from a normality assumption
I don't think that hyperfine makes any assumption about normality. Sure, we do report sample mean and sample standard deviation by default, but we also report sample minimum and the maximum. You can also easily export all the benchmark results and inspect in more detail with the supplied Python scripts.
> In fact, performance numbers (latencies) often follow a heavy-tailed distribution
So when is this really the case? In my understanding, if I am measuring the runtime of a deterministic program with the same input, the runtime should only be influenced by external factors that are out of my control (other programs being scheduled, caching effects, hardware-specific influences, ..). These are exactly the things that I want to "average out" by running the benchmark multiple times.
> What's worse is when these tools start to remove "outliers".
Hyperfine never removes outliers. What we do is to try and detect outliers. We do this by computing robust statistical estimates that specifically DO NOT assume a normal distribution (see https://github.com/sharkdp/hyperfine/blob/master/src/hyperfine/outlier_detection.rs https://github.com/sharkdp/hyperfine/blob/master/src/hyperfi... for details).
We perform this outlier detection to warn users about potentially interfering processes or caching effects.
Take a look at these results, for example: https://i.imgur.com/XRvE6Ys.png https://i.imgur.com/XRvE6Ys.png
I benchmarked a file-searching program. The underlying distribution, while probably not normal, seems to be "well behaved" and I think that the sample mean and the sample standard deviation could be quantities with a reasonably predictive power.
What you do NOT see in the histogram is a single outlier at 1.15 seconds, far outside the plot to the right. This was the first benchmark run where the disk caches were still cold. In such a case, hyperfine warns the user:
Warning: The first benchmarking run for this command was significantly slower than the rest (1.152 s). This could be caused by (filesystem) caches that were not filled until after the first run. You should consider using the '--warmup' option to fill those caches before the actual benchmark. Alternatively, use the '--prepare' option to clear the caches before each timing run.
In conclusion, I am not quite sure how your critisism applies to hyperfine, but I'd be happy to get further feedback.
- kqr 7y agoThank you for responding! I was curious to hear your thoughts on this. > Sure, we do report sample mean and sample standard deviation by default, but we also report sample minimum and the maximum. You can also easily export all the benchmark results and inspect in more detail with the supplied Python scripts. Hyperfine presents itself as a black-box tool where you plug in program and it is supposed to outputs usable numbers. For that type of tool, "but you can script up a custom report" is not how you defend wildly speculative reporting! I know trying to black-box performance measurements is very, very annoying, because things we are used to take for certain simply don't hold. But that is an inherent problem of the field, and not something one can wish away. > In my understanding, if I am measuring the runtime of a deterministic program with the same input, the runtime should only be influenced by external factors that are out of my control (other programs being scheduled, caching effects, hardware-specific influences, ..). These are exactly the things that I want to "average out" by running the benchmark multiple times. This is broadly correct. The common misconception regards how many runs are needed to successfully average out these external facts: probably more than you want your user to wait through. Sums of numbers drawn from heavy-tailed distributions converge very slowly to the normal distribution, to the point where a lot of people people won't wait for it to become even close to normal. > Hyperfine never removes outliers. What we do is to try and detect outliers. We do this by computing robust statistical estimates that specifically DO NOT assume a normal distribution (see ... for details). Sorry, that was my misreading. I'm glad you don't remove outliers! I like that you're using robust estimations, but I'm still not convinced they work as well as we would want to. I'm sure someone else could formalise this, but just based on very simple experimentation[1], I get some worrying results: When the cutoff value is chosen so that D > 3.5 is labeled as an outlier, the samples labeled as "outliers" reliably contribute around 0.75 to the expectation of the sample. Despite using robust estimations, the outliers completely dominate any expectations about the sample. > Take a look at these results, for example: https://i.imgur.com/XRvE6Ys.png https://i.imgur.com/XRvE6Ys.png > I benchmarked a file-searching program. The underlying distribution, while probably not normal, seems to be "well behaved" and I think that the sample mean and the sample standard deviation could be quantities with a reasonably predictive power. To me, that also looks like a heavy-tailed distribution where there simply aren't enough samples to reveal the extremal values that exist in the real population. If you still have the raw data, we could try a K-S test against the MLE fitting of some common heavy-tailed distributions to see if it's possible to rule them out, but I suspect we won't be able to do that. [1]: https://two-wrongs.com/pastes/outliers-529ad9.r.html https://two-wrongs.com/pastes/outliers-529ad9.r.html