5 ms·
Statistically rigorous Java performance evaluation
- filereaper 9y agoSPECjvm98 is an outdated measure of both system and JVM performance, the benchmark to look at is SPECjbb2015 which very aggressively taxes JVM subsystems like the GC and the JIT.
- efferifick 9y agoYes, but this paper is old so they did the right thing by using SPECjvm98 at the time the paper was written.
- scott_s 9y agoI prefer reporting the mean and the standard deviation - the paper advocates a confidence interval instead of standard deviation. Typically, I'm more concerned with the spread of obtained performance values than I am with how likely it is that our measured mean is the within some interval. I generally don't think of that spread of obtained values as noise or random errors, but as systematic consequences of using real computing systems. The reason I don't consider that systematic error is that the sources of variation in real computer systems are often the result of things like memory hierarchies and system buffers that will exist in practice. Real systems will have these things, so I want my experiments to have them as well - so long as our benchmark has them in the same way a real production system will have them. For example, see Table 2 in a recent paper I am a co-author on (page 8 of the pdf, page 73 using the proceedings numbering): http://www.scott-a-s.com/files/debs2017_daba.pdf http://www.scott-a-s.com/files/debs2017_daba.pdf In this paper, we care about latency, and we report the average latency along with the standard deviation. Here, a tighter standard deviation is more important than confidence that the mean falls within a particular range. And the variation in latencies is caused by both software and hardware realities of the memory hierarchy.
- chrisseaton 9y agoWhen you want to make a formal scientific claim - that technique A is faster than technique B, for example - how do you support that using a standard deviation as an error? I sometimes have used standard deviations for errors myself, but I'm not really sure I can defend it.
- scott_s 9y agoWhat I'm saying is that it's not error. I'm reporting a distribution that the values may take on. Ideally I would show all of the data, but that's not possible given both presentation and human cognitive limits, so we characterize that distribution with two dimensions: a mean (the "middle") and the standard deviation (how much it "spreads" from that middle). This is, I think, a philosophical difference from, say, measuring the charge of an electron. The electron has a charge. The mean of independent measurements is, we hope, very close to that true value. Deviations from that mean are indeed error. But when measuring performance in a computer system, all observable values are valid. We're trying to figure out not what the "true" value is (there isn't one), but what the range of values your performance can take on, and where in that range you're likely to fall. (I almost wrote something about excusing "real" errors such as memory corruption, disk failures and segfaults, but sometimes you want to include that! If you're doing performance analysis of a distributed system that scales to hundreds of thousands of compute nodes, real runs of an application are likely to encounter such things, so your system better be resilient to them, and they will have an impact on real performance.)
- slaymaker1907 9y agoFor microbenchmarks, I've seen it recommended that you just take the min value. The idea behind it is that any extra time from the min is very likely noise so the min is closest to what you want to measure.
- scott_s 9y agoI strongly disagree - I care about how the thing will perform in practice, which includes what people often consider "noise." The minimum value for a metric is just when all of the stars, sun, moon and planets lined up juuuuust right, and is often not representative of real performance. What if, for example, the technique performs great when everything is in the cache, but it also uses the data in such a way that it has horrible cache locality? Your min will be great, but your average, max and standard deviation will be terrible.
- deleted 9y ago[deleted]
- igouy 9y agoMore recently: "Quantifying performance changes with effect size confidence intervals" Tomas Kalibera and Richard Jones Technical Report 4-12, University of Kent, June 2012. https://www.cs.kent.ac.uk/pubs/2012/3233/ https://www.cs.kent.ac.uk/pubs/2012/3233/ Kalibera, Tomas and Jones, Richard E. (2013) "Rigorous Benchmarking in Reasonable Time" https://kar.kent.ac.uk/33611/ https://kar.kent.ac.uk/33611/