3 ms·
and then runs the function 1000 more times, measuring each run independently and reporting the average runtime at the end. When it comes to measuring the perfo
by csl 13y ago
and then runs the function 1000 more times, measuring each run independently and reporting the average runtime at the end.
When it comes to measuring the performance of code like this, averaging run times is not the way to do it.
To remove the noise caused by context switching, just run the code many times and report the single fastest run you get. This should be the value closest to running the code on an OS without preemption (i.e. you want to measure how fast the code runs on the bare metal without interruption).
Even Facebook's Folly library [0] changed their benchmarking code from using statistics to just providing the fastest run. As the comments say:
// Current state of the art: get the minimum. After some
// experimentation, it seems taking the minimum is the best.
return *min_element(begin, end);
This is explained in the docs [1]:
Benchmark timings are not a regular random variable that fluctuates around an average. Instead, the real time we're looking for is one to which there's a variety of additive noise (i.e. there is no noise that could actually shorten the benchmark time below its real value). In theory, taking an infinite amount of samples and keeping the minimum is the actual time that needs measuring.
[0]: https://github.com/facebook/folly/blob/master/folly/Benchmark.cpp#L142 https://github.com/facebook/folly/blob/master/folly/Benchmar...
[1]: https://github.com/facebook/folly/blob/master/folly/docs/Benchmark.md#a-look-under-the-hood https://github.com/facebook/folly/blob/master/folly/docs/Ben...