3 ms·
"benchmarks always tell you if there's a problem and exactly where it is, and squash any unneeded discussion." Unfortunately, always is too strong of a word. Ca
by 6keZbCECT2uB 5y ago
"benchmarks always tell you if there's a problem and exactly where it is, and squash any unneeded discussion." Unfortunately, always is too strong of a word. Cache eviction in one parts can cause memory stalls on another part, indirection in the caller can prevent speculation. Type erasure can prevent inlining resulting in the called function being blamed for problem in the caller.
Your problem might not even be CPU, if it's contention related, or timing related, overloaded queues, not pushing back at the right places, io bound, the bottleneck is work which is queued and executed elsewhere...
Causal profiling is a technique which is relevant specifically because profiling can miss the forest for the trees: https://github.com/plasma-umass/coz https://github.com/plasma-umass/coz
It's really easy to write a benchmark which measures a different scenario from what your application is doing. A classic example might be benchmarking a hashmap in a loop when that hashmap is usually used when cold.
I definitely agree about directing efforts to where you can make an impact and guiding that through measurement, but benchmarks can miss that there's a problem and blame the wrong part of the application.
If the difference is large enough, ms vs hours, you'd have to really screw up methodology to get the wrong result (I've done it almost that badly before).