4 ms·
In case anyone is not aware: This is a very small sample of microbenchmarks. When benchmarking very simple tasks like these performance tend to vary wildly betw
by NohatCoder 7y ago
In case anyone is not aware: This is a very small sample of microbenchmarks. When benchmarking very simple tasks like these performance tend to vary wildly between architectures.
For instance instructions are assigned to one of a handful of ports when executed, certain instructions may only be assigned to certain ports, what ports an instruction may be assigned to differ between architectures. If an inner loop use only a few different instructions one architecture may be unlucky in that most of the instructions need the same ports, and so it can execute fewer instruction overall.
For real benchmarking use lots of different complicated jobs. It is not perfect, but it is the best way we have of comparing different processors head to head.
- mrb 7y agoIndeed. Back in 1999 the AMD K7 was a full 3 times faster than Intel on microbenchmarks measuring the performance of ROR/ROL instructions, because the throughput per clock of these rotate instructions was exactly 3 times higher than on Intel. Obviously this did not mean that AMD was 3 times faster than Intel. Picking 1 or 2 random microbenchmarks like the blog post author did is not useful to categorize overall performance across all real-world workloads. If he had picked different ones, they might have shown AMD twice faster than Intel.
- BeeOnRope 7y agoExamples like that still exist: AMD popcnt throughput is 4x Intel's, for example (4/cycle vs 1).
- ncmncm 7y agoThe author appears to be benchmarking the specific operations that bottleneck their json parsing library when running on an Intel chip, which seems reasonable, on its face. It can fail if the library is limited by a different set of operations, on the different machine. But that is unlikely if the specific operations tested are slower.