2 ms·
I think if you look closely at the specs, I think the Xeon has about 2.5MB of L3 per core, while P8 has about 8MB of L3/core. There are bigger differences than
by PhuFighter 12y ago
I think if you look closely at the specs, I think the Xeon has about 2.5MB of L3 per core, while P8 has about 8MB of L3/core. There are bigger differences than this. But the examples above have both processors set at SMT2 (~30 seconds vs ~10 seconds). Differences are greater for SMT4 and SMT8. Of course, if your code fits nicely into the L3 - e.g. if your hot code consumes 3MB, then Power8 will be a bigger winner since there may be more thrashing for the x86 chip - unless, of course, the prediction routines is straight forward and the cacheline can be prefetched (think sequential access vs random access).
From what I saw elsewhere, highly parallel code runs faster on P8, where as single thread perf is faster on x86. So if your HPC app is basically single threaded on one of your compute nodes, then that would be faster. But if your HPC code is highly multithreaded on the core, then P8 may surprise you.