3 ms·
This seems to be a cache performance issue: - The difference in performance is too big to be a memory bus issue - Since the number of threads doesn't change,
by kmod 16y ago
This seems to be a cache performance issue:
- The difference in performance is too big to be a memory bus issue
- Since the number of threads doesn't change, just the number of processes, it's not an un-scalability issue of the program.
- It's most likely not a issue with which CPUs he chose; in a hyperthreaded system there can be a pretty large performance hit if he pinned the server to logical cpus that were on the same core. I don't know how Intel numbers its cores in hyperthreaded processors, but their general method is "make the cores that are closest to each other have the biggest difference in cpu ids", which would mean that ex cpus 1 and 5 are on the same core.
- I would guess that it's not a thread migration issue, since I wouldn't expect the Linux scheduler to be that bad.
Poor cache performance is the most likely cause. As a minimal example of what the author concludes, if you have two threads that sit in a loop incrementing the same counter, it is much faster to run those two threads on the same cpu than on their own cpus. This is because when running on separate cpus, that particular cache line will continually be ping-pong'ed between the two. Even if threads don't access the same addresses, if they access addresses that are in the same cache line, you'll still get ping-ponging ("false sharing"). This explanation is also supported by the fact that it disappears on i7 processors; one of the largest improvements with the Nehalem line is the improved cache architecture.
Poor cache behavior is usually the culprit in non-cpu-bound and non-io-bound systems.