4 ms·
A 2-socket AMD setup has the same NUMA topology as an 8-socket Intel machine, and because of that the performance is terrible on many workloads.
by ebikelaw 8y ago
A 2-socket AMD setup has the same NUMA topology as an 8-socket Intel machine, and because of that the performance is terrible on many workloads.
- zrm 8y agoThere are a lot of things that have to go the wrong way at once to get to the point where that really matters. The first is that your working set won't fit in the processor caches and has regular cache misses into main memory -- but most of the Epyc line has 64MB of L3 cache. Then the access pattern has to be random rather than sequential, which knocks out a major class of the applications satisfying the first criteria (all the ones that process big files in sequential order). Then the operating system scheduler has to fail to schedule the process on a core in the same node as its data, most commonly because you have a process with more active threads than there are threads per node. What you're left with is, basically, large databases. But large databases also benefit significantly from more cores, memory channels and I/O. Which factor dominates is going to depend on specific usage, e.g. a database with randomly accessed individual bits will be more sensitive to latency whereas one containing pictures or other medium-large blocks of data will be more sensitive to memory bandwidth. You can certainly find a worst-case usage pattern for one or the other but in general they're going to counterbalance each other.
- ebikelaw 8y agoI don't think you really understand the architecture of the machine. EPYC looks on paper like it has a large L3 cache, but it consists of separate L3 caches per "core complex" of which there are two per die and four dies per package. So what you've actually got is a bunch of redundant 8MB caches, which is not the same thing. Because of the baroque topology, especially when you have two sockets, access to main memory varies between almost-as-fast-as-xeon to way-way-slower-than-xeon. Combined with the small caches it's a total disaster.
- zrm 8y agoThe amount of L3 per thread is still 64MB/#threads. Where the difference you're describing most matters is for single-threaded code, where in theory the one thread could otherwise have the entire 64MB. But that isn't the circumstance with the NUMA latency anyway. If there is only one thread the OS can schedule it on the same node as its data. Most working sets fit in even 8MB (or less) -- the reason for 64MB is to provide for multiple threads. In which case if one thread isn't using its proportionate share there are seven others that can use it. Sharing with sixty-three instead would be "better" but at some point it's diminishing returns.
- philjohn 8y ago> Combined with the small caches it's a total disaster. I think you know that "total disaster" is pretty much hyperbole.