3 ms·
I don't think you really understand the architecture of the machine. EPYC looks on paper like it has a large L3 cache, but it consists of separate L3 caches pe
by ebikelaw 8y ago
I don't think you really understand the architecture of the machine. EPYC looks on paper like it has a large L3 cache, but it consists of separate L3 caches per "core complex" of which there are two per die and four dies per package. So what you've actually got is a bunch of redundant 8MB caches, which is not the same thing. Because of the baroque topology, especially when you have two sockets, access to main memory varies between almost-as-fast-as-xeon to way-way-slower-than-xeon. Combined with the small caches it's a total disaster.
- zrm 8y agoThe amount of L3 per thread is still 64MB/#threads. Where the difference you're describing most matters is for single-threaded code, where in theory the one thread could otherwise have the entire 64MB. But that isn't the circumstance with the NUMA latency anyway. If there is only one thread the OS can schedule it on the same node as its data. Most working sets fit in even 8MB (or less) -- the reason for 64MB is to provide for multiple threads. In which case if one thread isn't using its proportionate share there are seven others that can use it. Sharing with sixty-three instead would be "better" but at some point it's diminishing returns.
- philjohn 8y ago> Combined with the small caches it's a total disaster. I think you know that "total disaster" is pretty much hyperbole.