4 ms·
Sorry, what would AMD's or Intel's "latest and greatest" numbers for the same be?
by tandr 6y ago
Sorry, what would AMD's or Intel's "latest and greatest" numbers for the same be?
- sliken 6y agoHere's the M1: https://www.anandtech.com/show/16252/mac-mini-apple-m1-tested https://www.anandtech.com/show/16252/mac-mini-apple-m1-teste... Scroll down to the latency vs size map and look at the R per RV prange. That gets you 30ns or so. Similar for AMD's latest/greatest the Ryzen 9 5950X: https://www.anandtech.com/show/16214/amd-zen-3-ryzen-deep-dive-review-5950x-5900x-5800x-and-5700x-tested/5 https://www.anandtech.com/show/16214/amd-zen-3-ryzen-deep-di... The same R per RV prange is in the 60ns range.
- epistasis 6y agoCould this be coming from the page size being 4x as large for Apple Silicon versus x86? I don't fully understand the benchmark, but it appears to be accessing a variety of pages from the same first level TLB lookup? It's been a long time since I dealt with this stuff (wanted to get 1GB huge pages in Linux for some huge huge hash tables), so maybe I'm misunderstanding.
- sliken 6y agoCachelines, page sizes, and size of the TLB all play a role. But with tinkering you can see those effects yourself and I played with 1, 2, 4, 8, 16, and 32 "pages" which I assumed were 4KB each and didn't see much difference. Measured latencies do increase slowly, but you expect that as the TLB becomes progressively more of a bottleneck. If you use a 1GB array and see full random with much higher latency than a sliding window then you can be pretty sure that the page size is much less than 1GB. Getting the cacheline off by a factor of 2 does make a small difference since you get occasional cache hits instead of zero, but as long as the array tested is several times larger than cache the impact is small. But all in all the M1 has excellent memory bandwidth, excellent latency, and shows significantly better throughput on random workloads as you use more cores. Normal PC desktops have 2 memory channels (even the higher end i7/i9/ryzen7/ryzen9), only the $$$$ workstation chips like threadripper and some of the $$$$ Intel's have more. The little ole M1 in a mac mini, starting at $700 has at least 8 memory channels. So basically the M1 delivers on all fronts, larger and lower latency caches, wide issue, large reorder buffers, excellent IPC, and excellent power efficiency.
- tandr 6y agoThank you very much. So we are talking about doubling (or halving, depending what side you are looking from) the access times.
- andrei-at 6y agoHi, Andrei here. Just a clarification as to why the P per RV prange numbers are good: This pattern is simply aggressively prefetched by the region prefetcher in the M1, while Zen3 doesn't pull things in as aggressively.