4 ms·
What am I missing here? The marketing presentations spoke about 200GB/sec and 400GB/sec parts. Existing CPU's generally have 10s of GB/sec. But I see these part
by spitfire 5y ago
What am I missing here? The marketing presentations spoke about 200GB/sec and 400GB/sec parts. Existing CPU's generally have 10s of GB/sec. But I see these parts beating out existing parts by small margins - 100% at best.
Where is all that bandwidth going? Are none of these synthetic benchmarks stressing memory IO that much? Surely compression should benefit from 400GB/sec bandwidth?
This also raises the question how are those multiple memory channels laid out? Sequentually? Stripped? Bitwise, byte wise or word stripped?
- hajile 5y agoThey have 5.2 and 10.4 TFLOPS of GPU power to feed in addition to 10 very wide cores.
- kllrnohj 5y agoThe SoC having 400GB/s of memory bandwidth doesn't mean that any individual CPU core can saturate 400GB/s of memory bandwidth (or even if all the CPU cores combined can achieve that). The SoC's memory bus is also feeding the GPU, and GPUs tend to be _very_ memory bandwidth hungry (see other discreet GPUs pushing over 1TB/s of memory bandwidth). The CPU performance, even where memory IO limited, is more likely limited by how well it can prefetch memory and how many in-flight reads & writes it can do. A straight memcpy benchmark might be part of this suite, but that'd also be basically the only workload where a CPU core (or multiple) could come anywhere close to hitting 400GB/s of bandwidth. Otherwise memory _latency_ will be the bigger memory IO limitation for the CPUs, and that likely hasn't drastically changed with the new M1 Pros & Max's (there may be some cache layout changes which would shift some of the numbers, but DRAM latency is likely unchanged) For an example of this in a different product: https://www.anandtech.com/show/15578/cloud-clash-amazon-graviton2-arm-against-intel-and-amd/3 https://www.anandtech.com/show/15578/cloud-clash-amazon-grav... A single graviton2 CPU can "only" achieve 18-36GB/s of memory bandwidth even though the package in total can hit 200GB/s
- _kbh_ 5y agoA single firestorm (performance) m1 core in the m1 Mac mini can sustain ~60Gb/s read from ram. I have no doubt that 8 of them could come close to if not entirely saturate 400GB/s. https://www.anandtech.com/show/16252/mac-mini-apple-m1-tested https://www.anandtech.com/show/16252/mac-mini-apple-m1-teste...
- ben-schaaf 5y agoAs people stated in the announcement post, a high memory bandwidth doesn't really benefit CPUs much. That's also why AMD and Intel don't have high memory bandwidth on their cores, because it doesn't really help performance. Where it is beneficial is the GPU; for comparison AMD and Nvidia cards often exceed 400GB/s. A desktop RTX 3090 has 900GB/s.
- cormacrelf 5y agoI don’t think you could even top 150GB/s sustained if you ran memcpy on all ten threads at once. (Though that would be 300 total.)
- cormacrelf 5y agoThis turned out to be correct, look at https://www.anandtech.com/show/17024/apple-m1-max-performance-review/2 https://www.anandtech.com/show/17024/apple-m1-max-performanc... -- the "sustained" measurement is on the right hand side of the graph. It hits 243GB/s bandwidth with 8+2 threads, 224 with just the 8 performance cores.
- girvo 5y agoIntel is definitely looking at HBM for Sapphire Rapids though. https://www.anandtech.com/show/16795/intel-to-launch-next-gen-sapphire-rapids-xeon-with-high-bandwidth-memory https://www.anandtech.com/show/16795/intel-to-launch-next-ge...