3 ms·
> L1 cache is .. what.. 2 cycles? On Zen 4 CPU, I believe the typical latency of L1D is 4 cycles. However, if you (or your compiler) write AVX code which adds
by Const-me 1y ago
> L1 cache is .. what.. 2 cycles?
On Zen 4 CPU, I believe the typical latency of L1D is 4 cycles. However, if you (or your compiler) write AVX code which adds these floats, will still bottleneck on memory even if both inputs are in L1D cache. Each Zen 4 core can sustain two vaddps instructions per cycle, two loads per cycle, and one store per cycle. Due to the load and store bottlenecks, that kernel will only do one vaddps per cycle i.e. will waste 50% of theoretically available compute power.
> Intel architects were privately telling us: "compute doesn't matter any more; it's all about memory speed"
To be fair, that’s only true for automatically vectorized code, kernels like in the GP’s example. With sufficient efforts spent on software development, for some practical problems it’s possible to write codes which do saturate compute.
An example of such problem is multiplication of large matrices. A carefully written manually vectorized implementation should bottleneck on compute not memory, because theoretically required memory bandwidth scales as N^2, while theoretically required FLOPs scale as N^3 where N is size of the matrix.
That’s precisely what many BLAS libraries are doing under the hood. For the same reason GPU vendors report ridiculously high numbers of theoretical TFlops when multiplying low precision matrices with these special AI blocks, wmma/mfma instructions on AMD, tensor cores on nVidia.