3 ms·
SIMD (and especially the modern matrix extensions) can use as much bandwidth you can throw at it.
by fooker 11d ago
SIMD (and especially the modern matrix extensions) can use as much bandwidth you can throw at it.
- b112 11d agoWithout further clarification, that statement seems impossible. "As much" being unbounded and all. You should expand what you mean.
- fooker 11d agoThis will blow your mind, but it actually is pretty close to being unbounded. :) Consider the 'MMA N matrices' primitive modern CPUs are starting to support. For the current generation of CPUs, N is a constant like 16 or 32, but there's nothing preventing it from being 1024 or larger if we have more memory bandwidth. All this with a single instruction.
- articulatepang 11d agoSurely something prevents it being 1 quadrillion bits per instruction? Since that’s well within “unbounded”.
- pixl97 10d agoMost likely the chip running at the core temperature of the sun. We'll have to figure out how to read and right to the surface of a black hole to get speeds that high.
- imtringued 11d agoNah his point is a bit simpler. Vector units can process a limited amount of data per time unit. Memory can load a limited amount of data per time unit. If you have infinite memory bandwidth you just move the bottleneck back to compute so both have to grow simultaneously in lockstep. What you should have said is that CPUs have so much compute headroom for matrix vector multiplication that simply adding more memory bandwidth would make them faster so every improvement in memory bandwidth is welcome.
- fooker 11d agoAgreed. The "move the bottleneck back to compute" bit is changing rapidly though. The first time a major hardware company ships a PIM chip, you can push for a few orders of magnitude more data through without being compute bound.