3 ms·
The answer to these sort of questions (OP) is always the same: it was done because it made all the sense in the world to do. Most likely, 150 GB/s -> 200 GB/s
by reroute22 3y ago
The answer to these sort of questions (OP) is always the same: it was done because it made all the sense in the world to do.
Most likely, 150 GB/s -> 200 GB/s results in a fairly small improvement (when paired with a processor of M2 Pro / M3 Pro overall capability) only in a fairly small and specific subset of GPU applications. In particular, it's pretty much the matter of fact that that extra bandwidth achieves nothing in CPU-bound applications. It's also a matter of fact that only some of the GPU applications benefit, I just can not attest to exactly how big (or rather, small) and important that subset is.
With every new process node nowadays the following happens:
1. The cost per unit of area increases substantially. Decades ago the cost per unit of area was practically staying the same, resulting in 2x smaller node being 2x cheaper for the same design as the same design would take up 2x less area. Not anymore. The cost per unit of area is higher, therefore, if a portion of the design doesn't shrink much, it's actually more expensive on newer node. It takes a large shrink for the design to get cheaper or even merely stay the same in per-unit manufacturing cost.
2. IO shrinks very little. 4nm -> 3nm resulted in only 10% IO shrink. 1.25x SRAM, and 1.7x logic.
3. DRAM bandwidth is just a product of bus width * DRAM frequency, where bus width is really just a number-of-DRAM-controllers * 32.
4. DRAM controllers is IO. It barely shrinks in area going 4nm -> 3nm. But 3nm is more expensive per unit of area to manufacture. Therefore, DRAM controllers of the same design and the same count and the same bandwidth cost more money now on 3nm.
Most likely, that very marginal and situational performance benefit in a subset of GPU applications that M2 Pro saw going from 150 GB/s to 200 GB/s was still large enough to justify the relatively low (on 5nm) cost of 8 DRAM controllers (in traditional 32-bit-bus-per-controller terms). On M3 Pro that performance gain probably just dropped below the threshold and became unjustifiable against the increased cost of DRAM controllers, and the number of DRAM controllers was reduced to 6.