4 ms·
Isn't it misleading to just add up the output width of all SIMD ALU pipelines and call the sum "datapath width", because you can't freely mix and match when the
by crest 2y ago
Isn't it misleading to just add up the output width of all SIMD ALU pipelines and call the sum "datapath width", because you can't freely mix and match when the available ALUs pipelines determine what operations you can compute at full width?
- adrian_b 2y agoYou are right that in most CPUs the 3 or 4 vector execution units are not completely identical. Therefore some operations may use the entire datapath width, while others may use only a fraction, e.g. only a half or only two thirds or only three quarters. However you cannot really discuss these details without listing all such instructions, i.e. reproducing the tables from the Intel or AMD optimization guides of from Agner Fog's optimization documents. For the purpose of this discussion thread, these details are not really relevant, because for Intel and AMD the classification of the instructions is mostly the same, i.e. the cheap instructions, like addition operations, can be executed in all execution units, using the entire datapath width, while certain more expensive operations, like multiplication/division/square root/shuffle may be done only in a subset of the execution units, so they can use only a fraction of the datapath width (but when possible they will be coupled with simple instructions using the remainder of the datapath, maintaining a total throughput equal with the datapath width). Because most instructions are classified by cost in the same way by AMD and Intel, the throughput ratio between AMD and Intel is typically the same both for instructions using the full datapath width and for those using only a fraction. Like I have said, with very few exceptions (including FMUL/FMA/LD/ST), the throughput for 512-bit instructions has been the same for Zen 4 and the Intel CPUs with AVX-512 support, as determined by the common 1024-bit datapath width, including for the instructions that could use only a half-width 512-bit datapath.
- dzaima 2y agoWouldn't it be 1536-bit for 2 256-bit FMA/cycle, with FMA taking 3 inputs? (applies equally to both so doesn't change anything materially; And even goes back to Haswell, which too is capable of 2 256-bit FMA/cycle)
- adrian_b 2y agoThat is why I have written the throughput "for results", to clarify the meaning (the throughput for output results is determined by the number of execution units; it does not depend on the number of input operands). The vector register file has a number of read and write ports, e.g. 10 x 512-bit read ports for recent AMD CPUs (i.e. 10 ports can provide the input operands for 2 x FMA + 2 FADD, when no store instructions are done simultaneously). So a detailed explanation of the "datapath widths", would have to take into account the number of read and write ports, because some combinations of instructions cannot be executed simultaneously, even when there are available execution units, because the paths between the register file and the execution units are occupied. Even more complicately, some combinations of instructions that would be prohibited by not having enough register read and write ports, can actually be done simultaneously because there are bypass paths between the execution units that allow the sharing of some input operands or the direct use of output operands as input operands, without passing through the register file. The structure of the Intel vector execution units, with 3 x 256-bit execution units, 2 of which can do FMA, goes indeed back to Haswell, as you say. The Lion Cove core launched in 2024 is the first Intel core that uses the enhanced structure used by AMD Zen for many years, with 4 execution units, where 2 can do FMA/FMUL, but all 4 can do FADD. Starting with the Skylake Server CPUs, the Intel CPUs with AVX-512 support retain the Haswell structure when executing 256-bit or narrower instructions, but when executing 512-bit instructions, 2 x 256-bit execution units are paired to make a 512-bit execution unit, while the third 256-bit execution unit is paired with an otherwise unused 256-bit execution unit to make a second 512-bit execution unit. Of these 2 x 512-bit execution units, only one can do FMA. Certain Intel SKUs add a second 512-bit FMA unit, so in those both 512-bit execution units can do FMA (this fact is mentioned where applicable in the CPU descriptions from the Intel Ark site).
- dzaima 2y agoSo the 1024-bit number is the number of vector output bits per cycle, i.e. 2×FMA+2×FADD = (2+2)×256-bit? Is the term "datapath width" used for that anywhere else? (I guess you've prefixed that with "total " in some places, which makes much more sense)