4 ms·
https://images.anandtech.com/doci/13699/Ronak26.jpg https://images.anandtech.com/doci/13699/Ronak26.jpg https://www.anandtech.com/show/13699/intel-architecture
by hydroreadsstuff 7y ago
https://images.anandtech.com/doci/13699/Ronak26.jpg https://images.anandtech.com/doci/13699/Ronak26.jpg
https://www.anandtech.com/show/13699/intel-architecture-day-2018-core-future-hybrid-x86/2 https://www.anandtech.com/show/13699/intel-architecture-day-...
I believe this shows differences between FP and Integer units.
In order to achieve a certain performance goal you don't necessarily need another integer divider when you want a new adder/multiplier. So you add a slimmer unit instead.
I listed this in my original comment, because this is a giant can of worms for the compiler and decision-maker on where to execute what.
- tempguy9999 7y agoAh, thanks. The slide is interesting for extra reasons. With respect, I think you're misunderstanding. I thought you meant light/heavy versions of eg. adders, for some definition of light and heavy addition. I'm not an expert but... CPUs will put in extra execution units according to need (will typical code get faster with an extra X?) and cost. Shifters are typically very often used, and are simple. So are adders, though more complex. IIRC recent intel x64 will have several of of each[0]. Multipliers are less cheap so they have fewer (and often you can turn them into adds in certain cases such as progressive array lookups). Division is slow and very expensive in transistors, so they have 1 (division can often be turned into reciprocal multiplication anyway). Sqrt is even worse. And to repeat, I'm no expert and any corrections welcome. [0] <https://en.wikichip.org/wiki/intel/microarchitectures/coffee_lake> https://en.wikichip.org/wiki/intel/microarchitectures/coffee... If I'm reading this right, 2 shifters (2? I suppose they are fast so they are available soon after), 4 adders, 1 mult and 1 divider.
- deleted 7y ago[deleted]
- namibj 7y agoActually, barrel shifters are not that small. Added are significantly cheaper in terms of chip area.
- tempguy9999 7y agoSeriously?? It's the same bit pattern, err, shifted. Adders have got to carry, at each stage (ripple carry?). I am amazed, thanks.
- Robin_Message 7y agoThink of it this way: a barrel shifter has to be able to "carry" every bit to (potentially) every other bit.
- tempguy9999 7y agoYep, but also while aligned. Bits x and y are going to be the same distance apart, always, unless shifted off the end (where they can wrap or be lost). Thinking of this in the pub it seemed a butterfly thingy would be appropriate <https://en.wikipedia.org/wiki/Butterfly_network> https://en.wikipedia.org/wiki/Butterfly_network>. Further thinking suggested there'd be a ton of wires doing this, and perhaps it's the wiring that's taking up the silicon?
- namibj 7y agoIt's the multiplexers. Also, critical path is going towards limiting your speed: https://en.wikipedia.org/wiki/Barrel_shifter https://en.wikipedia.org/wiki/Barrel_shifter
- deleted 7y ago[deleted]
- hydroreadsstuff 7y agoRight, I mixed up GPUs and CPUs, here. Do ports (as in the picture) have independent pipelines, or do they execute certain pipeline stages of a big pipeline? I suppose, either way you can't issue to the same port in the same cycle. This paper sheds some light on how instructions are divied up between units on NVIDIA GPUs. http://www.stuffedcow.net/files/gpuarch-ispass2010.pdf http://www.stuffedcow.net/files/gpuarch-ispass2010.pdf Table IV. Notice that fp32 mul is in the SFU and SP, while others are not.
- tempguy9999 7y agoI'm the wrong guy to answer this but I believe the ports are available to one pipeline, so no independent pipelines, but instructions behind the 'head' can hop over the head if the head is stalled, if there are execution units free. Instructions can get reordered, over quite a wide window, something like 200 instructions (look up the reorder buffer, ROB, although there's another windows which affects this, something to do with retiring instructions). <https://en.wikipedia.org/wiki/Re-order_buffer> https://en.wikipedia.org/wiki/Re-order_buffer> Come to think of it, I don't know how the ports are used. I am entirely unqualified to answer this question :)