3 ms·
As far as I understand, with micro uop fusion the number of retired uops should be higher than the issued uops.
by pbsd 9y ago
As far as I understand, with micro uop fusion the number of retired uops should be higher than the issued uops.
- jnordwick 9y ago(Wrong stuff deleted) Edit: According the this pdf, macro/micro fusion happens very early on macro decoding, but I'm still not sure how it counts retired uops: http://pc.watch.impress.co.jp/video/pcw/docs/601/161/p21.pdf http://pc.watch.impress.co.jp/video/pcw/docs/601/161/p21.pdf And does this mean micro fusion cannot occur when you use the separate load/op/store instructions?
- pbsd 9y agoAccording to Agner when the uops go back to to the reorder buffer to get retired they are still treated as fused. However, if you look at the descriptions of the events [1, 2] a retired fused uop is counted as 2 retired uops. And yes, if you do mov r, m / add r, 1 / mov m, r instead of add m, 1 no such fusion happens. This is specific to CISCy instructions that do more than one thing at the same time. [1] https://software.intel.com/sites/products/documentation/doclib/stdxe/2013SP1/amplifierxe/pmn/events/uops_retired.html https://software.intel.com/sites/products/documentation/docl... [2] https://software.intel.com/sites/products/documentation/doclib/stdxe/2013SP1/amplifierxe/pmw_sp/events/uops_issued.html https://software.intel.com/sites/products/documentation/docl...
- BeeOnRope 9y agoMost recent archs have both fused and unfused counters for retired uops: UOPS_RETIRED.RETIRE_SLOTS is fused domain, while UOPS_RETIRED.ALL/UOPS_RETIRED.ANY and friends are unfused domain.
- BeeOnRope 9y agoAt least on Skylake the "retired uops" counters are counted in the "unfused" domain (i.e., an instruction that turns into 2 uops that get micro-fused counts as 2 not 1). I believe earlier archs had a similar counter with the opposite behavior (i.e., they counted in the unfused domain). Micro-fusion can happen only for stores, or when instructions have at least one memory source argument. On x86 this means either load-op instructions, which have a memory _source_ argument like: add eax, [rdx] ... or the so-called RMW instructions, which use memory as source and destination (sometimes with another non-memory source as in the first example) like: add [rdx], 1 ; memory source/dest, immediate 2nd source dec [rdx] ; memory source/dest only In these cases fusion happens because the memory access uop(s) can be grouped together with the ALU op, because the memory source is "anonymous" (doesn't get put into a visible register, so no-one can later use it - meaning it doesn't need a named register, hence renaming resources). The decomposition of the first add example into separate instructions looks like: mov ecx, [rdx] add eax, ecx This doesn't fit the pattern above so can't be micro-fused. in particular, the use of ecx as a temporary register here means you couldn't really fuse it: the value in ecx is live after this code segment, so it needs a named physical register. The third case of fusion is stores: mov [rdx], eax Internally this uses two uops (one address-generation and one store-data), but they are generally fused so it executes as one uop in the fused domain. This isn't super interesting since there aren't really alternatives for stores, they always have the above form, so micro-fusion is kind of less interesting (unlike the load-op ALU stuff, where you could use memory source, or not). Finally we get to RMW. This is logically a load-op combined with a store. So you have 4 unfused uops in total! Here we potentially get both types of fusion: the load-op fusion, and the store fusion, reducing the number of fused uops to 2. So to actually answer your question: if you implement a RMW has a separate load, ALU op, store like: ; instead of dec DWORD [rdx] mov eax, [rdx] dec eax mov [rdx], eax You would lose _half_ of the available fusion: the load-op part is now two instructions, so you don't get the fusion there, but you still get the store fusion. So you have 3 fused-domain uops and 4 unfused, versus 2 and 4 for the RMW case. On the other hand, for cases where you aren't writing the result back immediately to the same memory location, RMW doesn't apply, and you can get maximum fusion by using the "load-op" for of the instruction.