4 ms·
Also, there is the conceptual model of how an x86 computer works, and then how it actually works is another thing, for example I think the Sandy Bridge architec
by notaddicted 14y ago
Also, there is the conceptual model of how an x86 computer works, and then how it actually works is another thing, for example I think the Sandy Bridge architecture has 160 physical integer registers, if I'm interpreting this correctly: http://www.anandtech.com/show/3922/intels-sandy-bridge-architecture-exposed/3 http://www.anandtech.com/show/3922/intels-sandy-bridge-archi... .
- wtallis 14y agoRight, but the oddities and general paucity of the architectural registers makes it hard to actually put the large physical register file to use. Subtle differences in instruction selection and ordering can make a big difference in performance, eg. http://stackoverflow.com/questions/15349308/using-xmm0-register-and-memory-fetches-c-code-is-twice-as-fast-as-asm-only-u/15349403#15349403 http://stackoverflow.com/questions/15349308/using-xmm0-regis...
- ajross 14y agoThat's an interesting example, but it doesn't have anything to do with "putting the physical register file to use". Register renaming always helps as long as your algorithm is dependency-free. It's probably one of the single best general purpose optimizations available to a modern CPU design. The classic example of where register-poor ISAs hurt isn't about "subtle differences" at all. It's the fact that there are still only 8 (for i386) named registers, and so anything that needs to deal with a working set beyond that needs to do spill/fill to memory, and you can't "rename" memory accesses (though you can sort of cheat, as with the store forward optimization -- but that doesn't work nearly so well as renaming does).
- gchpaco 14y agoThe conventional x86 way of doing it was to spill to "the stack" and I'm told that their chips actually specially optimize that nowadays. But it's all very weird, deep magic in a lot of ways.
- brigade 14y agoThe only stack-specific special optimization that's done is fusing the decrement/increment of esp/rsp with the store µop. And that's done mainly since push/pop are one byte opcodes, unlike general load/store. Everything else is general memory optimizations that apply for everything like the aforementioned store forwarding. It's still expensive if the CPU can't use them (mismatched load/store size, incorrect speculation, etc.)
- brigade 14y agoThat's... not really subtle. That's like pipelining 101. Which is nearly 30 years old now... The only weird thing there is the compiler lucks out in generating a version that allows OoOE to save the programmer from themself.
- wtallis 14y agoIt's subtle compared to how that code would have been written if the assembler had access to a hundred registers. The compiler didn't get lucky there, it got complex in order to handle these situations. For a less trivial loop body that required a few more registers, it's harder to arrange things to allow register renaming to happen - much harder than just unrolling the loop to use several sets of registers. It should be clear that it's always harder to hint to the CPU that it can use a new register than it is to just explicitly use a new register. It's the same as why SSA form is useful for compilers.
- brigade 14y agoNo, it got lucky because it's emitting exactly what the author wrote; there's nothing complex going on. In fact, I'm sure the author is stunting where it got smart and hoisted val1+val2+val3+val4+val5+val6+val7 outside of the loop, and it can't do anything else (or even emit the hand-written version) because it would violate C++'s aliasing rules or the non-associativitiy of floating point. In effect, the compiler was asked to compile void foo(float &resVal, float &val1, float &val2, float &val3, float &val4, float &val5, float &val6, float &val7) { float tmp = val1+val2+val3+val4+val5+val6+val7; resVal += tmp; } whereas the author wrote asm doing void foo(float resVal, float val1, float val2, float val3, float val4, float val5, float val6, float val7) { resVal += val1; resVal += val2; resVal += val3; resVal += val4; resVal += val5; resVal += val6; resVal += val7; } Which isn't the same thing. And if the author unrolled it with a hundred registers, it would still be every bit as slow. And it should be obvious that it's slow not because it's messing up OoOE, it's slow because it's one incredibly long dependency chain with a stall between each add, due to it depending on the result of the previous add. OoOE happens to be able to fix the programmer's mistake for the first case, which is luck because the programmer obviously didn't intend for it to do so. But now I'm repeating myself...