4 ms·
Frequency tends to be somewhat orthogonal to instructions per watt. (It's easish to build hyperpipelined CPUs which have horrendous pipeline stall latency chara
by disposablezero 10y ago
Frequency tends to be somewhat orthogonal to instructions per watt. (It's easish to build hyperpipelined CPUs which have horrendous pipeline stall latency characteristics.) And CMOS tends to dissipate power (most as heat) proportional to the square of the frequency because zero crossing in CMOS logic is a virtual short-circuit. MIPS (the former Hennessy shop), SPARC and other RISC shops tended to ignore more/less the retail consumerization of the CISC MHz wars to bet on architecture fundamentals.
I highly recommend taking a digital logic and computer architecture class in a CS program to understand how pipelines/instruction units are laid out (microcoding, branch prediction, pipeline stalls, macro ISA, etc.), what makes them fast/slow and challenges to implementation.
- wyager 10y agoOne thing that surprised me while taking a processor design course is that pipelining is actually old technology; modern processors are essentially partially nonlinear graph-reduction machines. Your average CPU core might have 6-10 different logic units (multiple ALUs, FPUs, MUs, etc.) possibly shared across 2 "hyperthreaded" virtual cores. This means that if the incoming assembly code is fairly linear (or has no mispredicted branches) and has some partially independent computation, you could feasibly sustain 5+ full instructions per cycle. As a side note, the Tomasulo algorithm is a huge pain in the ass to do in hardware. Props to the processor engineers who are implementing modern OOO execution hardware in whatever awful proprietary HDL/IDE combo your company makes you use.
- dom0 10y agoOh yes they do. A really excellent example here are ARX (add rotate xor) algorithms in cryptography. Take BLAKE2 for example; the reference C code, compiled to plain x86-64 assembly mainly consists of long rows of addq/xorq/rorq instructions. A modern CPU manages to schedule these so efficiently that the performance difference to a hand-coded SSE or AVX version practically never matters. If you take a peek at Haswell's port map (eg. http://images.anandtech.com/reviews/cpu/intel/Haswell/Architecture/haswellexec.png http://images.anandtech.com/reviews/cpu/intel/Haswell/Archit... ) you'll see that there are four out of eight ports who can do integer arx operations. Since memory and addr ops got their own ports this means that even the straight addq/xorq/rorq code can almost fully utilize the core. This is no accident, though, the algorithm was designed to take advantage of designs like this :)