3 ms·
Yep, I've heard that actually something like 80% of the energy of the CPU is used by the front-end! Beyond decoding, you also have all the energy spent on score
by sakras 3y ago
Yep, I've heard that actually something like 80% of the energy of the CPU is used by the front-end! Beyond decoding, you also have all the energy spent on scoreboarding data dependencies and reordering the instructions and scheduling to execution ports. If you can do that 8x less as in the case for SIMD, that's a huge win!
- jsheard 3y agoMemory traffic too, assuming the SIMD code is properly optimized, you can slam nicely aligned blocks of memory straight into a SIMD register rather than loading each element one at a time.
- adgjlsfhk1 3y agothis is less important than you would expect because memory gets fetched into caches 512 bits at a time. As such, the difference in power use is only the cost of loading from L1 cache (since the first instruction will bring the line to l1)
- adgjlsfhk1 3y agoThat's definitely wrong (unless you use an incredibly expansive definition of "frontend"). The simplest way to prove that it's wrong is that heavily vectorized workloads can easily cause a 30% increase in power consumption over scalar workloads even though the instruction density is similar.
- totallyabstract 3y agoWouldn't that support the idea that most energy consumption is in decoding? If you're getting 2x, 4x, 8x ect as much value computation per instruction and yet only a 30% increase in power then clearly most the power is not used by computing the values.
- adgjlsfhk1 3y agono, because there's a lot the CPU does that is neither decoding nor execution. There's also caches, register renaming, branch prediction, inter core communication for atomics, and a dozen other things.
- totallyabstract 3y agoSorry poor terminology use on my part. I mean more broadly that most energy is used on frontend and middle end, rather than backend and that this is what vectorisation improves in regards to energy consumption. Register renaming and branch prediction energy consumption should be improved in the same factor as decoding. Caching probably less so (depending if we are talking instruction, data or combined). I don’t think inter-core communication is too relevant when comparing vectored and non-vectored on a single core, but definitely would be when batching across multiple cores.
- derefr 3y agoI wonder how efficient a modern x86-64 CPU could be — on a scale from “CISC” to “single-purpose ASIC” — if you could compile and load microcode directly into cache lines, pre-baking all of this just-in-time silicon allocation decision-making into a static microcode loop, so that the frontend can shut off entirely. (Presumably doing this would only be possible in the context of a single-threaded uninterruptible unikernel workload with system-management-mode functions disabled — but I’m sure a lot of “one powerful single-core SoC”-type embedded systems would be happy to make that trade off!)
- Out_of_Characte 3y agoForgive me if i'm misunderstanding, CPU's have micro-op caches that will bypass decoding instructions altogether if its in that cache. This means that variable length instructions will be decoded to a micro-op in the cpu and will get reused for as long as you have cache hits. Which does mean you"ll have ghost performance penalties if your code/gcc doesn't respect the size and limitations of hidden caches. Otherwise you can pretend that any instruction length gets shortened to its micro-op.
- derefr 3y agoI think I was imagining a useful operation level lower than the regular uOp — one less akin to programming a RISC processor, and more akin to the direct control over individual RTL signal lines — similar to the "instruction word" of signal-line states directly encoded by a row of the op-decoder PLA of a 6502. Right now, AFAICT, even when executing a uOp stream from L0 cache, several "planning" systems are still active — juggling caches around and deciding routing between them; renaming registers; switching between power "license" states; etc. None of these decisions are explicit even in uCode for modern CPUs, so they have to be made, over and over again, even when running from uOp cache. Which means the silicon that's making these decisions can't ever go dark, even when running from uOp cache. A "nanocode" would be a version of microcode that burns in all these "planning"-system decisions — and which can thereby put all the "planning" silicon to sleep. There would be "nanoOps" for explicitly wiring registers to other registers or cache buffers, for adding precise numbers of delay cycles, etc. And these would all happen at Nth-of-a-cycle-precise points in the execution stream, indicated by either entire nOp instructions, or pragma bits on nOp instructions. (Given this, to usefully function, these "nanoOps" would either likely need to be bytecode and be decoded+executed at some ridiculously high frequency relative to the regular CPU clock — or they would need to be some arcane parallelized packing of what was originally a serial stream of uOps, such that you get VLIW nOps where all the pragmas you want to have happen for the next CPU cycle are indicated in the instruction-word at once as individual signal-line bits. Just like the 6502 PLA output "instruction word", actually!) I know this probably sounds like nonsense — the instruction stream would be so much more bloated than microcode that it probably wouldn't be worth it. But that assumes an instruction stream that needs to live in RAM and get shuttled through layers of caches to reach the decoder. But what if the instruction stream could be loaded into a reserved SRAM area within the CPU itself — an SRAM that would effectively act as EEPROM (in that it would be loaded from CPU NVRAM on boot); and which you could directly drop the instruction pointer into (in nCode decode mode), skipping RAM+caching entirely? I mention this, because I've always had a vague hypothesis that Intel x86 CPUs specifically already have something akin to this implementation of a "nanocode" + reserved SRAM area that holds some of it — specifically for the purpose of programming custom dynamic RTL for instructions after CPU release to hotfix CPU errata. It's something they had to learn the necessity of over and over again, after releasing CPUs with bugged instructions, all the way back to the FDIV bug.
- twoodfin 3y agoGiven the massive # of instructions that an M3 core can keep in flight, I suspect that Apple’s CPU engineers could (but won’t) write some really interesting papers on this challenge.