3 ms·
If you want to build yourself an order of magnitude of how much performance is left on the table in CPUs, consider the following: * A single bit Adder with Car
by BenoitP 2y ago
If you want to build yourself an order of magnitude of how much performance is left on the table in CPUs, consider the following:
* A single bit Adder with Carry can be made with 9 transistors, and has to wait for 4 stages for the signal to propagate into. Then you have to stack 32 of them in series for a 32 bit carry. And wait for the signal to propagate.
* Then you first have to route the adding instruction, which is code treated as data. To do a single addition, you have a whole meta pipeline to run. With a lot of side-support: prefetching, cache coherency, loading, writing, instruction pointer, stack pointer, etc.
* Then you have to have the data available, or else you'll be waiting for RAM. That can be as long as 200 cycles, all the while you could be doing several additions in parallel; wasting 650 picoJoules, when the addition in itself is just a few picoJoules.
As silicon scaling laws hit their respective walls (latency, clock limits have been hit; bandwidth, memory sizes remain growing) the only way we'll be able to eek out more performance is to go ASIC on a lot of stable code.