3 ms·
> @Intel: how about thinking about 10GHz single core CPUs? They tried that -- the Pentium 4. The key design goal was to push clock speed as high as possible, a
by cfallin 10y ago
> @Intel: how about thinking about 10GHz single core CPUs?
They tried that -- the Pentium 4. The key design goal was to push clock speed as high as possible, and they used some crazy tricks, like 30-some-stage pipelines, a really long instruction scheduling loop with a pretty long lookahead, a double-pumped ALU with staggered 16-bit half-adds, etc. Fascinating from a microarchitecture perspective, but the big lesson was that high clock speeds sacrifice efficiency. The tricks needed to get there cause a lot of performance outliers/bad cases too -- e.g. the instruction scheduling replay system on lookahead conflicts was notorious for "tornados" which would kill IPC. And the branch prediction of the day was somewhat suboptimal for the pipeline lengths involved.
> Why are we stuck in 2006 era CPU singe core performance?
We aren't! The core microarchitecture teams at Intel and their competitors have made a lot of incremental progress -- Intel gets about 15% single-thread performance per generation, for example. No one is evilly scheming and holding back from turning a knob higher. There's just a lot of really hard engineering. Clock speed has topped out due to power limitations so we're at the point of looking for better branch prediction algorithms, cache replacement algorithms, and lots of little tricks everywhere to optimize bad cases. It's hard work (this was my job for a bit).
I'd recommend looking at, e.g., the proceedings of ISCA and MICRO conferences in the 2000-2006 timeframe -- this was when the industry and associated academia figured out that chasing clock speed was a losing battle after a certain point.
- sirsar 10y agoWhat does it mean to have a "double pumped" ALU? Google is failing me.
- cfallin 10y agoThe ALU performs an operation on half of the machine word (16 of the 32 bits) each half-cycle, i.e., one on the rising edge and one on the falling edge of the clock. Basically they split the carry chain across two (half-)pipe stages and then run it twice as fast. See Hinton et al., "The microarchitecture of the Pentium 4 processor" [1] for all the nifty details -- pp 8-9, and Fig 7 in particular. [1] http://www.ecs.umass.edu/ece/koren/ece568/papers/Pentium4.pdf http://www.ecs.umass.edu/ece/koren/ece568/papers/Pentium4.pd...
- sirsar 10y agoMakes sense, thanks!
- userbinator 10y agoThat also makes for some very interesting timings, where the instruction runs faster in the case that there's no carry between the two halves. From section 2 of this: https://gmplib.org/~tege/x86-timing.pdf https://gmplib.org/~tege/x86-timing.pdf "Pentium F0-F2 can sustain 3 add r, i per cycle for -32768 <= i <= 32767, but for larger immediate operands it can sustain only about 3/2 per cycle."
- tcas 10y agoIt runs at 2x the clock rate as the main chip. So if the chip is running at 2Ghz, the ALU is running at 4Ghz. While this might not make sense at the surface (the next stage of the pipeline won't be ready for the data a half clock cycle earlier), you can minimize logic area by doing 2 quick 16 bit operations. Additionally, for some operations that have dependent instructions, for example a super scalar processor executing two integer instructions at once, if you have an ADD that depends on another ADD, you can forward the result from the first cycle on the first ALU to the second cycle on the second, although I'm not sure if this is actually done. Another case where this is useful is if the result of an ADD operation needs to be used on a LOAD instruction earlier, you can forward the result a half clock cycle sooner.
- deleted 10y ago[deleted]
- TillE 10y agoRight, even single CPU performance of mainstream Intel chips has more than doubled in 7 years (Nehalem to Skylake). Many important things which affect performance have been stagnant for a very long time, especially cache size and RAM latency. But there's been steady progress in raw CPU power, coupled with a significant reduction in TDP.
- gpderetta 10y agoDo you have a reference for the doubling? How much does the boost from AVX affect the score? I'm asking because I recently "upgraded" my desktop to a cheap westmere 6 core xeon [1] as in most benchmarks it is competitive with recent i7s. [1] x5670 @3ghz, but should easily overclock well into the 4ghz range. Edit: autocorrect
- robocat 10y agoPrimatelabs provide benchmarks for: 32 bit 1 CPU, 32 bit threaded, 64 bit 1 CPU, 64 bit threaded. https://browser.primatelabs.com https://browser.primatelabs.com Useful information if your workload is single threaded. Note that their usage of the word browser is nothing to do with www browsers.
- gpderetta 10y agoOh, Geekbench which can be 'won' by implementing a single encryption instruction in hardware. Still, even looking at the scores for less broken sub benchmarks, like (possibly) lua and dijkstra show a %50 increase in performance from Westmere to Skylake for the same frequency, which is more than I was expecting. And skylake should potentially clock much higher although I believe the top is currently still 4ghz. Of course for anything that can take advantage AVX2 (or AVX3 for xeons) skylake would smoke the old xeon. edit: reword