5 ms·
Yep. RISC was interesting when gate budgets for CPU pipelines were seriously limited. It was interesting because before RISC the industry had been merrily spend
by abainbridge 3y ago
Yep. RISC was interesting when gate budgets for CPU pipelines were seriously limited. It was interesting because before RISC the industry had been merrily spending the gate budget increase on adding lots of use-specific instructions. The RISC people pointed out that if you removed support for all the fancy instructions you had enough gate budget for the ALU to be nicely pipelined, and then you could wind up the clock rate greatly and this was worth much more than the fancy instructions.
For decades now we've had enough gate budget to have nicely pipelined designs with complex instruction sets, so that's what everyone does. RISC solves a problem that no longer exists.
- thesz 3y agoYou are not quite right about pipelined design being faster. At least, not without substantial effort. https://en.wikipedia.org/wiki/R2000_microprocessor https://en.wikipedia.org/wiki/R2000_microprocessor "The R2000 is a 32-bit microprocessor chip set developed by MIPS Computer Systems that implemented the MIPS I instruction set architecture (ISA)..." "The R2000 was available in 8.3, 12.5 and 15 MHz grades..." https://en.wikipedia.org/wiki/I386 https://en.wikipedia.org/wiki/I386 "The Intel 386, originally released as 80386 and later renamed i386, is a 32-bit microprocessor introduced in 1985..." "Max. CPU clock rate: 12.5 MHz to 40 MHz" As you can see, 80386 was released a year earlier than R2000 and was about 1.5 times faster than MIPS implementation from the start. The critical path is, usually, in addition/subtraction, which should be complete in one cycle in both 80386 and in R2000. To pipeline addition you need a superpipelined CPU, one that has several stages for computation. Even seemingly simple computation of condition codes can make clock cycle 10% longer (SPARC vs MIPS) if your CPU is just simply pipelined. BTW, some Pentiums did computed 32-bit addition in two cycles, all in name of higher clock frequencies.
- abainbridge 3y agoInteresting. An R2000 did run programs faster than a 80386, right? This was a few years before my time. From a quick google now, it looks like the R2000 was about 3x better than the 80386 at Dhrystone MIPS/MHz. I guess an accurate comparison of how the R2000 and 80386 spent they gate budget and what they got in return would involve a lot of detail. I remember my compsci professor giving us the computer architecture course in about 1997, and he dispaired at how all the clever RISC stuff in the Patterson and Hennessey seemed irrelevant when Intel could just throw money at the implementation (and fab, I guess) and produce competitive chips despite their (allegedly) inferior architecture.
- thesz 3y agoMy point is that you cannot get design much faster in terms of clock frequency by just pipelining. Pipeline unrolls state machine and overlaps different executions of the state machines. But the bottleneck, which is addition, is there in all designs and you need additional effort to break it. (also MIPS has [i]ntelocked [p]ipeline [s]tages - that "IPS" in MIPS; I implemented it, I know - exception in execution should inform other stages about failure) By the 1997 Intel has already bought Elbrus II design team, lead by Pentkovski [1]. That Pentkovski guy made Elbrus 2 a superscalar CPU with a stack machine front-end. E.g., Elbrus 2 executed stack operations in a superscalar fashion. You can entertain yourself by figuring out how complex or simple can that be. [1] https://en.wikipedia.org/wiki/Vladimir_Pentkovski https://en.wikipedia.org/wiki/Vladimir_Pentkovski So at the time your professor complained about Intel's inferior architecture being faster, that inferior architecture implementation has a translation unit inside it to translate x86 opcodes into superscalar-ready uops.
- peterfirefly 3y agoAddition was not the bottle neck for the 386. It had a FO4 delay of 80+ per clock. An adder is much faster. Maybe you meant that it was (one, just one!, of many of) the bottle neck(s) in an optimized implementation?
- thesz 3y ago> Addition was not the bottle neck for the 386. It is a bottleneck for MIPS, SPARC, Alpha and not for 386. How so?
- peterfirefly 3y agoThe 386 wastes so many FO4 gate delays on other things. I thought I made that extremely clear?
- thesz 3y agoCan you elaborate on where the delays came from?
- smcin 3y agoIt's not apples-to-apples to compare raw clock rates between semiconductor processes; Intel's 386 was intially fabbed on 1.5μ then shrunk to 1.0μ process (Intel CHMOS III and IV), whereas MIPS R2000 was 2.0μ, fabless and relied on Sierra, Toshiba, then in 1987 LSI, IDT and other licensees [0][1]. Back in the 1980s/90s/2000s, Intel was consistently a process generation or two ahead of competitors. That was one of their main sources of advantage. Just imagine if MIPS had been able to fab on Intel process. [0]: https://www.righto.com/2023/10/intel-386-die-versions.html https://www.righto.com/2023/10/intel-386-die-versions.html [1]: https://en.wikipedia.org/wiki/R2000_microprocessor https://en.wikipedia.org/wiki/R2000_microprocessor
- thesz 3y agoAnd you are also confirm that in order to have higher clock frequency you need more than just pipelining. Thank you. I also think that 1.5x difference in clock speeds cannot be directly attributed to the difference between node size (lambda): difference in lambdas 1.3(3)=2.0/1.5 at the introduction of the 80386 and R2000 is noticeably less than 1.47=12.5/8.5.
- smcin 3y agoSmaller transistors are faster, but the relationship between clock frequency and 1/feature size isn't necessarily linear like you're assuming. https://cs.stackexchange.com/questions/27875/moores-law-and-clock-speed https://cs.stackexchange.com/questions/27875/moores-law-and-...
- thesz 3y agoMy assumption is that speedup is less than lambda's ratio.
- bananabiscuit 3y agoIs there something about RISC that is still makes it better than CISC when it comes to per-watt performance? Seems like nobody has any success making an x86 processor that's as power efficient as ARM or RISC.
- card_zero 3y agoHave there been recent attempts? Maybe it's just, like, speciation, by this point in time.
- simne 3y ago> Is there something about RISC that is still makes it better than CISC when it comes to per-watt performance? CLASSIC CISC was micro-coded (for example, IBM S/360 have feature, you could make your custom microcode for compatibility with your inherited equipment, like IBM-1401 machines or IBM-7XXX series, or for other purposes), and RISC was with pipeline from birth. Second thing, as I understand, many CISC existed as multiple chips board or even as multiple boards, so have great losses on wires, but RISC appear in 1990s as one die immediately (only external cache added as additional IC), but I could mistake on this. > nobody has any success making an x86 processor that's as power efficient as ARM or RISC Rumors said, Intel Atom (essentially CMOS version of Pentium first generations) was very good in mobiles, but ARM far succeed it on software support of huge number of power saving features (modern ARM SOC allows to turn off near any part of chip any time and OS support this), and because of lack of software support, smartphones with Intel have poor time on battery. More or less official info said, that Intel made bad power conversion circuit, so Atom consumes too much in mode between deep sleep and full speed, but I don't believe them, as this is too obvious mistake for hardware developer.
- peterfirefly 3y ago> CLASSIC CISC was micro-coded Sometimes. Far from always. Some would have a complicated hardwired state machine. Some would have a complicated hardwired state machine and be pipelined. Some would have microcode and be pipelined (by flowing the microcode bits through the pipeline and of course dropping those that have already been used so less and less microcode bits survive at each stage).