7 ms·
Comparing a RISC and a CISC with similar hardware organization (1991)
- jabl 1y agoNeeds a [1991] tag. Needless to say, in the 34 years since that article was published, a lot has changed. Thanks to massively increased transistor budgets, a more complex decoder with accompanying microcode ROM that might have been a big detriment in 1991 would today be a small speck of dust on the processor floor plan. At the same time, memory access performance hasn't increased to the same extent as compute performance, thus putting a relatively bigger emphasis on code density. All this being said, RISC "won" in the sense that many RISC principles have become the "standard" principles of designing an ISA. Still, choosing "RISC purity" over code density is arguably the wrong choice. Contemporary high performance RISC architectures (ARMv9, say) are very un-RISC in the sense of having a zillion different instructions, somewhat complex addressing modes, and so forth.
- drob518 1y agoExactly. If you’re going to design a new ISA, you’d be foolish to make it a classical CISC design and would definitely choose RISC (e.g., RISC-V). But if you have a CISC ISA and you want to keep it running fast, then a virtually unlimited transistor budget allows you to create a sophisticated decoder that dispatches micro-ops to a RISC-like core and bridge the gap. That paper really take me back to working on PA-RISC designs at HP during that timeframe.
- hajile 1y agoThis seems to have its own issues and the proof is in the final cores. ARM’s entire gross profits are about half of AMD’s R&D budget, but ARM cores have soundly beat x86 in IPC for years now (since around A78) and the most recent generations seem to be beating them in total performance, perf/watt, and core size. We now have all three of the big ARM-based cores (Apple, ARM, and Qualcomm) beating x86. Apple you could maybe write off as unlimited money, but all three isn’t just coincidence. If that weren’t enough, ARM designers are releasing new cores every year instead of every other year meaning they are doing around twice as many layout and validations despite the massively lower budget. Before I get the “ARM only makes cores” excuses, I’d note that ARM announced that they’ve been working on their own server chips and that work is obviously having to fit within their same (comparatively tiny) budget. It seems fairly obvious. Spring legacy garbage drives up complexity and cost (eg, ARM reduced A-series decoder size by 75% when they dropped 32-bit mode which was still way less complex than x86). This complexity drives up development time and cost. It also drives up validation cost and time. There’s also a physical cost. Large, high frequency uop caches and cache controllers are better than just decoders on x86, but worse than not needing them at all is better still. Likewise, you hear crazy stuff like the x86 overly-strict memory model not mattering because you can speculate it away. That speculation means more complexity, more power, and more area. Once you’re done with enough of these work-around, you get a chip that is technically as fast, but it cost more to design, costs more to validate, costs more to fab, costs more to buy, costs more to operate, and carries an opportunity cost from taking so much longer to get to market.
- drob518 1y agoI think you’re just making my point.
- ndesaulniers 1y ago> ARM reduced A-series decoder size by 75% when they dropped 32-bit mode Interesting. Got a source for that?
- hajile 1y agohttps://fuse.wikichip.org/news/6853/arm-introduces-the-cortex-a715/ https://fuse.wikichip.org/news/6853/arm-introduces-the-corte... It was one of the biggest features of A715 going by ARM’s slides.
- ajross 1y ago> you’d be foolish to make it a classical CISC design and would definitely choose RISC I think that's arguable, honestly. Or if not it hinges on quibbling over "classic". There is a lot of code density advantage to special-case CISCy instructions: multiply+add and multiplex are obvious one in the compute world, which need to take three inputs and don't fit within classic ALU pipeline models. (You either need to wire 50% more register wires into every ALU and add an extra read port for the register store, or have a decode stage that recognizes the "special" instructions to route them to a special unit -- very "Complex" for a RISC instruction). But also just x86 CALL/RET, which combine arithmetic on the stack pointer, computation of a return address and store/load of the results to memory, are a big win (well, where not disallowed due to spectre/meltdown mitigations). ARM32 has its ldm/stm instructions which are big code size advantages too. Hardware-managed stack management a-la SPARC and ia64 was also a win for similar reasons, and still exists in a few areas (Xtensa has a similar register window design and is a dominant player in the DSP space). The idea of making all access to registers and memory be cleanly symmetric is obviously good for a very constrained chip (and its designers). But actual code in the real world makes very asymmetric use of those resources to conform to oddball but very common use cases like "C function call" or "Matrix Inversion" and aiming your hardware at that isn't necessarily "bad" design either.
- drob518 1y agoI’m talking about a VAX-like system with large instructions, microcode-based, etc. In the same way that CISC adapted since 1990, RISC has also adapted to add “complex” instructions where they are justified (e.g. SIMD/vector, crypto, hashing, more fancy addressing modes, acceleration for tensor processing, etc.). Nothing is a pure play anymore, but I’d still argue that new designs are better off starting on the RISC side of the (now very muddled) line, rather than the CISC side.
- ajross 1y agoRight, that's "quibbling over 'classic'". You said "you'd be foolish to design CISC" meaning the hardware design paradigm of the late 1970's. I (and probably others) took it to mean the instruction set. Your definition would make a Zen5 or Arrow Lake box "RISC", which seems more confusing than helpful.
- hajile 1y agoRISC is more about the time each instruction takes rather than how many instructions because consistent timing reduces bubbles and complexity. In this sense, RISC has completely won. Complexity of new instructions in the main pipeline is very restricted by this limitation and ISAs like x86 break down complex instructions into multiple small instructions before pushing them through. ARMv9 still has very few instruction modes with far less complexity when you compare it with x86 or some other classic CISC ISA. > memory access performance hasn't increased to the same extent as compute performance, thus putting a relatively bigger emphasis on code density. The problem isn't RAM. The problem is that (generally speaking) cache is either big or fast. x86 was stuck at 32kb for a decade or so. Only recently have we seen larger 64kb caches included. Higher code density means more cache hits. This is the big reason to care about code density in modern CPUs. RISC-V shows that you can remain RISC and still have great code density. Despite arguably making some bad/terrible decisions about the C instructions, RISC-V still generally beats x86 in code density by a significant margin (and growing as they add new instructions for some common cases).
- buildbot 1y agoNot really. A RISC design can have very complex timing and pipelines, with instructions converted to uOPs (and fused to uOPs!) just like X86: https://dougallj.github.io/applecpu/firestorm.html https://dougallj.github.io/applecpu/firestorm.html Caches can be fast and very expensive (in area & power)! I have an HP PA-RISC 8900 with 768KB I&D caches. They are relatively fast and low latency, given the time-frame of their design. They also take up over half the die area.
- hajile 1y agoI don't know how this has anything to do with what I said. The original intent of uops in x86 was to break more complex instructions down into more simple instructions so the main pipeline wasn't super-variable. If you look at new designs like M-series (or even x86 designs), they try very hard to ensure each instruction/uop retires in a uniform number of cycles (I've read[0] that even some division is done in just two cycles) to keep the pipeline busy and reduce the amount of timings that have to be tracked through the system. There are certainly instructions that take multiple cycles, but those are going to take the longer secondary pipelines and if there is a hard dependency, they will cause bubbles and stalls.
- brucehoult 1y ago> thus putting a relatively bigger emphasis on code density. > choosing "RISC purity" over code density is arguably the wrong choice You appear to be under the incorrect impression that CISC code is more dense than RISC code. This seems to be a common belief, apparently based on the idea that a highly variable-length ISA can be Huffman encoded, with more common operations being given shorter opcodes. This turns out not to be the case with any common CISC ISA. Rather, the simpler less flexible operations are given shorter opcodes, and that is a very different thing. A lot of the 8 bit instructions in x86 are wasted on operations that are seldom or never used and that could, even in 1976, have safely been hidden in some secondary code page. The densest common 32 bit ISAs are Arm Thumb2 and RISC-V with the C extension. Both of them have two instruction lengths, 2 bytes and 4 bytes, as did many historical RISC or RISC-like machines including CDC6600 (15 bits and 30 bits), Cray 1, the first version of IBM 801, Berkeley RISC-II. The idea that RISC means only a single instruction length is historically true only for ISAs introduced in the brief period between 1985 (Arm, SPARC, MIPS) and 1992 (Alpha) out of the 60 year span of RISC-like design (CDC6600 1964, the fastest supercomputer of its time). And, as an outlier, Arm64 (2011), which I think will come to be recognised as a mistake -- they thought Amd64 was the competition they had to match for code density (and they did) but failed to anticipate RISC-V. In 64 bit, RISC-V is by far the densest ISA. > Contemporary high performance RISC architectures (ARMv9, say) are very un-RISC in the sense of having a zillion different instructions, somewhat complex addressing modes, and so forth. Yes, ARMv8/9-A is complex. However there is no evidence that it is higher performance than RISC-V in comparable µarches and process nodes. On the contrary, other than their lack of SIMD SiFive's U74 and P550 are faster than Arm's A53/A55 and A72, respectively. This appears to continue for more recent cores, but we don't yet have purchasable hardware to prove it. That should change in 2026, with at least Tenstorrent shipping RISC-V equivalent to Apple's M1.
- clausecker 1y agoARM64 has a trick up its sleeve: many instructions that would be longer on other architecturea are instead split into easily recognisable pairs on ARM64. This allows for simple inplementations to pretend it's fixed length while more complex ones can pretend it's variable length. SVE takes this one step further with MOVPRFX, which can add be placed before almost all SVE instructions to supply masking and a third operand.
- wang_li 1y ago> All this being said, RISC "won" in the sense that many RISC principles have become the "standard" principles of designing an ISA. I disagree. Maybe many RISC chip design ideas may have taken over, but only because there are massive transistor budgets. I'd like to see a RISC chip that actually has a basic instruction set. As in, not having media instructions, SIMD instructions, crypto primitives, etc. If anything, Moore's Law won and the RISC v CISC battle became meaningless and they can just spend transistors to make every instruction faster if they care to.
- fweimer 1y agoARMv9 also has read-modify-write memory instructions, so does any usable RISC-V implementation. It turns out that LL-SC (which would avoid those) does not permit efficient implementations. (LL-SC does look like a rather desperate attempt to preserve a pure RISC register-register architecture.) I get the impression people believe that instruction density does not matter much in practice (at least for large cores). For example, x86-64 compilers generally prefer the longer VEX encoding (even in contexts where it does not help to avoid moves or transition penalties), or do not implement passes to avoid redundant REX prefixes.
- clausecker 1y agoLL/SC is performant, it just doesn't scale to high core counts. The VEX encoding is actually only rarely longer than the legacy one, and frequently it is shorter.
- deleted 1y ago[deleted]
- sylware 1y agoRISC-V is the better sweet-spot, and has no strong IP locks like arm or x86_64. Not to mention the silicon of nowdays changes everything: you avoid silicon design complexity as much as possible since it will be more than performant enough for the bulk of the programs out there.
- aleph_minus_one 1y ago> RISC-V is the better sweet-spot See https://gist.github.com/erincandescent/8a10eeeea1918ee4f9d9982f7618ef68 https://gist.github.com/erincandescent/8a10eeeea1918ee4f9d99... for an ex-ARM's engineer's critic of RISC-V. HN discussion: https://news.ycombinator.com/item?id=24958423 https://news.ycombinator.com/item?id=24958423
- sylware 1y agoOfc arm people will go after risc-v: it is a death sentence for them... Come on... But a real-life ISA doing a good enough job, without any global strong IP locks like x86-64 or arm... yummy.
- thesz 1y ago> Highly unconstrained extensibility. While this is a goal of RISC-V, it is also a recipe for a fragmented, incompatible ecosystem and will have to be managed with extreme care. This is taken straight from the ISA specification. For example, Intel can brand their own chips as RISC-V-ZIntelx8664, if they want to, because RISC-V-Zxxx implementation can be as incompatible with any other RISC-V implementations (and even specification) as one wants.
- tux3 1y ago>RISC-V-Zxxx implementation can be as incompatible with any other RISC-V implementations (and even specification) as one wants. It can't. Extensions have a lot of freedom, but still have to follow the core spec. You can't just submit amd64 as a riscv extension for a myriad of reasons, it clearly conflicts with the base ISA. This is also why there is a trademark. If you or Intel wanted to deliberately create nonsense to be disruptive, you are allowed to create it, but then you can't call it RISC-V without receiving some fan-mail from lawyers.
- drob518 1y agoYep. And then we learned that transistors were almost free, could be purchased in lot quantities of 1 billion, and could be used to create a translation layer between a CISC instruction stream and a RISC core.
- rjsw 1y agoNot for VAX, see the recent thread on it [1]. [1] https://news.ycombinator.com/item?id=45378413 https://news.ycombinator.com/item?id=45378413
- drob518 1y agoI’m not sure I get your point? Are you saying it is impossible to accelerate the VAX instruction set via the same technique used on x86? If so, you’ll have to explain why. Now, whether you’d want to or not is another question.
- p_l 1y agoYes, it's not possible to do it the same way, because what made x86 successful in their application is that x86 was remarkably RISC-y in actual behaviour compared to 68k or VAX. The main reason for that is that x86 code, outside of being register poor leading to lots of stack etc. use, decomposes most "ciscy" operations into LEA (inlined in pipeline) + memory access + actual operation. VAX (and to lesser extent, m68k and others) had multiple indirections just to get the operands right, way more than what is essentially single LEA instruction. The most complex VAX instructions could have been ignored as "this is super slow one and rarely used", but the burden of handling the indirections remain, including possible huge memory latency costs.
- clausecker 1y agoAlso, the VAX instruction encoding is a class of horror above that of x86.
- 1y ago
- DowsingSpoon 1y agoRISC vs CISC is nonsense. Pre-RISC CPU designs were pragmatic responses to the design constraints of their time. (expensive memory, poor compilers) RISC was a pragmatic response to the design constraints of its time. (memory becomes less expensive, transistor budgets are tight, and compilers are a little better) Post-RISC designs of today are, also, only pragmatic responses to the design constraints of today. Those constraints are different than they were in the 80s and 90s. The supposed dichotomy is just utter horseshit. It was invented as a marketing campaign to sell CPUs. It was canonized by the most popular text books being written by /Team RISC/. I wish, as an industry, we’d just get over it, move on, and stop talking about it so much.