5 ms·
The author fundamentally misunderstands what RISC is, what CISC is, what SIMD is and what VLIW is apart from misunderstanding every major computer architecture
by teton_ferb 7y ago
The author fundamentally misunderstands what RISC is, what CISC is, what SIMD is and what VLIW is apart from misunderstanding every major computer architecture concept
RISC was invented as an alternative approach in an era when processors had really complex instructions, with an idea that high level languages could be efficiently compiled to them and assembly programmers would be efficient if they can do many things with one instruction. RISC philosophy was to make simpler instructions, let compilers figure out how to map high level languages to simple instructions, and therefore fit the processor on one die (yes, "processors" used to be several chips) and therefore run it at high clock speeds. RISC is not a dogma, it is a design philosophy.
On top of that exception handling in complex instructions is hard. Implementing complex instructions in hardware consumes considerable design and validation effort. RISC has won for these reasons.
Some things have changed, we can fit really complicated processors on a single die. Memory access is the bottleneck. The downsides of RISC in this reality is well known: It takes many more instructions to do the same thing, which means instruction cache is used inefficiently (anyone remember the THUMB instruction set of ARM?). It might be useful to add application-specific hardware acceleration features, because we now have the transistors to do it. How does that make RISC unscalable?
Many CISC machines (eg Intel's) are CISC in name only. The instructions are translated to micro-ops in the hardware. The micro-ops and the hardware itself, is RISC.
Register re-naming was invented to ensure that we enjoy the benefits of improvements in hardware without having to recompile. Let us assume you have a processor with 16 registers. You compiled software for it. Now we can put in 32. What do you do? Recompile everything out there or implement register re-naming?
VLIW failed because they took the stance that if we remove hardware-based scheduling, the extra transistors can be used for computation and cache. Scheduling can be done by compilers anyway. The reason they werent successful is because if a load misses the cache, you wait. Instead of superscalars which would have found other instructions to execute. On top of it, if you had a 4-wide VLIW and then you wanted to make a 8-wide one, you had to recompile. And oh, the "rotating registers" in VLIW is a form of register re-naming.
Poorly informed article.
- ajross 7y ago> The reason they werent successful is because if a load misses the cache, you wait. Instead of superscalars which would have found other instructions to execute. I'm not sure I quite buy that. In practice, what would happen is that a suitably large, optimized VLIW core would start fetching more than one wide instruction in a cycle and issuing the resulting ops in parallel with interlocks for dependencies, etc... Effectively, that is, VLIW would drop the explicit promise of the instruction set and turn into a speculative RISC core internally. And the cost for that translation would have been very comparable with what we see in all the existing very successful x86 CPUs. But this never happened, because we never got that far. VLIW failed for other reasons. This particular problem had an obvious safety safety valve.
- gpderetta 7y agoDenver (and Crusoe before it) implements OoO behaviour on top of a VLIW core via dynamic translation. It performed fairly competitively, but in the end the complexity moved from one place to the other and was hardly worth it.
- CalChris 7y agoThe Denver team came from Transmeta but they didn't bring VLIW with them. Denver is a very wide (7+ wide) in-order superscalar pipeline. https://en.wikichip.org/wiki/nvidia/microarchitectures/denver https://en.wikichip.org/wiki/nvidia/microarchitectures/denve... I can't even understand how OOO VLIW would work.
- gpderetta 7y agoThe rumour I have heard is that Denver is indeed an VLIW machine (in order superscalar does not exclude that). It might be wrong of course. Regarding the OoO part, the dynamic scheduling is supposedly done by a JIT layer in firmware.
- CalChris 7y agoEverything I've read/heard is that it's superscalar. They've gone from 7 wide to 10 wide with Carmel. That's a lot of width to schedule and/or waste. The other big difference is that Denver uses a HW ARM decoder vs Transmeta's SW decoder. That was smart since they could license the IP from ARM whereas Intel would fight them every step of the way.
- gpderetta 7y agoIt seems that we had the same discussion 2 years ago on HN :). I can't find any authoritative source so I'll concede. The guys at RWT are pretty sure it is an VLIW though.
- deleted 7y ago
- hybrids 7y ago> Many CISC machines (eg Intel's) are CISC in name only. The instructions are translated to micro-ops in the hardware. The micro-ops and the hardware itself, is RISC. This is a half-truth at best. Unless you work at Intel/AMD/etc. as a chip designer odds are you do not know what is "really happening" behind the scenes, so the underlying implementation is whatever they "want it" to be. The underlying implementations can even change from microarchitecture to microarchitecture. So internally we might make a guess that microcode must look more "RISC-y" after transformation, given the fact that "some transformation" must be happening. The larger internal (renamed-to) register file was a common trait of RISC architectures, after all, and a lot of microcoded systems historically have been regarded as "RISC-y." But the existence of modern optimizations like macro-operation fusion suggest that the internals of CPUs today renders things much more ambiguous today with regards to "RISC-ness" vs. "CISC-ness." Fun fact: the original ARM1 processor in 1985 was microcoded - https://en.wikichip.org/wiki/acorn/microarchitectures/arm1#Decode https://en.wikichip.org/wiki/acorn/microarchitectures/arm1#D...
- adwn 7y ago> Many CISC machines (eg Intel's) are CISC in name only. The instructions are translated to micro-ops in the hardware. The micro-ops and the hardware itself, is RISC. This is an oft-repeated but incorrect statement. Modern x86 CPUs perform macro-op fusion and micro-op fusion. As an example of the former, a comparison and a jump instruction can be fused into a single micro-op [1], which is decidedly non-RISC. As for the latter, some micro-ops perform a load from memory and an arithmetic operation with the retrieved value – also very non-RISCy. Modern x86 CPUs are CISC above and below the surface. [1] https://en.wikichip.org/wiki/macro-operation_fusion#x86 https://en.wikichip.org/wiki/macro-operation_fusion#x86
- monocasa 7y agoYeah, vertical microcode always looked pretty RISCy if you're not used to it. And FWIW, I've heard that there's two distinct uOp formats inside a.single core these days for quite a few of the uArchs. There's the frontends view which is concerned with amortizing decode costs (so wide fixed purpose instructions, that other wise look pretty CISC), and the backend's uOps that's concerned with work scheduled on functional units. A lot of the fusion happens on the front end, and a lot of the cracking happens on the backend, and the frotbejd tends to be two address, and the backend three address. So like a frontend's and rax, [rbx, addr] is something like ld_agu t0, rbx, addr ld t1, t0 and rax, rax, t1 on the backend.
- CalChris 7y agoI'd add that horizontal microprogramming well predates RISC. x86 μops are wide.
- CalChris 7y agoRISC was invented as an alternative approach in an era when processors had really complex instructions Yes, RISC was simpler than the VAX-11 (which was simpler than the 432). But what did RISC do that Cray hadn't already done? It was a Cray on a chip sans vector. Even the CDC 6600 (also co-designed by Cray) was load/store.
- zaarn 7y ago>Register re-naming was invented to ensure that we enjoy the benefits of improvements in hardware without having to recompile. Let us assume you have a processor with 16 registers. You compiled software for it. Now we can put in 32. What do you do? Recompile everything out there or implement register re-naming? Just run the software? Nothing forces a software to use all available registers, if yours was compiled for 16 and the CPU has 32, it uses 16 registers. Modern CPUs have plenty of additional instruction codes too and you don't need to recompile your software because one CPU has the RDRAND instruction and some other doesn't have it (minus AMD breaking RDRAND but that's a different topic). >The micro-ops and the hardware itself, is RISC. I would argue that µOps are VLIW wearing a RISC hat, it's very VLIW-y what's happening under the hood in x86, just less coherent. >On top of it, if you had a 4-wide VLIW and then you wanted to make a 8-wide one, you had to recompile. And oh, the "rotating registers" in VLIW is a form of register re-naming. As mentioned above, in that case the software doesn't perform as well as it could but there is no reason it would stop working if you properly designed it.
- imtringued 7y ago>Many CISC machines (eg Intel's) are CISC in name only. The instructions are translated to micro-ops in the hardware. The micro-ops and the hardware itself, is RISC. No, that's just a popular HN myth. First of all microops are not part of the public ISA which is what the entire RISK vs CISC debate is about. If the public ISA does not matter and you can always convert to whatever is best then RISC loses because it's raison d'etre is to enable optimizations through a betzer ISA. Second using microcoding (not the same as microops) to implement a huge number of instructions is an integral core of CISC, if you are saying microcode is RISC then CISC did RISC way before RISC even existed. RISC was all about removing microcoding to simplify CPU designs so clearly if your CPU has extensive microcode it is not RISC and definitively not "CISC in name only". If CISC is RISC then why even bother with RISC? After all CISC does "RISC" better.