3 ms·
I can definitely see his argument, although I still do believe RISC-V did a lot of things better than x86... I really do hope that the arch is eventually able
by ethin 2mo ago
I can definitely see his argument, although I still do believe RISC-V did a lot of things better than x86...
I really do hope that the arch is eventually able to fix this. Better that there be an open ISA than them all be closed IMO.
- wmf 2mo agoBetter than x86 is a low bar when ARMv8 exists.
- phire 2mo agoAnd personally, I'm not even sure it crosses that bar. RISC-V somehow manages to be more fragmented than x86 (which is impressive), and just can't compete on instruction density. I think a large part of the issue with RISC-V is that it predates (public knowledge of) ARMv8 by a year or two, so it couldn't use it as inspiration. If you compare RISC-V to 32-bit ARM, the comparisons are much more favourable.
- monocasa 2mo agoEverything I've seen is that rv64gc is very competitive with aarch64 wrt code density.
- wmf 2mo agoThe article makes the case that RISC-V achieved code density the wrong way. Instead of compressed instructions, ARM has fixed-size instructions with richer semantics.
- phire 2mo agoThe fact that it's only "competitive" with aarch64's code density is a solid black mark against RISC-V. The only reason it's "competitive" is the compressed instructions, which means it's paying all the costs of variable length instructions, yet only getting marginal benefits. IMO a modern ISA taking advantage of variable length instructions should be able to absolutely smash the code density of a fixed width ISA like aarch64. At minimum, it should be competitive with x86 code density, if not smashing that too (because x86 has a lot of legacy baggage) Compressed instructions aren't a bad idea for very small cores. They give you a decent code density boost with minimal added complexity. But for large cores you either want to go full fixed length (like AArch64 and Qualcomm's proposal, which bought non-compressed RISC-V into the range of AArch64) or adopt a much more complex variable length scheme that can actually beat x86 on code density.
- monocasa 2mo agoThere's a huge difference between 2/4 byte variable density and 1-15 byte variable density. And as I've said in other places, my experiments showed that it ended up being kind of across the board less than half a pipeline stage to handle C instructions, kind of orthogonally to decode width. It is a different front end design, so that's why Qualcomm didn't want to reengineer their aarch64 core more than they had to, but the rest of the riscv community was right to not embrace it. Not to mention that a lot of the aarch64 derived pieces in the proposed qualcomm extension are almost certainly patent encumbered. Qualcomm can absolutely handle just about any patent fight, but other risc-v companies can't.
- phire 2mo agoI agree that 16-bit/32-bit variable length would struggle to beat x86. But I suspect it could have gotten close, simply because x86 wastes a huge amount of its advantage on legacy cruft. The important point is that there is no reason why a 16-bit/32-bit encoding shouldn't have smashed Aarch64's 32-bit only code density. My secondary point, is that why should RISC-V limit itself to just 16-bit/32-bit? It has the encoding space set aside for 6 bytes, 8 bytes, 10 bytes and all the way up to 24 bytes (which is overkill). If it's already paying the variable length tax, it should be making better use of it. IMO, a 2, 4, 6, 8, 10... byte scheme should be able to massively improve on x86's code density.
- monocasa 2mo ago> I agree that 16-bit/32-bit variable length would struggle to beat x86. But I suspect it could have gotten close, simply because x86 wastes a huge amount of its advantage on legacy cruft. I'm saying the opposite. Maybe some theoretical CISC-V would leave RISC-V behind, but x86(and -64) makes wild choices for instruction density, and RV64GC already clearly beats x86-64 in .text density. > My secondary point, is that why should RISC-V limit itself to just 16-bit/32-bit? It has the encoding space set aside for 6 bytes, 8 bytes, 10 bytes and all the way up to 24 bytes (which is overkill). If it's already paying the variable length tax, it should be making better use of it. IMO, a 2, 4, 6, 8, 10... byte scheme should be able to massively improve on x86's code density. There's nonlinear issues as you add more options. A 16-32 decoder is pretty simple, a 16-32-48 isn't the worse thing in the world (and a 32bit immediate might make it worth it), but you start to hit weird explosions in gate count once you go much past that. Hence x86's splitting into essentially multiple front end banks in modern designs, and even then typically only has one decoder per bank that can decode everything, and even that takes multiple cycles for some instruction sequences, even just to discover the length. The larger lengths in the RISC-V spec are more targeted towards bespoke stuff like GPGPU that's maxing out issuing a single instruction per instruction stream anyway. When you look at shader machine code, it's clear density was essentially an afterthought, but they love them some 64bit wide instructions. Which unsurprisingly is pretty much the same width of vertical microcode in archs that still do such a thing.
- Someone 2mo ago> think a large part of the issue with RISC-V is that it predates (public knowledge of) ARMv8 by a year or two Does it? https://people.eecs.berkeley.edu/~krste/papers/EECS-2016-1.pdf https://people.eecs.berkeley.edu/~krste/papers/EECS-2016-1.p... has a section on ARMv8 (section 2.5) It says they became aware of it a year after they started the RISC-V project, but that’s five years before that paper was published.
- phire 2mo ago2015 is when it started to gain steam as a community run project. But version 1.0 of the spec [1] was released all the way back in May 2011, and the first RISC-V chip was taped out at the same time. This is 5 months before ARMv8 was even announced, and we didn't start seeing actual aarch64 chips until late 2013. And TBH, I'm not sure anyone realised just how good of an ISA aarch64 is until quite a bit later. RISC-V 1.0 isn't binary compatible with modern RISC-V, they hadn't frozen the encoding, but rough design is all there. [1] https://www2.eecs.berkeley.edu/Pubs/TechRpts/2011/Archive/EECS-2011-62.pdf https://www2.eecs.berkeley.edu/Pubs/TechRpts/2011/Archive/EE...
- kjs3 2mo agoYeah...risc-v can learn from 50 years of x86 (among others). And yet.......
- hn_submit 2mo agoHe has good points, except he misses the goal posts completely.
- 201984 2mo ago>I still do believe RISC-V did a lot of things better than x86... Such as? I can't think of anything it does better for high performance cores.
- _chris_ 2mo ago1B through 15B variable length instruction mess, for one. Which still yields a worse than average 4-5B per instruction average.
- 201984 2mo agoThat is a strength, not a weakness. It allows for things like 64-bit immediate loads, 32-bit branch offsets, and nearly unlimited future extensibility. With RISC-V, multiple instruction workarounds are needed for all of the above, and those sequences are usually sequentially dependent ones so they can't be run in parallel. i.e. the insanity of loading a 64-bit value through repeated 12-bit immediates with shifts, using multiple instructions to compute branch offsets, and RVV needing setvli instructions everywhere due to not having opcode space to encode vector length/type. AArch64 is better, but still has problems with limited opcode space when it comes to future extensions. They've had to make "start mode" and "end mode" for SME to save on opcode space, and future compromises will likely be necessary.
- Symmetry 2mo agoMore to the point x86 instruction streams aren't self-synchronizing. There are cases where you can read one valid stream of x86 instructions starting at byte X, but another completely different one starting at byte X+1. Apart from the security implications this makes wide decode on x86 notably harder than it has to be, though in practice you can make it work by just starting a decode your fetch window at every byte boundary the fist time you're executing something and throwing away the unused decodes, and then mark the invalid positions in the instruction cache so you don't waste that power again.