19 ms·
RISC-V: They Should Have Known Better
- random__duck 2mo agoI wonder if they will be inviting him to the next RISC-V design committee meeting.
- dmitrygr 2mo agoFor a friendly meeting, like Julius Caesar had on March 15, 44 BC.
- random__duck 2mo ago"This time its different".
- 80x86 2mo ago100% agree with dmitrygr. I was excited when I heard about the project just after it started. However, past experiences taught me to wait before getting excited about the new 'shiny thing'. I did it differently with RISCV. I waited. I am glad I did. It took a long time for actual silicon to appear. Also, the silicon today has all the facepalming special cases mentioned in the article. Its almost like those old soviet era cpus that had the list of bad instructions handwritten on the package. Overall, RISCV was a minor spin on MIPS, but without really learning from other processors. So why is everyone still pushing for it? It has the words 'open' on it. People pattern match on that marketing. As part of that marketing, they also pushed this attitude from the project... 'RISC won'. I think Chester Lam said it best when he wrote his essay stating that RISC didn't win... OoO archs won. I couldn't articulate that nearly as well as he did. If you haven't read it, I recommend it. So, yeah, here we are. Many people will follow the bandwagon, but they will find that RISCV will not make a significant difference. I am glad we still have Arm (in all its many forms), x86, and others. (btw, despite my username, I don't think x86 is the best either :-) Also, if you aren't trying to ship a product, you can experiment with ISAs on an fpga. Yes, fpgas are a lot slower, but they are also a lot more fun. Especially with the great work done to create open source toolchains. Heck, if you are really serious (slighly crazy), you can build your own chip. For the foreseeable future ASIC shuttles are available at prices under $10k. (again, you have to be a little crazy)
- random__duck 2mo ago> slighly crazy What a lovely euphemism. Signed: someone slightly crazy.
- p_l 2mo agoI'd say RISC won, when you consider how "RISCy" x86 is[1] compared to the ur-CISCs (68k, VAX) that RISC projects were in opposition to. [1] Not because of often-called "risc like" microcode engine, but because the most complex addressing mode on x86 usually decodes two microinstructions, and decodes in single cycle. In comparison VAX needed separate pipeline for instruction decoding.
- rnvannatta 2mo agoThe two winning instruction sets are the RISCiest CISC, x86, and the CISCiest RISC, arm.
- peterfirefly 2mo ago> In comparison VAX needed separate pipeline for instruction decoding. That's how the VAX 9000 and NVAX did it. It's not the only way. It is absolutely possible to decode a number of normal VAX instructions in parallel using a pipelined decoder similar to x86 decoders. It is also possible to use a µop cache similar to many x86 and ARM implementations. Fallbacks are only needed for the weirder addressing modes and for instructions that positively beg to microcoded (system calls/protection level transitions, some bit vector stuff, COBOL decimal stuff, block copy/scan/fill/compare, probably POLY) and startup and interrupt/exception handling. DEC never did this but they absolutely could have.
- p_l 2mo agoX86 does not really need pipelined decoders like NVAX did. The complex decode for x86 is the fast path for NVAX without going into CSU. And CSU is where all the more complex addressing modes on VAX end up going - the complex instructions are executed, yes slowly, but in separate unit once CSU finishes the decode for them (and even for packed decimal stuff I-box can theoretically decode in nearly one cycle if all operands are register or immediate). uop cache I'd admit could help for some cases, but still leaves you with even a simple ADD instruction possibly expanding into ~7 uops, maybe 3-4 if we assume big fused equivalent of LEA but then 2 of those will still stall with memory requests. DEC didn't try to parallelize the decoder further because it already could face 56 bytes for a single instruction, and the NVAX design was costly as hell. x86 in comparison has limit of max 15 bytes per instruction, and most instructions in x86 code fall in 4 bytes
- dzaima 2mo agoRandom minor-ish notes: - A big problem with extension detection RISC-V has is that there's no central authority mandating vendors to not overlap things (obviously, given RISC-V being an open standard), so basic bitmasks for supported extensions is generally rather problematic (and of course even if you collected a standardized bitmask of all extensions from all vendors, it'd grow quite massive quite quickly); you'd at least want some grouping/marking by vendor, if not full extension strings. That said, it would be nice to at the very least have some standard in-memory blob format if nothing else, that you could query from any OS/libc. (which maybe somewhat-exists to some extent with a C API meant for libc, but as-is still doesn't attempt to figure out vendor extensions). - many, if not the vast majority, of aarch64 TBZ/TBNZ are probably branching on a boolean; something RISC-V can also of course do in one instruction. Generally, comparing instruction frequencies across ISAs is messy if not approximately meaningless due to different sorts of things existing for solving the same tasks. - "Having this happen means that instead of a clearly-understandable crash you get ... well ... anything." - RISC-V will do you one better - it doesn't even guarantee a crash when an instruction isn't defined at all! Overlapping extensions is definitely messy for disassembly, sure, but that's also just basically unavoidable as long as RISC-V is open (see my first point). (perhaps there could've been stricter rules for reserved-for-standard encodings than reserved-for-vendor ones? of course still doesn't help vendor encodings, nor non-compliant vendors)
- dzaima 2mo agoSome more: > The spec says that bit must be zero, and yet no encoding uses the space opened up by that bit being one. The spec says "the code points with shamt[5]=1 are designated for custom extensions.", so the space is specifically reserved for custom vendor extensions. So, if I wanted to add a custom "dzaima.c.clear_top_n_bits rd, imm5" instruction, that's space I could safely put it in, knowing that no future standard instruction will be added there that I may regret overlapping. So while that space goes unused in the standard, its existence helps with the overlapping encoding problem! > For I-type instructions, bit 1 [...], bit 11 Of course, that's cherry-picking two of the 25% of bits that have multiple positions they come from, and specifically 11 as it's the worst one. Full stats: 1 position: 24 bits: (everything that's not listed below) 2 positions: 7 bits: 0, 1, 2, 3, 4, 12, 20 3 positions: 1 bits: 11 (the single worst case) So that's like 9 muxes for merging all immediates to the same place (or less of course if the different encodings' immediates go to different places), the rest is just wires. Obligatory note is that some of the funkiness is to place the sign-extended bit in the same bit position, so some saved muxes from that. Now, I am a "software person who's never written verilog", but I highly doubt a 3:1 mux is as cheap as a 2:1 mux in silicon, so even if you always need to merge in the sign bit, reducing the number of cases is still beneficial. Compressed does make it a ton more ugly though (combining both 32-bit and 16-bit instruction encodings, placing the 16-bit ones in the low 16 bits): 1 position: 13 bits 2 positions: 7 bits: 10, 13, 14, 15, 16, 17, 20 3 positions: 4 bits: 3, 4, 9, 12 4 positions: 5 bits: 0, 1, 2, 5, 11 5 positions: 3 bits: 6, 7, 8 looking at aarch64 on https://asmjit.com/asmgrid/ https://asmjit.com/asmgrid/: tbz Xt, #imm, #relS*4 imm:1|0110110|imm:5 | relS:14 |Rt lsl Xd, Xn, #n 1 1010011|01|immr:6|imms:6|Rn|Rd Fun! (lsl being a subset of the bitfield extract instrs is neat; tbz's similar-functionality 6-bit field is just entirely-differently placed though. Also.. using the Rd slot for an input-only Rt? that's one thing RISC-V doesn't do, even across compressed and 32-bit instrs!)
- kev009 2mo agoIt's basically MIPS all over again The conclusion is honest, and you can of course brute force any ISA into any role. I used to loathe x86 for that reason, but now that I'm older I respect the game.
- kjs3 2mo agoI think MIPS is a great example, and even there I don't think there's the bizarre bifurcation of ISA options RISC-V brings to the table. As a fellow olderster, I can't help but think that after almost 50 years of "ISA X is sooooo much better than x86 it's obvious ISA X is the future and x86 will be dead Real Soon Now (for whatever todays version of x86 is)" I can only shake my head ruefully and say "ping me when that happens". Controversial Take (that history proves isn't): Software matters; ISAs don't.
- kevin_thibedeau 2mo agox86 chips don't truly exist anymore. They only use it as a compressed ISA for a more capable internal representation that can be freely updated at any time.
- kjs3 2mo agoI keep seeing this line of reasoning and have no idea why it's relevant. You don't program that 'internal representation'. The software people want to run only care if that software doesn't run. Cyrix, Transmeta,NexGen, Centaur, WinChip, etc., etc, theoretically had "more capable internal representation". The only thing that actually matters is "does it run the exact same x86 software I bought X many years ago" and "does it run it at a decent price/performance ratio". Everything else is dick measuring. Today, we have Intel and AMD, and some bit-player embedded folks.
- monocasa 2mo agoSort of. They always had a much cleaner instruction set internally, going back to the 8086.
- jcranmer 2mo ago
- brcmthrowaway 2mo ago> What does a cheap microcontroller core need? Let's inspect what they are used for. Typical use cases are to interface with and quickly reconfigure hardware blocks in a larger chip, eg in an MP3 player, an SD card, or a USB stick. The hard work is done by custom IP and the CPU core is just there to occasionally prod a register or configure something. He forgot electronic cigarettes (vapes)
- __d 2mo agoSo … use RISC-V as the strawman, and create a community-based RISC-6 that doesn’t have these weaknesses? Better to get in now before it becomes too solidly entrenched.
- inigyou 2mo agoYou can't make a community-based ISA, it's not possible unless you have a community-based fab. He who makes the chips makes the rules.
- IshKebab 2mo agoLikely impossible unless you somehow come up with something vastly better (unlikely). None of these things are remotely bad enough to make the downsides of using another ISA palatable.
- NetMageSCW 2mo agoAnther ISA like ARM? It seems pretty palatable to just about everyone not academic.
- duskwuff 2mo ago
- mappu 2mo agoRVA23 hardware is available (e.g. SpacemiT K3)
- Joel_Mckay 2mo agoSome are already on RVA23.1 even before the standard made it to more than 4 manufacturers product lines. The meme joke about standards is sadly relevant for riscv. =3 https://xkcd.com/927/ https://xkcd.com/927/
- d-us-vb 2mo agoAs I’ve come to understand it, standards simplify intensionally, not extensionally. For those who select a part that is compliant with a standard, more standards to choose from is better because engineers are able to make better tradeoffs; they’re not forced to select a part that does way more than the application needs thus making the product more expensive if there are lots of “competing” standards: some do less some do more. For RV, a litany of standardized modules creates a system where each capability that the module provides will have a standard interface. No manufacturer is forced to invent extensions bespoke to their implementation, but they’re not forced to support everything the most powerful models do either. Just my two cents.
- ngl999 2mo agoThat is given, vendors actually _know_ what exact practical applications they are building for.
- Joel_Mckay 2mo agoSure, the constellation of features is no longer a general purpose computer in the retail context, but rather an ASIC appliance the ends up incompatible/useless rather quickly. Maybe Gentoo could tame that level of chaos... or people just buy ARM64 again knowing the software ecosystem already works. =3
- camel-cdr 2mo ago
- ethin 2mo agoI can definitely see his argument, although I still do believe RISC-V did a lot of things better than x86... I really do hope that the arch is eventually able to fix this. Better that there be an open ISA than them all be closed IMO.
- wmf 2mo agoBetter than x86 is a low bar when ARMv8 exists.
- phire 2mo agoAnd personally, I'm not even sure it crosses that bar. RISC-V somehow manages to be more fragmented than x86 (which is impressive), and just can't compete on instruction density. I think a large part of the issue with RISC-V is that it predates (public knowledge of) ARMv8 by a year or two, so it couldn't use it as inspiration. If you compare RISC-V to 32-bit ARM, the comparisons are much more favourable.
- monocasa 2mo agoEverything I've seen is that rv64gc is very competitive with aarch64 wrt code density.
- wmf 2mo agoThe article makes the case that RISC-V achieved code density the wrong way. Instead of compressed instructions, ARM has fixed-size instructions with richer semantics.
- phire 2mo agoThe fact that it's only "competitive" with aarch64's code density is a solid black mark against RISC-V. The only reason it's "competitive" is the compressed instructions, which means it's paying all the costs of variable length instructions, yet only getting marginal benefits. IMO a modern ISA taking advantage of variable length instructions should be able to absolutely smash the code density of a fixed width ISA like aarch64. At minimum, it should be competitive with x86 code density, if not smashing that too (because x86 has a lot of legacy baggage) Compressed instructions aren't a bad idea for very small cores. They give you a decent code density boost with minimal added complexity. But for large cores you either want to go full fixed length (like AArch64 and Qualcomm's proposal, which bought non-compressed RISC-V into the range of AArch64) or adopt a much more complex variable length scheme that can actually beat x86 on code density.
- exmadscientist 2mo ago> After being asked for the Nth time to explain, I decided to put it all down in one place so that I could simply link to it when asked next. Bookmarked, because I've needed the same. The worst part of all this is that they really should have known better by now. In 1980 you could make these kinds of mistakes, because this was pretty new territory. In 2020, doing this just makes you stupid. Or ignorant. Or both.
- NetMageSCW 2mo agoI’m not so sure - the 6502 existed in 1980 and showed the way.
- bsder 2mo ago6809 is a better exemplar, but, yeah, we knew this stuff way back when. The problem is that everybody around RISC-V wants to sell IP instead of a chip. Most of the worst brain damage follows from that. The rest of the brain damage follows from "We want to compete with ARM A-Series cores." No. Just ... no. Nobody willing to spend that much on a processor gives one iota of damn about ARM licensing fees. So, the semiconductor market wants a cheap, consistent chip that operates in the deep embedded space while the RISC-V ecosystem considers the mere thought of that to be icky beyond reason. And China will push on this like Longsoon and pray that somebody figures out how to make it not suck (Prediction: they won't succeed.) And, the worst part is that RISC-V has basically lost its window. The single possible advantage that RISC-V had was that as people converged to a shared tooling ecosystem it would create lockout. Unfortunately, that convergence never happened so, at best, we got some shared compilers. And, now, AIs can basically one shot all your other tools around it and probably the compiler not far behind. And there goes your ecosystem lockout.
- hn_submit 2mo agoBecause selling "bits" is very lucrative, whilst actual hardware can lead to huge losses if it doesn't sell. Just ask Microsoft. It's no wonder Microsoft is pulling out of the game console market and handing it over to PC manufacturers to make the actual hardware.
- brcmthrowaway 2mo agoIt's clear that RISC-V started as an academic exercise (albeit from a group with esteemed credentials) and they had to bolt on these hacks to make it work in industry. Sad.
- bigyabai 2mo agoARM was also rooted in an academic exercise. A lot of the drawbacks for modern ARM PC platforms stem from the aversion to actually advanced features like SVE/SVE2 and UEFI. It's sad, but it was also wildly successful. RISC-V has already replaced ARM in highly-custom embedded spaces like Nvidia's GPU controllers, and it likely won't stop unless ARM finally changes their tune vis-a-vis licensing.
- fidotron 2mo ago> ARM was also rooted in an academic exercise. Where do people get ideas like this from? Just nonsense.
- bigyabai 2mo agoARM, the ISA, is wholly rooted in academic exercises like Berkeley RISC. What do you think happened? RISC-I and RISC-II never existed, ARM means "Automated Reasoning Mechanism" and the ISA was never RISC whatsoever? Talk about nonsense, damn...
- sehansen 2mo agoNo, ARM means "Acorn RISC Machine" and was developed by Acorn Computers, a British computer manufacturer, because the 8-bit MOS 6502 they had based their previous designs on was obsolete, available 16-bit processors "a bit crap"[0] and available 32-bit designs too expensive. So the ARM ISA was developed by a computer manufacturer for use in their own chips in their own products, quite far from an "academic exercise". 0: https://www.youtube.com/watch?v=Hf67JYkUCHQ https://www.youtube.com/watch?v=Hf67JYkUCHQ at 7:35
- hn_submit 2mo agoWhy is he complaining about everything being optional in RISC-V? Isn't that the whole idea of RISC-V? The market can sort it out for themselves. RISC-V is already dominant in the MCU space despite its flaws, and many of them will be solved in due time. Most MCUs are used for dead-simple solutions, like electric blankets and microwaves with segment displays or LEDs. Whether their interrupts are handled in 44 or 22 cycles doesn't really matter that much. And RISC-V does have a link register, making returning much faster when the parameters for the interrupt can all fit in registers and no external memory access is needed, as is the case with most MCUs which put the stack in RAM. To fetch the return address an external memory access is always needed even if there are no parameters.
- tsukikage 2mo agoHe explains, at length: there is no sane way to determine what the hardware you are running on actually supports, and so there is no sane way to ship compiled code that is both compatible and performant. We already had the mystery meat CPU wars several decades ago. We know how to make sane ISAs now and should be past that.
- hn_submit 2mo agoYou don't need to probe what hardware you're running on because you know being the manufacturer. The code is bespoke for your solution and nothing more. No foreign code is going to run on it. Different problems require different solutions. An electric blanket doesn't need a barrel shifter for multiplication or even floating point hardware. The ISA can change depending on what's needed to solve a particular problem, not to provide an "one size fits all" solution.
- kjs3 2mo agoI don't think I've read a more "doesn't actually know anything about how software is produced, but with absolute confidence knows everything about it" post in a very long time.
- 2mo ago
- brcmthrowaway 2mo agoWhat happened to the Rivos accelerator cores?
- tsukikage 2mo agoMeta acquired Rivos last year.
- IshKebab 2mo agoThey got bought by Meta who then fired half of them.
- andrekandre 2mo ago> fired half of them always a winning strategy...
- brcmthrowaway 2mo agoThanosbook
- IshKebab 2mo agoI think a lot of this criticism is completely true. However it's also overblown. I do think the ISA matters, but little mistakes like these definitely don't matter enough to preclude making M-series class chips. The reason it hasn't happened yet is simply time. It takes a really really long time to build up to that level of performance. They've definitely gone overboard on the optionality stuff though. I don't think it matters too much for the actual CPU design but it makes verification and writing portable software a huge pain. Profiles definitely help but still... Oh also I feel like you could probably come up with an equally compelling list about any other ISA. It's not like the fact that something has flaws means it's bad.
- NetMageSCW 2mo agoI don’t think making optional what optional features are available is a little mistake. It is a torpedo to the waterline.
- IshKebab 2mo agoIt's not. In practice you have two scenarios: 1. You have a microcontroller. You're compiling code yourself and the docs tells you what features are available and which compiler flags to use. 2. You are writing application code. In that case you simply target RVA23. The edge case is the same edge case where you use CPUID on x86, I.e. you want to target say RVA23 and RVA28 in the same binary. In that case you do have to use the OS APIs to discover what is supported... which is slightly annoying, but in practice you're just calling a different function. In theory `mconfigptr` will eventually make this a lot nicer but nobody has put in the effort to define how it works yet (last I heard they were looking at ASN.1 sick emoji).
- yjftsjthsd-h 2mo ago> You are writing application code. In that case you simply target RVA23. You're allowed to not handle a majority of extant Linux-capable machines, but it seems like an awkward position.
- wren6991 2mo ago
- Retr0id 2mo agoI wrote an RV64IMA emulator recently. I just needed a virtual CPU core that could boot linux, and RV64IMA seemed like the simplest way to do that - and I think that's more or less true. But then I wanted to be compatible with off-the-shelf toolchains and binaries, and I found myself needing to extend the ISA profile to RV64GC. Not a huge lift, but it involved pulling in a softfloat library. That got me as far as booting Alpine linux. And then I wanted to be able to boot Ubuntu, which needed RVA23, which was comparatively a much bigger lift, involving the vector instruction set among many other things. At this point I think I'd have been better off just emulating aarch64.
- brucehoult 2mo agoUbuntu 24.04 LTS exists and needs only RV64GC and will be supported and enhanced for many more years. Debian has no plans to require more than RV64GC. RVA23 is a very good thing in certain markets, but nothing forces you to support it for a personal project.
- boredatoms 2mo agoI cant see debian remaining on rv64gc later on when enough rva23 boards are purchasable
- brucehoult 2mo agoWhy? It'll still run fine. Just like Debian still runs on original x86-64-v1 from 1999, not x86-64-v3 (needs AVX2,FMA, BMI1, BMI2, LZCNT) or even x86-64-v3 (needs AVX-512). Similarly, Debian for arm64 still requires only ARMv8.0-A from 2011 not even ARMv8.2-A (everything from A75/A55 to A78/N1/V1) let alone ARMv9-A (A710, A510, X2 and on). Why would they do in the RISC-V world what they totally haven't done in amd64 or arm64?
- camel-cdr 2mo agoNo, debian requires ARMv8.0-A + FP + NEON, as those are optinal extensions (even optional in ARMv9.0-A)
- wren6991 2mo agoRISC-V is... fine. It satisfies my two requirements for an ISA as a hobby CPU designer, which are: 1. Supported in mainline LLVM and GCC. 2. I can implement it without lawyers sending me a love letter. Everything else, I can fix in post. There are enough good ideas spread across the extensions that I can assemble a reasonably put-together, curated embedded ISA with competitive performance and code density that admits a simple implementation. I think Dmitry's points are largely on-target, though I have filed my usual statutory complaint that every rant that includes a bitfield diagram for the RISC-V J format should accompany it with a similar diagram for the Arm T32 BL encoding.
- zephen 2mo ago> RISC-V is... fine Exactly. > It satisfies my two requirements for an ISA as a hobby CPU designer... You probably have some unstated requirements as well, such as available toolchains and "vetted well enough to actually be able to run code." Risc-V now occupies the Schelling point for people who, for whatever reason (rent-seeking and security top the list) want to leave the x86 and Arm ecosystems.
- andrewflnr 2mo agoThey did explicitly specify: > 1. Supported in mainline LLVM and GCC. Which pretty well encapsulates the ecosystem requirements.
- brucehoult 2mo agoLuke is too modest. Something in the region of 5 million chips containing his hobby CPU have shipped since launch on August 8, 2024.
- gblargg 2mo agoJust noting, even if instructions were 100000000000000 bits long, reserving a single bit for 16-bit encoding would waste 50% of the instruction space.
- brucehoult 2mo agoIt's not wasted when it makes programs overall smaller, as it does.
- chrisjj 2mo agoIt is wasted if its reduction is less than that of alternative uses for that instruction space.
- brucehoult 2mo agoSuch as? There is still plenty of unused 32 bit (30 bit) opcode space.
- eek2121 2mo agoStarted reading, however I wanted to add this in: a lot of people expect RISC-V to do too many things, and nearly all of those things are "beat every other architecture out there in every way/shape/form, while also being open". The reality? The fastest "available" RISC-V CPUs don't match the best chips in terms of speed, power consumption, or die area. "available" obviously means the chips that have been released to the public and can be independently benchmarked. I do think that is okay, however I also think that those involved with RISC-V aren't helping much, and current attempts at standardizing seem to be just creating a bigger problem. That being said, RISC-V does seem to perform well in specific niches.
- bjornnn 2mo agothe significance and allure of risc-v, the reason china is investing heavily in it right now, has little to do with the technical details of how it works under the hood, it's the fact that it is an open standard not encumbered by intellectual property law. even if it isn't technically the best general-purpose processor architecture, it sets an important precedent by proving that it is possible to develop an open public architecture that the world can use to build computing devices without being extorted by a multinational corporation charging licensing fees or a geopolitical superpower enacting tariffs and sanctions.
- zephen 2mo ago> it's the fact that it is an open standard not encumbered by intellectual property law. There are actually many of those. But Risc-V has become, through effective marketing, the Schelling point for anybody who wants to avoid the x86 and Arm ecosystems, both for the rent-seeking behaviors you mention, and also, in some instances, for security reasons. And, as others have mentioned, the ISA doesn't really matter. As long as it's agreed upon, then the CPU vendors can optimize on one side, and the compiler writers on the other side. Sure, Risc-V has its warts, but you can certainly say the same about all the rest.
- mhh__ 2mo agoThe ISA not mattering I think isn't as true when you account cost e.g. in a huge OOO cpu all the fusions and so on are afaict fairly doable but if you are on a cheaper / worse CPU all those extra bytes in the instruction stream do add up.
- wren6991 2mo agoThe RISC-V fusion arguments from back in ~2018 didn't really pan out. A lot of those fusion opportunities are just instructions now. slli + add? Zba (sh*add). slli + srli? Zbb (zext.*). slli + srai? Believe it or not, also Zbb (sext.*). Look at that pair of RVC instructions you used instead of a single 32-bit opcode. They are: * Taking up valuable compressed instruction space; each compressed codepoint has an opportunity cost of 64k uncompressed ones. * Limited in which registers they can use (usually x8..x15). * Often clobber their input operand instead of giving a free move. Also consider that the frequency data that drove the RVC compression decisions was driven by the lack of architecturally fused instructions like sh*add, so any arguments you derive from that data are circular. An instruction can be a good uarch fusion target because it's compressed, and a good compression target because you didn't fuse it in the architecture. I think designing for uarch fusion in your ISA is coming at it from the wrong end. Fusion is something uarch designers do to make up for shortcomings in the ISA.
- UncleOxidant 2mo agoIs there a RISC-VI in the works where they try to learn from the RISC-V mistakes to make improvements?
- dmitrygr 2mo agoGiven the amount of learning that could have been done before RISC-V and wasn’t, I wouldn’t have such high hopes.
- phire 2mo agoConsidering just how many of the problems seem to come from RISC-V being a clean-sheet design, I suspect we would be better off not doing another. What I am interested in is the idea doing an AArch64 style revamp of the ISA, were much of the non-encoding semantic stuff is kept, but the entire instruction encoding (plus all the CSRs, and other things) are reworked to be sane. You might even do two reworkings in parallel, with one variable-width encoding optimised for microcontrollers, thumb-style; And the other being a fixed-width encoding optimised for wide out-of-order cores. And at the same time, you make a bunch of extensions mandatory, and unify others into bigger chunks; Code compiled to one of these two encodings would know it had access to a much wider range of instructions. The idea would be that any C code targeting RISC-V can be compiled to this encoding with close to zero changes, and that mechanical translation of exiting RISC-V binary code should be "possible", as none of the underlying semantics have changed. And the same would help any core wanting to natively support both (or all three) encodings, you would only need a front-end translator.
- acutelittlebox 2mo agoI feel like that's largely mitigated by profiles. RVA23 is really looking like it'll be the modern base target used for high performance application processors and it makes mandatory pretty much everything you'd want for those use cases, and other comments by people familiar with designing RISC-V CPUs mention that the variable length encoding can be dealt with in a very simple manner that doesn't even add another pipeline stage so it doesn't seem like it's all that big of a deal while also bringing in benefits in code size reduction. Not everyone is adopting it, but several major players have set the stage by mandating it.
- retired_account 2mo agoAlways good to see stuff from Dmitry; his presentation (Linux/4004) at last year’s Teardown was awesome.
- kazinator 2mo ago> Say you want to store a byte to a register plus offset. What range of offsets can a [compressed] 16-bit instruction encode? Zero through three. If a compressed instruction could load or store a word to a word-scaled offset 0-3, relative to a register base address, that would be quite useful. It could be used for accesses to all structures four words or smaller.
- dmitrygr 2mo agoIn thumb, it can encode 0..31
- kazinator 2mo agoHonestly, I would feel uncomfortable if I were designing an instruction encoding and came up with some addressing mode format where there are two bits for a displacement. I would pull myself aside and have a word with myself. That's just me, though.
- brucehoult 2mo agoAnd Arm dropped a T16-like encoding entirely from their 64 bit instruction set. If they did everything exactly the same they would be the same ISA not different ISAs. It's just as easy to point to things that RVC can do that T16 can't. You need to look at a far larger picture to decide on who made the better decisions overall.
- deleted 2mo ago[deleted]
- monocasa 2mo ago> The second category for big-compute is actual desktops and SBCs that do interactive computation, browsing, gaming, and other such "desktop work". I do not expect RISC-V to be a serious player at the top of this market. Simply put, the architecture is not designed for it, as pointed out above. Additionally, this market has the margins to afford licensing a much-better-designed aarch64 core from ARM, and gain proper support from a much larger corpus of software. Before you get your megaphone to shout about "openness", please note that the openness of the RISC-V spec is not relevant here at all, because an open spec does not magically materialize a well-designed out-of-order core for you for free. And if someone were to design a good out-of-order core, they would not be giving it away for free. An open spec does not mean every implementation is free. I basically disagree with this. Not because this isn't the current state of things (it absolutely is), but because we're at a bit of an inflection point where mooore's law has proved itself to be an scurve, and we're very clearly well into the top half of it. From that, gate counts per core will also start to ossify, and that means the longer latency for getting an open core design off the ground initially will also start to make sense.
- dmitrygr 2mo agoWhom do you expect to work for free to design you a state-of-the-art core?
- monocasa 2mo agoThe same kind of people that 'worked for free' to develop Linux.
- dmitrygr 2mo agoIf those people build cores like linux kernel is built design-wise, i will PAY to watch the spectacle. You do realize that Linux got basic SMP support 3 years after NT, and it was shaky for a while after? It still does not have reliable sleep-wake. And it only added native async file i/o in 2019, while NT has had it on the same hardware since 1993? So.. i'll expect an in-order core with an IPC south of 0.5 that cannot exit low power sleep 30% of the time in a decade or so.
- Neywiny 2mo agoI think I get it. I've tried microblaze-v for a while now. And just look at their interrupt handler. https://github.com/Xilinx/embeddedsw/blob/master/lib/bsp/standalone/src/riscv/trap_handler.S https://github.com/Xilinx/embeddedsw/blob/master/lib/bsp/sta... . With the FPU enabled at compile time, that's > 128 memory ops per interrupt. That's insane, especially without an NVIC and chaining and all that. My latency was astronomical, and my maximum interrupt frequency was pitiful. Ended up doing the work (sw and hardware options) to get it to operate more like arm-m, but arm-m doesn't need that work to be done. NVIC is always NVIC, and NVIC is good
- wren6991 2mo agoYeah, this is a bug. They should only be saving the FP state if it's dirty. Also this is one of the reasons I think Zfinx is a better option for embedded (i.e., the standard FP instructions operate on x registers instead of f registers): 31 registers is plenty to hold a mixture of integer and floating-point values, and you avoid the worst-case context save penalty.
- Neywiny 2mo agoNot a bug, just not optimized. Because I think there's a csr to read it the fpu is dirty but... That requires csr extension. It would also increase jitter, which in some cases is more important. At best it should be still there as an option, but also optionally improved
- wren6991 2mo agoFair enough. It's a performance issue but not a functional correctness issue. > Because I think there's a csr to read it the fpu is dirty but... That requires csr extension Yes, and they already unconditionally read that CSR :-) The "CSR extension" is an almost 100% theoretical concern. It was the spec authors being defensive in case the privileged ISA was so flawed they had to throw it out in future, while keeping the base ISA. I don't see that happening at this point. The only exception is deeply embedded cores that drop even basic IRQ and exception support. These are always going to exist and I think they're a sufficiently separate class of processor that they don't really factor into the compatibility equation, because such processors usually only run one program in their entire lives.
- atomicUpdate 2mo agoIt’s kind of funny that all of the complaints about optionality apply equally to Vulkan. Google even created the same profile solution with “Android Vulkan Profiles (AVP)”. I suspect Vulkan suffers from the same design by committee problem, which similarly caused it to miss seemingly basic features in the base spec that then need to be filled in with extensions and also made it too difficult for developers to want to move too.
- HexDecOctBin 2mo agoWell, graphics programmers were used to the mess from OpenGL days. Now compiler writer and hardware designers get to share the sorrow.
- panic 2mo agoWayland too!
- nc55g3g 2mo ago[dead]
- camel-cdr 2mo agoMy disagreement with the article is mostly the following: RISC-V is not an ISA, but an ISA generation framework. If RISC-V would've standardized aarch64 1-to-1, the end result would've still been a huge extension mess, because a lot of people (RVI member) have different requirements and a very happy to build their own subsets, which would then be upstreamed because multiple vendors want the same subsets and compatibility between them. Obviously it would've been better, similar to if RISC-V spawned with RVA23 done, but development takes time and RISC-V International started, because people where already using RISC-V. RISC-V also is the most DOSed ISA, with people proposing crazy stuff. Just the other day somebody proposed an instruction that would do up to 2^30 16-bit comparisons in one instruction at the largest VLEN. Because they wanted to improve their string processing usecase. --- In my experience RVA23 matches aarch64 and x86 in uop count (without fusion), code density is better, instruction count is slightly higher. The biggest impact on the instruction count advantage of aarch64 over RVA23 is a single instruction, load-pair, which gets cracked at decode in every high-performance implementation, because it writes to to registers. The Arm approach to code density is using multiple writeback instructions that have to be cracked and the RISC-V one is RVC. Both prohibit simple linear scaling of parallel decoding, so code density seems to have mattered to Arm enough to make the tradeoff worth it.
- camel-cdr 2mo agoMy disagreement with the article is mostly the following: RISC-V is not an ISA, but an ISA generation framework. If RISC-V would've standardized aarch64 1-to-1, the end result would've still been a huge extension mess, because a lot of people (RVI member) have different requirements and a very happy to build their own subsets, which would then be upstreamed because multiple vendors want the same subsets and compatibility between them. Obviously it would've been better, similar to if RISC-V spawned with RVA23 done, but development takes time and RISC-V International started, because people where already using RISC-V. RISC-V also is the most DOSed ISA, with people proposing crazy stuff. Just the other day somebody proposed an instruction that would do up to 2^30 16-bit comparisons in one instruction at the largest VLEN. Because they wanted to improve their string processing usecase. --- In my experience RVA23 matches aarch64 and x86 in uop count (without fusion), code density is better, instruction count is slightly higher. The biggest impact on the instruction count advantage of aarch64 over RVA23 is a single instruction, load-pair, which gets cracked at decode in every high-performance implementation, because it writes to to registers. The Arm approach to code density is using multiple writeback instructions that have to be cracked and the RISC-V one is RVC. Both prohibit simple linear scaling of parallel decoding, so code density seems to have mattered to Arm enough to make the tradeoff worth it.
- NooneAtAll3 2mo agowdym by "gets cracked at decode"?
- camel-cdr 2mo agoThe decoder decodes them into two or more internal instructions (uops). Take for example a post increment load, which does a=mem[b++], notice how this writes to two registers. Handeling two writes (up to 4) would explode the stage after decode (rename). So high performance arm implementations generate two uops for this. But since the number of decoders is fixed and the number of rename slots as well, you now have alnost the same problem as in RISC-V with compressed instructions: the nth input to the rename stage can come from a variaty of outputs of the decode stage, so you need a large shuffle network, and propagate the uop counts from start to end. Cracking is a lot cheaper, if you can do it later in the pipeline. E.g. the cheapest is if you can simply "replay" the instruction. That is, instead of removing the entry from the issue queue, when it starts executing, you decrement a counter and keep the entry to do something else next. But as I mentioned that doesn't really work with multiple write back.
- phendrenad2 2mo agoThings are generally defined by the neccessities that led to their creation. x86 was designed for home PCs and has been forced to evolve with PC technology. ARM was designed to take advantage of RISC architecture, and were forced to evolve with the mobile industry. What was RISC-V invented for, and what external forces have acted on it since then?
- xiphias2 2mo agoIf RISC-V was good enough for AMD to use it in their controller for their GPUs and it became cheaper than ARM, and NVIDIA is using it in many places, it was better to build upon than getting a change in ARM/x86 licensed and approved by Jim Keller, it's good enough. It turns out that the cost of waiting years for an ISA change is more costly than fixing whatever problems it has.
- erichocean 2mo agoWhat I like about RISC-V is not the ISA per se, but the ecosystem that has developed around it, particularly Chisel and CIRCT. Specific choices for instruction encoding is less interesting, especially in the age of AI.
- bhewes 2mo agoAh rants from a non designer. So Patterson and crew, don't know what they are doing? Yeah hard pass.
- aappleby 2mo agoHaving written a few RISC-V cores, worked on a chip design project that used RISC-V cores, and generally being OK with the architecture in real-world use cases: What the heck is this guy's problem? Just about every thing he mentioned as a problem is not a problem in practice. Too many options? Who cares, you're not trying to write code that runs on every possible configuration. Either you're writing embedded firmware and know exactly what core you're using, or you're writing an application that runs in an operating system and that system has a minimum ABI like RVA20 or whatever. Array accesses take an extra instruction? Either you're in a tight loop walking a tiny array and you don't do the full offset calculation per step, or you're walking over an array in RAM and you're bottlenecked by the memory bus. Hell, 90% of his arguments are "You can't detect X at runtime from user code without relying on some extension" - Yes, that is totally fine. Either you know your target CPU, or you don't - and then you ask your OS for details. This is not some dealbreaker. From the article - "For example, if you are writing a kernel and want it to support all RISC-V cores" - NOBODY IS DOING THAT. You target a platform spec, not the combinatorial explosion of everything from RV32E to RVA22 or whatever the latest is. You want to distinguish S mode from M mode? WHY DO YOU NOT ALREADY KNOW THIS? Instruction encoding is weird? WHO CARES, the decoding is like eight lines of Verilog. "Who can predict how their binary will act when a floating point store silently becomes a double-register move or a jump instruction, or vice-versa?" - THIS DOES NOT HAPPEN IN PRACTICE. Guhhhhh, I don't get it. This guy has some vendetta and either has not shipped any risc-v code or is just in love with his own personal favorite instruction set.
- Lord-Jobo 2mo agoYeah the OP post read to me like someone throwing the baby out with three drops of bath water. If this was presented more like “minor gripes with risc V” I’m guessing I wouldn’t feel that way
- bfrog 2mo agoThe encoding being oddball does have some effects on linkers/loaders though I imagine? Not that linking/loading is a super hot path people generally worry about.
- 2mo ago
- baron3dl 2mo agoThis feels like Andy Tennenbaum's LINUX is OBSOLETE post from 30 years ago.
- theamk 2mo agodon't see how? the last section is pretty explicit: > None of this is to say that RISC-V is doomed. As I said, I fully expect it to take over the space currently occupied by [...] Much like the linux kernel -- the price is right.
- desterothx 2mo agoLast section?! You think people actually read these before commenting?!
- baron3dl 2mo agoIt's interesting that this quote closed it for you, because that quote is what triggered my take. I read that as a hat tip to the legendary argument, and a partial adoption of Linus's rebuttal, “Linux wins heavily on points of being available now.” Besides the general tone, I guess.
- adrian_b 2mo ago+++ Excellent and well written description of the RISC-V ISA.
- daishi55 2mo agoWe are using RISC-V for AI accelerators to great success https://ai.meta.com/blog/meta-mtia-scale-ai-chips-for-billions/ https://ai.meta.com/blog/meta-mtia-scale-ai-chips-for-billio... RISC-V was a great choice due to being so customizable and extensible.
- djha-skin 2mo agoHe addresses this use case in the article.
- deleted 2mo ago[deleted]
- esseba-dev 2mo ago[flagged]
- thayne 2mo agoIf only ISAs weren't protected (or protectable) by patents.
- g8oz 2mo agoWhen writing a spec, every single thing you make optional, you split the possible implementations into two incompatible groups. Do this enough times and you end up with your spec being meaningless. I felt that.
- nullc 2mo agoI wonder how many of the obvious design shortcomings in RISC-V are from IPR avoidance / making IPR problematic parts optional.
- unfocso 2mo agoRefreshing style of writing. I know nothing about ISAs, but the rant was so fun
- segmondy 2mo ago[flagged]
- jack_h 2mo ago> What does a cheap microcontroller core need? Let's inspect what they are used for. Typical use cases are to interface with and quickly reconfigure hardware blocks in a larger chip, eg in an MP3 player, an SD card, or a USB stick. The hard work is done by custom IP and the CPU core is just there to occasionally prod a register or configure something. This is not the only reason to use a microcontroller or 75% of microcontroller vendor (e.g. STM) offerings would have no customers. Not everyone has custom IP that does all the work either, that’s actually fairly rare. It’s odd to pigeonhole microcontrollers like this just to go on a fairly lengthy rant about interrupt latency as if that somehow makes RISC-V unsuitable to what is an incredibly diverse application space. Maybe the rest of their post has better arguments, but I’m not impressed enough by the first one to keep reading.
- ACCount37 2mo agoThis is a "microcontroller core", not a "microcontroller". We're talking "deep embedded" applications - where an ASIC is designed for a very specific purpose, and that design just so happens to call for a programmable CPU core to be included in it. This is the kind of design that lives in your keyboard, your mouse, your USB stick, your USB hub, your HDD, your SSD, your eMMC chip, your memory card and more. Remember: you're never more than 3 meters away from an 8051 core. I do agree that most of this piece is nitpicking - poking at ultra low level things that are largely irrelevant to the tried and true "deep embedded" exercise of Just Ship It. No one really gives a shit if an operation takes one instructions or two, or which instruction sets are consistently present in different cores. What "deep embedded" people give a shit about is not having to work with ancient 8051 tooling and 8 bit ALUs and memory banked 64kb spaces while writing code for the one core they happen to actually have. And RISC-V got that. The piece actually agrees with that sentiment.
- random3 2mo agoYet it's royalty-free and good enough for Espressif (maker of ESP32) to move exclusively to the RISC V open-source instruction set architecture [1]. "Good enough ISA plus zero licensing cost" beats "perfect ISA plus royalties" in the embedded space. Also, let's not forget that the reason the world is built on the von Neumann architecture is that it was made available for free. [1] - https://www.eenewseurope.com/en/espressif-moves-exclusively-to-risc-v/ https://www.eenewseurope.com/en/espressif-moves-exclusively-...
- inigyou 2mo agoWhy didn't they make their own ISA long ago? Then they could have zero royalties. AFAIK ESP8266 was already its own architecture.
- hmry 2mo agoESP8266 uses an Xtensa CPU like the ESP32 (just a non-customizable preset)
- desterothx 2mo agoyeah thats kind of the conclusion if you read the article entirely...
- climate_denier_ 2mo agoNear the end of the essay, the author mentions that the folks at Berkeley considered OpenRISC. Would that have been a better path do go down, to throw a bunch of work, money, and R&D after, or is there anything inherently bad about that design besides delay slots? I kinda feel that even the smartest people will build great things on crumbling foundations as long as those foundations are available. I'm thinking of NASA embracing RISC-V or anyone who decided to write secure-by-design software in C. *edit - rephrased question for clarity
- weakhead 2mo agoI'm amused that the story doesn't even mention the 4k pages - way too small for anything but embedded systems today.
- pulse7 2mo agoSo it's mostly the "Optionality". Like USB. And yet USB is everywhere...
- inigyou 2mo agoAnd USB-C is known as a compatibility mess.
- izacus 2mo agoIs it though? Is it really? Outside the HN rant circles which want to return back to times where you needed an adapter for every single laptop model or be shit out of luck for connecting your mouse or projector?
- desterothx 2mo agoyes, hope this answers the loaded question :)
- inigyou 2mo agoI don't remember ever needing a USB-A adapter to plug in a mouse to a laptop. Until laptops started coming with only USB-C.
- deepsun 2mo ago> So what does it even mean to comply with the spec then, if everything is optional? Similarly, I kept saying it for long that a file/wire format's usefulness is not in what it supports, but in what it forbids. A binary file supports any type of data, but it's not useful.
- sylware 2mo agoThis guy is weird: there is no perfect ISA, only compromises and tradeoffs. He is looking for _his_ perfect. Won't happen, unless lucky, namely your perfect aligns with RISC-V tradeoffs. There are also sweet spots, and I am writting RISC-V assembly almost every day and that "hits" them often. I currently use at 99.99% the core ISA (I have a few muls and divs here and there). I don't even use the bit manipulation extension... There are millions of RISC-V chips out there. Performant microarchitectures are getting there, but the access to the latest silicon process is gated by the other ones, hogging production capacity (and they probably don't want RISC-V to "get there"...). And most of all, hardware manufacturer/designers won't have a lawyer ringing at their door: this is so much critical, this will make them tolerate a lot of RISC-V tradeoff choices they dislike. And ofc, big mistakes WILL BE MADE AND WILL HURT BAD. Expecting anything else is thinking like a teenager. It seems the current biggest mistake is the compressed instruction extension. It seems the complexity it adds for high performance is not worth it (arm removes the thumb instructions for reasons). I have suspicions on some microarchitectures designed around the compressed instructions (16bits) having a negative performance impact on core ISA 32bits instructions (and many compiler optimizations are friendly to the way compressed instructions are, namely the destination register is one of the source register, that due to the legacy x86_64). BTW, Intel APX something, is basically RISC-V for x86_64....... Another aspect people tend to forget while dealing with RISC-V, many of those design choices were made for the simplest way to implement performant CPU microarchitectures. Some say thats why on 'out-of-order' CPUs, you don't want a status flag register (there is none in RISC-V).
- artemonster 2mo ago"What range of offsets can a 16-bit instruction encode? Zero through three. Not thirty three, not three hundred and three. Three! Well, maybe it is better for storing a halfword? Nope... zero or two. What even? Why" What the fuck is this criticism? Its sanely specced, who would want arbitrary unaligned offsets, like for anything? Supporting such obscure idiot cases is too much unnecessary pain, so cut it off on spec level
- newsre4der 2mo agoMIPS or PowerPC is also free. We can use them too.
- Narishma 2mo agoMIPS has switched to RISC-V.
- speed_spread 2mo agoNot free enough. If your stuff becomes popular, you'll see patent holders popping and ask for money.
- phendrenad2 2mo agoThe PowerPC G5 was released in 2003. Patents last 20 years, so they expired 3 years ago. If a G5 isn't powerful enough for your use case, I don't know what is.
- enricotal 2mo ago[flagged]
- GeorgeTirebiter 2mo agoOne of the things that drives me nuts about RV is the plethora of Zextensions. And on top of those, manufacturers add on their own proprietary extensions. That is one of the selling points. But... Interestingly, in the Olden Days, machines often had custom instructions. The pdp-1/D had a tad instruction for 2's complement addition (it was normally 1's complement machine), and there were also new pdp-1 instructions for timesharing. For the IBM 1401, there were all sorts of add on 'features', e.g. the Multiply / Divide Feature (sped up * and / in HW), High-Low-Equal Compare feature, Advanced Programming (added Index Register, Subroutine Linkage), Move Record feature (allows right-to-left movement until hitting a Record Mark, useful for Tape), Expanded Print Edit Feature (float $ sign, automatic * insertion, etc), various memory sizes from 1,400 characters up to 16,000 characters. The point is that manufacturer-distributed software had to deal with having features (or not); and this was handled e.g. in the assembler by having a CTL card that listed (coded) the features, e.g. "CTL 31110" says this is a 12K machine, with Automatic Multiply / Divide, the High-Low-Equal Compare, the Move Record Feature, but NOT the Expanded Print Edit Feature -- that told the Macro Generator what it needed to know to properly expand macros for a specific HW configuration. (Certain Features did not require SW mods, e.g. the notorious "Print Overlap" feature that would speed up print operations by having a hardware buffer to hold the Print Line, so the CPU did not need to stall. The feature was notorious because the sub-rack of HW required to implement it connected into probably 75% of the machine's instruction decoder & execution units; when it failed, it was incredibly painful to find the fault(s)). And, there have been Writeable Control Stores like, forever. It was a feature on the Burroughs B1700, where different control stores could be loaded on-the-fly depending if you were executing COBOL or FORTRAN or Pascal - the 'instruction set' would be optimized for running that particular language. The pdp-11 had some version of this, and CMU's custom C.MMP had something like this. So, the desire for certain customers to have machines specifically honed to their use cases was normal. It has only been this brief period of homogenization of single-chip(ish) CPUs (8008, 8080, Z80, 8086... amd64 etc) that introduced new instructions in tranches. The End-Users now get their custom instructions in other ways (e.g. PCIe and GPUs)
- ur-whale 2mo agoHere's an interesting ChatGPT conversation posted in the comments of the article: https://chatgpt.com/share/6a81dbb7-1934-83ea-8a47-6d1d1b4c0f17 https://chatgpt.com/share/6a81dbb7-1934-83ea-8a47-6d1d1b4c0f...