13 ms·
This may add a bit of colour to a technical debate going on right now in the RISCV Profiles working group. Qualcomm have proposed dropping the C (16-bit compres
by gchadwick 3y ago
This may add a bit of colour to a technical debate going on right now in the RISCV Profiles working group. Qualcomm have proposed dropping the C (16-bit compressed instructions) extension from the RV23A profile (effectively the set of things to support if you want a 'standard' high performance RISC-V core). They have two main reasons
1. The variable length instructions (currently 16 bit or 32 bit but 48 bit on the horizon) complicate instruction fetch and decode and in particular this is a problem for high performance RISC-V implementations.
2. The C extension uses 75% of the 32-bit opcode space, this can be put to better use.
They're saying the benefits from the C extension don't outweigh the costs. They're also saying that if you move forward with the C extension in RVA23 now there's no real backing out of it. As the software ecosystem develops removing it once it's baked in just won't be possible. However adding it back in later is more feasible.
SiFive strongly disagree. They believe the C extension is worth the cost and that it doesn't prevent you from building high-performance cores. They also say that there's lots of implementations with C in already, so backing out of it now disadvantages those implementations.
It could end up in being the first major fragmentation in the eco-system Qualcomm go one way and SiFive the other (other companies also sit on one side or the other of this debate but Qualcomm and SiFive are driving it). Indeed the latest proposal from Krste is to do just that with a new 'RVH23' profile..
I wonder how much this development has been driving SiFive's thinking here? Clearly they are under pressure to deliver to their investors so you can see why they want to keep things as they are rather than consider a big change. Good for SiFive, but good for the long-term RISC-V ecosystem?
Edit: If you want the details check out the publicly readable tech-profiles list: https://lists.riscv.org/g/tech-profiles/messages https://lists.riscv.org/g/tech-profiles/messages they've got recordings of the last two meetings that discussed the issue and presentations, all available via that message archive.
- klelatti 3y agoJust to add that Qualcomm have also proposed a new extension that will help keep code size down without using the C extension. It includes new load/store addressing modes, pre-post increment load/stores and load/store pairs amongst others. It would seem to take RISC-V closer to AArch64 in approach?
- gchadwick 3y agoYes exactly, the 'existence proof of a competitive architecture using exclusively 32-bit instructions' has often been reference. Qualcomm's proposal is all instructions are aligned to their size. Initially that means everything is a 32-bit instruction, now with a lot more green-field encoding space to play with (so less need to have larger instructions). 64-bit instructions would be introduced (aligned on a 64-bit boundary) when needed with the expectation they'd be used for rare operations and 48-bit instructions wouldn't happen. The SiFive (and original RISC-V architects view) is RISC-V is meant to be a variable length instruction set and a mix of 16/32/48 provides better static code size along with better dynamic code size meaning smaller icaches needed, smaller buffers in fetch units etc. Interesting that the architecture that was meant to be a 'purer' RISC implementation than ARM is pushing towards the more CISC style variable length instructions. In a sense Qualcomm are trying to keep it closer to the RISC ideal!
- _a_a_a_ 3y agoWhy would you need a 64 bit instruction; what kinds of things are going to be used for it? What does 'rare' mean here, does it mean rare in execution, or rarely appears in code? (The difference being that something might only appear once in your code but be part of your hot loop so be executed any number of times) If they are rare in execution, what is their value over composing them of 32-bit instructions, where the (rare) overhead of doing so would be typically a amortised away? (The only thing I can think of that 64 bit instruction seem suited to is some kind of internal CPU management instructions, but context switches etc. are relatively rare & very expensive anyway so... I don't know)
- londons_explore 3y ago> The variable length instructions (currently 16 bit or 32 bit but 48 bit on the horizon) complicate instruction fetch and decode and in particular this is a problem for high performance RISC-V implementations. I want to see variable length instructions, but a requirement for instruction alignment. Ie. every aligned 64 bit word of RAM contain one of these: [64 bit instruction] [32 bit instruction][32 bit instruction] [16 bit instruction][16 bit instruction][32 bit instruction] [32 bit instruction][16 bit instruction][16 bit instruction] [16 bit instruction][16 bit instruction][16 bit instruction][16 bit instruction] That should make decode far simpler, but put a little more pressure on compilers (instructions will frequently need to be reordered to align - but a review of compiler generated code is that that frequently isn't an issue)
- camel-cdr 3y agoThis is basically what qualcomm proposes, 32 bit instructions and 64 bit aligned 64 bit instructions. I don't think we have real data on it, but I suspect that the negative impact of this would effect 16/32/48/64 way more than just 32/64.
- londons_explore 3y agoI would like to also have aligned 16 bit instructions. And maybe even aligned 8 bit instructions for very common things like "decrement register 0" or "test if register 0 is greater than or equal to zero". "Jump back by 10 instructions", etc. Those instructions get widely used in tight loops, so might as well be smaller.
- adgjlsfhk1 3y ago8 bit instructions are a really bad idea. you only get a tiny number of them and they significantly increase decode complexity (and massively reduce the number of larger instructions available)
- gchadwick 3y agoAs the other reply states that is effectively the Qualcomm proposal though note the 16-bit instructions likely gobble up a large amount of your 32-bit instruction space. You have to have something to identify an instruction as 16-bit which takes up 32-bit encoding space. The larger you make that identification (in terms of bits) the less encoding space it takes up but then the fewer spare bits you have to actually encoding your 16-bit instruction. RISC-V uses the bottom two bits for this purpose, one value (11) indicates a 32-bit instruction, the others are used for 16-bit instructions. So you're dedicating 75% of your 32-bit encoding space to 16-bit instructions.
- fidotron 3y agoRISC-V architectural purism was never going to survive any major effort to deploy it. Either you make changes like what Qualcomm suggest here or you aren't competitive. The major question is how well RISC-V will manage disputes over this sort of thing without some group such as Qualcomm deciding to just release their version anyway.
- mort96 3y agoI'm wondering what part you call "architectural purism" here. Spending a whole lot of opcode space on a set of compressed instructions doesn't strike me as an especially purist solution to the code size problem, and if what camel-cdr suggests in https://news.ycombinator.com/item?id=37997077 https://news.ycombinator.com/item?id=37997077 is correct, then Qualcomm's solution is also pretty much a set of 16-bit compressed instructions, but where the 32-bit instructions must be 32-bit-aligned, which strikes me as neither significantly more nor significantly less pure than the current C extension. To me, this looks like a reasonable argument over design decisions, where there are clear advantages and disadvantages to either side. It's basically a trade-off between code size and front-end complexity. Can you detail where exactly you see the purism thing being an issue?
- fidotron 3y agoI believe Qualcomm are proposing dropping 16 bit instruction support, exactly like Aarch64.
- deleted 3y ago[deleted]
- mort96 3y agoYou seem to be right. I had interpreted some other responses in this thread to mean that Qualcomm has their own alternative 16-bit encoding that doesn't have the 32-bit instruction alignment issue, but it seems like they instead have a whole bunch of new 32-bit instructions which have memory operands and a bunch of addressing modes. I see now what you mean by posing this as a conflict between ISA purists (only provide load/store all other instructions have register or immediate operands, only provide one store and one load instruction, add compressed instructions to combat binary bloat) and ISA pragmatists (add new special-case instructions with memory operands and useful addressing modes).
- camel-cdr 3y agoI think James comment summaries the problem quite well, as both aises have segnificant self interest/sunk cost for their prefered approach: https://lists.riscv.org/g/tech-profiles/topic/rva23_versus_rvh23_proposal/102127876 https://lists.riscv.org/g/tech-profiles/topic/rva23_versus_r...
- londons_explore 3y agoI'm not sure either side has all that much sunk cost. So far, there isn't much binary RISC-V code thats distributed in binary form and expected to run on future processors. So far, nearly all RISC-V is in the embedded space where everything is compiled from scratch, and a change to the ISA wouldn't have a huge impact. Far more important to get it right for RISC-V phones/laptops/servers, where code will be distributed in binary form and expected to maintain forward and back CPU compatibility for 10+years.
- Narishma 3y agoThe sunk cost here I think refers to the existing CPU designs of the respective camps. Qualcomm's ARM-based cores don't support an equivalent of the C extension and adding it would presumably require major and expensive rework.
- phkahler 3y ago>> This is basically what qualcomm proposes, 32 bit instructions and 64 bit aligned 64 bit instructions. Well that's only sunk cost if they assumed from the start that they were going to change the design to RISC-V AND drop the C extension. In that case, it was a rather risky plan from the start - assuming they can shift the industry like that. I'm guessing RISC-V was a change of direction for them and this would make things easier short term.
- dezgeg 3y agoDebian has already started compiling it, and even more will be by the time this new incompatible ISA would hit the shelves. The time to 'get it right' has already passed IMO. If a hard ISA compatibility break happens at this stage, who is going to trust that it won't happen again?
- mort96 3y agoHmm. The C extension has played a very important PR role for RISC-V, with compressed instructions + macro-op fusion being the main argument for why the lack of addressing modes is no big deal. It would be interesting to see how big of a difference it actually makes to binary sizes in practice though.
- monocasa 3y agoWe know, it's around 30% larger binaries. That's why qualcomm also added a bunch of custom extensions.
- mort96 3y agoThis is an important point which I didn't realize after reading only gchadwick's comment. It's a discussion of how best to design a compressed instructions extension, not whether to have a compressed instructions extension. Does Qualcomm have a concrete proposal for how their version of compressed instructions would work, or is the idea more or less just "the C extensions but 32-bit instructions must be 32-bit aligned"? Have they published details somewhere?
- jabl 3y agoAFAIU qualcomms proposal for extra 32-bit instructions is https://lists.riscv.org/g/tech-profiles/attachment/332/0/code_size_extension_rvi_20231006.pdf https://lists.riscv.org/g/tech-profiles/attachment/332/0/cod... It adds new addressing modes, and things like load/store-pair instructions.
- mort96 3y agoAh I had assumed they proposed their own alternative compressed instructions without the alignment issues, but they're actually proposing more addressing modes and adding instructions which operate directly on memory. That makes sense I guess.
- 3y ago
- denotational 3y agoFrom what I remember, Krste has advocated for compressed instruction sets with macro-op fusion in the uarch front-end for a while, and the design of the RV ISA is heavily inspired by this, so it’s not particularly surprising that SiFive (i.e. Krste’s (and others’) company) is opposing Qualcomm’s proposals. It will be very interesting to see what happens.
- aseipp 3y agoI think an example is something like opcodes crossing I-cache lines, re: fetch and decode complication; instructions are 16-bit aligned when C is present, so you can have a 32-byte instruction cross cache lines easily. At minimum it will definitely require a bunch of extra verification to handle those cases, and that's often the longest part of the whole development process anyway, so I see the reasoning for not wanting it. It doesn't matter how high performance something is or can be, if you can't prove it works to some tolerance level. I know there's the big discussion about macro-op fusion. But in hindsight, I think a big motivator for C -- implicit or not -- was the fact that on the very low-end microcontroller or in the softcore (FPGA) world, you typically have disproportionately low amounts of SRAM available versus compute fabric. Those were the initial deployment targets (and initial successful deployments!) for RISC-V, since you need tons of extra features for "Application Class" designs. These cores often have a short pipeline and are completely in-order, so their cost and verification effort are much lower. These are (very likely) not going to implement macro fusion, at least on the medium-low end. So, increasing the effective size of the I-cache through smaller opcodes is often a straight win to increase IPC. On the other hand, Application Class designs today are typically OoO, so they achieve high IPC while still hiding miss latencies pretty effectively; smaller instructions are still good but the benefits they provide aren't as prominent. And it does use a ridiculous amount of opcode space, yes. I wonder if they would have just been better off copying one of ARM's design principles from the very start: actual design families akin to the -M, -R, and -A series of ARM processors, created for different actual design spaces. These could actually be allowed to have (potentially large!) incompatibilities between them while still sharing a lot of the base instruction set and privileged e.g. PMP extensions could probably exist among all of them. I'd be happy to have an "Application Class" "-A series" RISC-V processor that could run Linux but didn't have compressed instructions or whatever; likewise I would probably not miss e.g. Hypervisor extensions on a microcontroller. EDIT: Clipped an incorrect bit about ABI compatibility with the C extension. I was misremembering some details about a specific implementation!
- mort96 3y ago> Another big issue for Application Class systems IIRC -- unrelated to all this -- is that I don't think hardware implementing C can actually run binaries compiled without it. I believe this is incorrect? I believe RV{32,64}-with-C is simply a superset of RV{32,64}-without-C. Now I have only implemented RV32I, so I'm not that familiar with the C extension or other extensions for that matter, but in my digging through the various RV specs, I haven't found anything which suggests that implementing C requires breaking code compiled without the use of C. Do you have any details?
- monocasa 3y agoThere's a bit more context rumbling under the surface. Not too long ago, Qualcomm bought NUVIA, a designer of high performance arm64 cores that can theoretically compete with Apple cores on perf. Arm pretty much immediately sued saying that the specifics of the licenses that Qualcomm and NUVIA have mean that cores developed under NUVIA's license can't be transferred to Qualcomm's license.[0] Qualcomm obviously disagrees. Whatever happens those cores as they exist today are going to be stuck in litigation for longer than they're relevant. Qualcomm's proposal smells strongly like they're doing the minimum to strap a RISC-V decoder to the front of these cores. For whatever reason the seem hell bent on only changing the part of the front end that's the 'pure function that converts bit patterns of ops to bit patterns of micro-ops'. Arm64 is only 32bit aligned instructions, so they don't want to support anything else. At the end of the day, the C extension really isn't that bad to support in a high perf core if you go in wanting to support it. The canonical design (not just for RISC-V but high end designs like Intel and AMD too) is to have I$ lines fill into a shift register, have some hardware on whatever period your alignment boundary is that reports 'if an instruction started here, how long is it', and a second stage (logically, it doesn't have to be an actual clock stage) that looks at all of those reports generates the instruction boundaries and feeds them into the decoders. At this point everything is also marked for validity (ie. did an I$ line not come in because of a TLB permissions failure or something). [0] - https://www.reuters.com/legal/chips-tech-firm-arm-sues-qualcomm-nuvia-breach-license-trademark-2022-08-31/ https://www.reuters.com/legal/chips-tech-firm-arm-sues-qualc...
- ekiwi 3y agoThe 32-bit aligned instruction assumption is probably baked into their low-level caches, branch predictors etc. That might mean much more significant work for switching to 16-bit instructions than they are willing to do.
- monocasa 3y agoI don't think anyone bakes instruction alignment into their caches since the early 2000s, and adding an extra bit to the branch predictors isn't that big of a deal. It's got to be the first or second stage of their front end right before the decoders.
- deleted 3y ago[deleted]
- JonChesterfield 3y agoThe compressed instruction extension was described somewhere as overfit to a naive gcc implementation which seems plausible. It does have a significant cost to a 32bit opcode space. Getting rid of that looks right to me, have some totally different 16 bit ISA if you must, but don't compromise the 32 bit one for it.
- Pet_Ant 3y ago> They're also saying that if you move forward with the C extension in RVA23 now there's no real backing out of it. That doesn't seem correct. I think adding and dropping C for desktop/server workloads would be relatively easy. Most of what will be run on it is either open source (Linux, Apache et al) or Java/Python/Go/.Net. Either way, I'd expect Oracle or somebody to support both with a single installer. This isn't x86 where there is a lot of binaries with no source, or lots of janky code that assumes x86 that we need backward compatibility. (Note: IIRC RV32A is the "application" profile, not for embedded where hand tuned assembly is a real thing, things are much more fragile there). That said, just like Linux supported multiple x86 based platforms (PC-98), I'd imagine Debian and others would support non-C processors with distros, so I don't think it would really hurt Qualcomm if it's kept and they don't include it. > They also say that there's lots of implementations with C in already, so backing out of it now disadvantages those implementations. Ugh. So we should hold hold onto something, even if it is a bad idea just because other people wasted time on it? That seems like a very crab-bucket mentality. Not saying it should be removed, but the decision should be technical, not based on favoring certain players.
- dezgeg 3y agoI doubt many people would consider it easy. How many people running their stuff on EC2 would like to hear at some point that to upgrade the newest instance type you need to remake your VMs/containers?
- Pet_Ant 3y agoI mean it's irritating, but we all have CI/CD pipelines now don't we? I'd see it being a project for a single team for a month to get it changed, tuned, and verified. We regularly build both x86 images and ARM images for devs on Macs. It's really not that difficult when you aren't redoing manual steps.
- StillBored 3y agoDebian maybe, but then again maybe not. The distros don't like supporting a lot of separate arch revisions because they tend to behave like different arches. That is why most of them have dropped 32-bit arm support, it was from a distro perspective a completely separate arch despite being able to run on much the same HW as the 64-bit arm distro. Given most arm devices made in the past ~decade have been 64-bit it was an obvious choice. People with 32-bit binary apps can run them on the 64-bit distro, and the maintainers don't have to keep building/testing/fixing an entirely separate set of machine images. So, if someone forks the arch such that two different distros are required based on HW, its just going to fragment the distro's too because some of them will just pick one or the other profile.
- ribit 3y agoWhat I find particularly interesting is that SiFive never actually built a high-performance CPU core. Their highest performing IP offers IPC closer to ARM's current efficiency cores. And SiFive always kept very vague about other metrics about their CPU cores (like power consumption).
- brucehoult 3y agoSiFive announces a new core with significantly increased performance mostly every October(ish), and has been doing so every year(ish) since U74 (the core in the fastest currently shipping SoCs) succeeded U54 in October 2018. They have never build a core comparable to the fastest current Arm, x86 cores because they haven't been moving up the performance curve for very long. Just five generations at this point. October 2017: U54, almost A53 competitive despite being single-issue October 2018: U74 (dual issue), A55 class October 2019: U84 (OoO), A72 class June 2021: P550, A76 class December 2021: P650, A78 class October 2023: P870, Cortex-X3 class SiFive can't talk a lot about power consumption because that depends not only on the core design but the entire SoC, the process node, the corner of the process node, the physical design and many other things that are under the control of SiFive's customers, not SiFive.
- phkahler 3y agoI've always disliked the RISC-V instruction encoding. The C extension was an after thought and IMHO could have been done better if it were designed in from the start. I'm also a fan of immediate data (after the opcode), which for RISC-V I would have made come in 16,32,64 bit sizes. The encoding of constants into the 32-bit instruction word is really ugly and also wastes opcode space. After all the Vector, AI, and graphics stuff is hashed out I'd like to see a RISC-VI with all the same specs but totally redone instruction encoding. But maybe that's just me.
- loup-vaillant 3y ago> The C extension was an after thought and IMHO could have been done better if it were designed in from the start. Not sure what you mean here: opcode space has to have been reserved for the C extension from the start, that part can’t have been an afterthought. It may have been badly designed still, but if so that must be for other reasons (working from a bad code sample is often cited). > The encoding of constants into the 32-bit instruction word is really ugly and also wastes opcode space. It kinda has to be to minimise fanout, and with it propagation delays and energy consumption. As a software guy I recoil in horror, but I can’t argue against faster and more efficient decoders. https://www.youtube.com/watch?v=a7EPIelcckk https://www.youtube.com/watch?v=a7EPIelcckk > I'm also a fan of immediate data (after the opcode), which for RISC-V I would have made come in 16,32,64 bit sizes. So was I, before I read the RISC-V specs. One possible disadvantage of separate immediate data is wasting instruction space (many constants are so much closer to zero than 127), making alignment issues even worse, and it could increase decoding latency. I would definitely do this for bytecode for a stack machine meant to be decoded by software, but for a register machine I want to instantiate in an FPGA or ASIC, I would think long and hard before making a different choice than RISC-V.
- oconnor663 3y ago> The C extension was an after thought I understand that "afterthought" can be more of a subjective comment on the design than a concrete claim about the order of events, but still I'll quote directly from the RISC-V Instruction Set Manual: > Given the code size and energy savings of a compressed format, we wanted to build in support for a compressed format to the ISA encoding scheme rather than adding this as an afterthought
- u320 3y agoThe RISC-V community seeing an influx of large interests, used to ARM-style architectures might create conflicts. But it is really a sign of RISC-V winning.
- pclmulqdq 3y agoAs a chip designer who has made a few RISC-V cores (including one open-source one that nobody uses), I personally hate the C instructions, and I am on Qualcomm's side here. There are just too many of them, and they really muck up instruction decoding without providing large benefits for anything but the smallest MCUs. Maybe I should weigh in on this issue in the official channels.
- hajile 3y agoConsider the Linux kernel code getting 50% larger when you move from compressed to uncompressed instructions[0]. That puts RISC-V as among the least efficient ISAs out there and would make it unsuitable for most applications. [0] https://people.eecs.berkeley.edu/~krste/papers/EECS-2016-1.pdf https://people.eecs.berkeley.edu/~krste/papers/EECS-2016-1.p...
- pclmulqdq 3y agoAt the scale of Linux, you don't care about code size very much, but you care a lot about working set size. 2-20% seems to be the range of working set size reductions you see in the literature, and if you compensate with other instructions, you can get back a lot of that code size. The analysis from the SiFive folks generally doesn't include that compensation factor: it just involves a straight find-and-replace in the binary.
- hajile 3y agoYour assertion has two major issues. First, you context switch a lot to and from the Linux kernel, so decreased cache pressure does matter. Second, if you have proof that loops predominantly consist of 32-bit instructions, prove your case. To my mind, a loop is likely to use fewer registers and likely to have shorter branches and smaller immediate values. all of these seem to favor compressed instructions actually favoring working code even MORE than general code.
- pclmulqdq 3y ago
- OhMeadhbh 3y agoNot going to happen. All the tech working groups in the foundation report up through Yunsup who's the benevolent tech dictator for life. Changing the spec wouldn't go through without his approval. And compressed instructions check a check box in the low end cores SiFive is selling. Code density is pretty crappy otherwise. Or at least it's crappy compared to 8051s which is where they're competing against in some corners. The low-end RV16 or RV32 are a little more competitive against Cortex M0/3/4s, but even then only when you have the compressed instruction extension.
- bonzini 3y agoThose low end cores do not implement the RVA or RVB profiles.