5 ms·
Yes exactly, the 'existence proof of a competitive architecture using exclusively 32-bit instructions' has often been reference. Qualcomm's proposal is all ins
by gchadwick 3y ago
Yes exactly, the 'existence proof of a competitive architecture using exclusively 32-bit instructions' has often been reference.
Qualcomm's proposal is all instructions are aligned to their size. Initially that means everything is a 32-bit instruction, now with a lot more green-field encoding space to play with (so less need to have larger instructions). 64-bit instructions would be introduced (aligned on a 64-bit boundary) when needed with the expectation they'd be used for rare operations and 48-bit instructions wouldn't happen.
The SiFive (and original RISC-V architects view) is RISC-V is meant to be a variable length instruction set and a mix of 16/32/48 provides better static code size along with better dynamic code size meaning smaller icaches needed, smaller buffers in fetch units etc.
Interesting that the architecture that was meant to be a 'purer' RISC implementation than ARM is pushing towards the more CISC style variable length instructions. In a sense Qualcomm are trying to keep it closer to the RISC ideal!
- _a_a_a_ 3y agoWhy would you need a 64 bit instruction; what kinds of things are going to be used for it? What does 'rare' mean here, does it mean rare in execution, or rarely appears in code? (The difference being that something might only appear once in your code but be part of your hot loop so be executed any number of times) If they are rare in execution, what is their value over composing them of 32-bit instructions, where the (rare) overhead of doing so would be typically a amortised away? (The only thing I can think of that 64 bit instruction seem suited to is some kind of internal CPU management instructions, but context switches etc. are relatively rare & very expensive anyway so... I don't know)
- hajile 3y agoPersonally, I like the idea of doubling the instruction length every time -- 16, 32, 64, 128, etc. There's a big use case on the longer instruction end for VLIW/DSP/GPU applications.
- Pet_Ant 3y agoAFAIK you want short instructions for VLIW because you want to pack multiple of them into a single word.
- deleted 3y ago[deleted]
- camel-cdr 3y agoFrom the RVI thread on 48 bit instructions, 64 bit ones would probably look similar: > There are several 48-bit instruction possibilities. > 1. PC-relative long jump > 2. GP-relative addressing to support large small data area, effectively giving GP-relative access to entire data address space of most programs > 3. Load upper 32-bits of 64-bit constants or addresses > 4. Or lower 32-bits of 64-bit constants or addresses > 5. And with 32-bit mask > 6. More effective ins/ext of 64-bit bit fields Another thing thats offten discussed is moving the vtype and setvl into each vector instructions, I'm not sure if that requries 48 or 64 bit instructions.
- _a_a_a_ 3y agoI was really asking about 64-bit instructions specifically, but going with what you've put, if you don't mind... > 1. PC-relative long jump My understanding is that these are rare > 2. GP-relative addressing to support large small data area, effectively giving GP-relative access to entire data address space of most programs What is 'GP' here? but "...access to entire data address space of most programs" In this case you are just going to be bouncing all over the address space, substantially missing any level of cache much of the time, surely?. Maybe you get a little extra code density but you aren't going to get any extra speed to speak of. > 3. Load upper 32-bits of 64-bit constants or addresses > 4. Or lower 32-bits of 64-bit constants or addresses > 5. And with 32-bit mask Well yeah, but how common is this? I understand the alpha architecture team looked at this and found it uncommon which is why they were okay with less-than-32-bit constants. If it really speeded things up you might build a specific cache to store constants (a kind of larger, stupider, register set). It would seem a simpler solution. I'm not sure what you mean with 6, and I'm not familiar with vtype/setvl
- classichasclass 3y ago> My understanding is that these are rare Depends on how many bits you had to start with. On Power ISA they aren't common either, but when they happen you need up to seven instructions (lis, ori, rldicl, oris, ori, then for branches mtctr/b(c)ctr) to specify the new address or larger value. Most other RISCs are similar when full 64-bit values must be specified. This is a significant savings.
- benj111 3y agoWell you can embed longer immediates directly in the opcode. You could have a lot more registers. The first example, I'm not sure you'd want a full 64bit encoding space. You still aren't going to be able to load a 64bit immediate directly so I'd rather see an instruction that uses the next instruction as the immediate. But then 50% of the time you're still going to be padding this to 64bit alignment, so it's unclear to me that this is a benefit over 2 lots of the same but with 32bit immediates. The second option is interesting. But if you've got 256 addressable registers say, what use are the 32 and 16 bit instructions that can only address a tiny proportion of those registers.
- hajile 3y agoThe sweet spot for scalar code is about 24 registers, but that leads to weird offset-bits (there's an ISA that does this, but I forget what it's called), so 32 registers is easier to implement and provides a mild improvement in the long tail of atypical functions. On the flip side, the ability to have more registers is very good for SIMD/GPU applications.
- benj111 3y agoAbsolutely, I'm not saying a 64bit instruction length with 5/6/7/8 bits of registers would be bad per se. In fact I'd be interested to see where it leads. But if you have a processor that also uses 16 bit instructions those extra registers become unusable. Thumb can't encode all registers in all instructions so you have the high registers that are significantly less useful than the low registers. X86 is the same, never really done 64bit ASM so I don't know if they improved that. So then you may aswell just divide up the registers so you've got 16 general purpose registers and 16 registers for simd or whatever.
- Joker_vD 3y agoHow do you even use all those registers? Serious question. I've toyed with a couple of 256-register ISAs, and the moment you hit function calls/parameter passing you realize that to utilize those efficiently, you really need some way to indirectly refer to registers, be it register windows, or MMIX's register slide, or Am29k's IPA/IPB/IPC registers; the only other option seems to be to perform global register allocation but that hardly works in scenarios with separate compilation/dynamic code loading.
- classichasclass 3y agoPower10 added "prefixed" instructions, which are effectively 64-bit instructions in two 32-bit halves (the nominal instruction size). They are primarily used for larger immediates and branch displacements. https://www.talospace.com/2021/04/prefixed-instructions-and-more-in.html https://www.talospace.com/2021/04/prefixed-instructions-and-...
- TanjB 3y agoMIPS had load const to high or low half. More that 40 years ago Transputer had shift-and-load 8 bit constants. Lots of ancient precedents for rare big constants.
- classichasclass 3y agoSo does classic PowerPC, SPARC, and many other ISAs. It's the most common way to handle it on RISC. The Power10 prefixed instruction idea just expands on it.
- camel-cdr 3y ago> Interesting that the architecture that was meant to be a 'purer' RISC implementation than ARM is pushing towards the more CISC style variable length instructions. In a sense Qualcomm are trying to keep it closer to the RISC ideal! The initial idea of RISC-V was pretty much, a variable length RISC isa, but sane and easy to decode. That is not the x86, "we need to add yet another prefix".
- bunnie 3y agovariable length instructions == CISC is not quite a correct equivalence. The compressed 'C' extention is designed to re-use a lot of the existing decode infrastructure. On RV32 the C instructions are a strict subset of the full length instructions, so at least on RV32 it is very light weight, and it adds barely any logic to a core. It's almost always worth it to turn on C extensions versus making the cache bigger or trying to speed up main memory. In my experience I-cache pressure is real, especially on lightweight implementations that don't have multiple levels of cache hierarchy and huge amounts of associativity to reduce the impact of an instruction cache miss. I have played with both C and non-C variants, and also played with compiler tuning that saves code size versus 'performance' (which includes loop unrolling and thus more I cache misses). Generally smaller code size is better for power and system complexity, while keeping performance at par. Of course if you aren't as restricted on power or complexity (as is the case on a high end CPU), the calculus is different. This kind of simplicity to me embodies the heart of RISC. If your CPU is hitting cache lines more often, you don't have to speculate as deep, don't have to re-order as much, thus less logic, less complexity, less power, higher clock rates and fewer side channels. On the other hand, I suppose if you are already committed to deep speculation and out of order, compressed instructions might extract a disproportionate cost, maybe less so in decode and more so in precise exception handling and in tricks like register renaming.
- gchadwick 3y ago> variable length instructions == CISC is not quite a correct equivalence. Yeah I'd tend to agree, in particular x86 variable length encoding is a lot more complex than the RISC-V encoding! What I'm really getting at is CISC and RISC aren't well-defined things and it's interesting seeing how the design of RISC-V is getting pulled in different directions. > so at least on RV32 it is very light weight, and it adds barely any logic to a core. It's almost always worth it to turn on C extensions versus making the cache bigger or trying to speed up main memory. Definitely, and the Qualcomm proposals are that things should stay that way for RV32/low end in general. It's high-end RV64 they care about. > On the other hand, I suppose if you are already committed to deep speculation and out of order, compressed instructions might extract a disproportionate cost, maybe less so in decode and more so in precise exception handling and in tricks like register renaming. This is the root of it. It's easy enough to do a study demonstrating changes in static code size, also easy enough to build a low-end RISC-V processor and examine the trade-offs. It's all a lot more complex at the higher-end especially as high-end RISC-V cores are far from mature.
- ilyt 3y agoThey want to have the cake and eat it too with same instruction set fitting "small" (few hundred kB of flash/RAM at most) and "big" (Linux kernel running devices and up) ones. IMO it's futile effort that unnecesarily taxes the big codes.
- nickik 3y agoIf this is such a big problem, why have the other RISC-V high performnace people never made this into a big issue? This really just seems to be Qualcomm wanting it to be more like ARM so they can use their existing cores. That seem pretty clear from what they are proposing.