5 ms·
I'm heavily invested in RISC-V, both personally and professionally, and I think the story is much more complicated than this makes it out to be, but I'm not goi
by FullyFunctional 5y ago
I'm heavily invested in RISC-V, both personally and professionally, and I think the story is much more complicated than this makes it out to be, but I'm not going to rehash the discussion yet again.
However, I do want to point out that a real issue (especially with legacy code) is the scaled address calculation with 32-bit unsigned values. Thankfully the Zba extension adds a number of instructions that help a lot, but still would require fusion to get complete parity with Arm64
For
int update(int *base, unsigned index) { return base[index]++; }
We get
update:
sh2add.uw a1,a1,a0
lw a0,0(a1)
addiw a5,a0,1
sw a5,0(a1)
ret
Zba is included in the next Unix profile and will _likely_ be adopted eventually by all serious implementations.
EDIT: grammar and spacing
- abainbridge 5y agoI'm guess that your assembly code is RISC-V with the Zba extension. Is the non-Zba version worse than Arm64? Compiling your function with Godbolt, I get: RISC-V (no Zba) Clang - 7 instructions - https://godbolt.org/z/7znnrzxKq Arm64 Clang - 7 instructions - https://godbolt.org/z/Trv8scxad Annoyingly I can't see the code size for the Arm64 case because no output is generated if I tick the "Compile to binary" option in "Output". I have to use GCC instead: RISC-V (no Zba) Clang - 20 bytes - https://godbolt.org/z/eWfPaorcj Arm64 GCC - 24 bytes - https://godbolt.org/z/bzsPzov5h
- FullyFunctional 5y agoEDIT: Hmm, I seem to have picked a bad example. Try this one: int get(int *base, unsigned index) {return base[index];} Arm64: update: ldr w0, [x0, w1, uxtw 2] ret RV64GC (vanilla): update: slli a5,a1,32 srli a1,a5,30 add a0,a0,a1 lw a0,0(a0) ret RV64GC+Zba: update: sh2add.uw a0,a1,a0 lw a0,0(a0) ret Arm64 is able to do some indexed loads in a single instruction that might take two in RISC-V w/Zba (and up to 4+ in regular RISC-V). However, calling that a win for Arm64 is not so clear as the more complicated addressing modes could become a critical timing path and/or require an extra pipeline stage. However, as a first approximation, for a superscalar dynamically scheduled implementation, fewer ops is better so I would say it's a slight win. I don't understand the obsession with bytes. 25% fewer bytes has only very marginally impact on a high-performance implementation and the variable length encoding has some horrendous complications (which is probably why Arm64 _dropped_ variable length instructions). Including compressed instruction in the Unix profile was the biggest mistake RISC-V did and I'll die on that hill. ADD: Don't forget that every 32-bit instruction is currently wasting the lower two bits to allow for compressed, thus any gain from compress must be offset by the 6.25% tax that is forced upon it.
- saagarjha 5y ago> 25% fewer bytes has only very marginally impact on a high-performance implementation Instruction cache doesn't come for free, and is usually pretty small on most shipping processors. It's not a big deal for smaller benchmarks, but in real-world programs this can become a problem.
- FullyFunctional 5y agoI am obviously aware and I'm here to tell you that the overhead of variable length instructions matters more. Arm agrees. M1 has a 192 KiB I$ btw. ADD: had RISC-V just disallowed instructions from spanning cache lines and disallowing jumping into the middle of instructions then almost all of the issues would have gone away. Sigh.
- saagarjha 5y agoI actually had Apple's chips in mind when talking about "most shipping processors" because they have historically invested heavily in their caches and reaped benefits from it. But not all the world's an M1, and also I'll have you know that Apple themselves cares very much about their code size, even with their large caches. Don't go wasting it for no reason! (I should also note that I am pretty on board with you with regards to variable-length instructions, this is just independent of that.)
- avianes 5y agoVariable instruction sizes have a cost, but with only 2 instruction sizes like current RISC-V that cost remains very low as long as we don't have to decode a very large number of instructions each cycle, and it gives a huge code density advantage.
- KerrAvon 5y agoHave the ARM AArch64 designers ever commented on this? They intentionally left out any kind of compressed instructions, and certainly Apple at least cares a lot about code size.