6 ms·
EDIT: Hmm, I seem to have picked a bad example. Try this one: int get(int *base, unsigned index) {return base[index];} Arm64: update: ldr
by FullyFunctional 5y ago
EDIT: Hmm, I seem to have picked a bad example. Try this one:
int get(int *base, unsigned index) {return base[index];}
Arm64:
update:
ldr w0, [x0, w1, uxtw 2]
ret
RV64GC (vanilla):
update:
slli a5,a1,32
srli a1,a5,30
add a0,a0,a1
lw a0,0(a0)
ret
RV64GC+Zba:
update:
sh2add.uw a0,a1,a0
lw a0,0(a0)
ret
Arm64 is able to do some indexed loads in a single instruction that might take two in RISC-V w/Zba (and up to 4+ in regular RISC-V). However, calling that a win for Arm64 is not so clear as the more complicated addressing modes could become a critical timing path and/or require an extra pipeline stage. However, as a first approximation, for a superscalar dynamically scheduled implementation, fewer ops is better so I would say it's a slight win.
I don't understand the obsession with bytes. 25% fewer bytes has only very marginally impact on a high-performance implementation and the variable length encoding has some horrendous complications (which is probably why Arm64 _dropped_ variable length instructions). Including compressed instruction in the Unix profile was the biggest mistake RISC-V did and I'll die on that hill.
ADD: Don't forget that every 32-bit instruction is currently wasting the lower two bits to allow for compressed, thus any gain from compress must be offset by the 6.25% tax that is forced upon it.
- saagarjha 5y ago> 25% fewer bytes has only very marginally impact on a high-performance implementation Instruction cache doesn't come for free, and is usually pretty small on most shipping processors. It's not a big deal for smaller benchmarks, but in real-world programs this can become a problem.
- FullyFunctional 5y agoI am obviously aware and I'm here to tell you that the overhead of variable length instructions matters more. Arm agrees. M1 has a 192 KiB I$ btw. ADD: had RISC-V just disallowed instructions from spanning cache lines and disallowing jumping into the middle of instructions then almost all of the issues would have gone away. Sigh.
- saagarjha 5y agoI actually had Apple's chips in mind when talking about "most shipping processors" because they have historically invested heavily in their caches and reaped benefits from it. But not all the world's an M1, and also I'll have you know that Apple themselves cares very much about their code size, even with their large caches. Don't go wasting it for no reason! (I should also note that I am pretty on board with you with regards to variable-length instructions, this is just independent of that.)
- avianes 5y agoVariable instruction sizes have a cost, but with only 2 instruction sizes like current RISC-V that cost remains very low as long as we don't have to decode a very large number of instructions each cycle, and it gives a huge code density advantage.
- KerrAvon 5y agoHave the ARM AArch64 designers ever commented on this? They intentionally left out any kind of compressed instructions, and certainly Apple at least cares a lot about code size.
- deleted 5y ago[deleted]
- klelatti 5y agoTry this at 34:30 - from Arm’s architecture lead Richard Grisenthwaite. Earlier he says that several leading micro architects think that mixing 16 bit and 32 bit instructions (Thumb2) was the worst thing that Arm ever did. https://m.soundcloud.com/university-of-cambridge/a-history-of-the-arm-architecture-and-the-lessons-learned-while-building-it https://m.soundcloud.com/university-of-cambridge/a-history-o...
- brucehoult 5y agoHe explicitly specifies that those micro-architects are at companies OTHER than ARM. His own opinion appears to be that the worst thing ARM ever did was T2EE, designed for JIT compilers and compilers for dynamic languages. He says that by the time the chips came out compiler technology had advanced to the point that it was no longer useful and no one else used it. A couple of other points picked up in the talk: - He reverses Hennessy and Patterson wrt SPARC and MIPS. - A64 effort started in 2007. So it took 5 years to freeze/publishing, the same as RISC-V. - A64 architects thought code density is no longer important. Some people definitely disagree with that. At the time they probably thought amd64 was the only competition and matching/beating that was good enough. - he seems to be regretting the 2nd operand shift because it fell naturally out of the 1985 micro-architecture, but it's a burden now. And yet it was included in A64 -- presumably because the initial processor pipelines had it anyway, because they supported A32. But now we have A64-only CPUs. - LL/SC was the wrong thing to do.
- snek_case 5y agoNot only that, but CPUs have a maximum number of instructions they can dispatch per cycle (typically 4 or 6). Even in microbenchmarks, the difference there could show up.
- marcosdumay 5y agoThe bottleneck on that is data interdependency of your algorithm. If you break it in 6 or 10 instructions, the data dependency stays the same. (Of course, you can add unnecessary dependency with a badly designer ISA. But it's not a necessary condition.)
- snek_case 5y agoThere's still a limit to how many instructions you can decode and dispatch every cycle, even with zero dependencies. There's also definitely dependencies in the example where you're computing a memory address to access a value.
- avianes 5y agoWhy using an unsigned? It is obvious here that RISC-V without Zba takes 4 instructions because it manages special cases related to unsigned. If you use a simple int for index: slli a1,a1,2 add a0,a0,a1 lw a0,0(a0) And isolating this code in a small function puts constraints on register allocation, but if we remove this constraint then we can write: slli a1,a1,2 add a1,a1,a0 lw a1,0(a1) Which is very suitable for macro-op fusion and C extension > Including compressed instruction in the Unix profile was the biggest mistake RISC-V did and I'll die on that hill. This is so wrong. The C extension is one of the great strengths of RISC-V, it is easy to decode, very suitable for macro-op fusion, and it gives a huge boost in code density
- JonChesterfield 5y agoIirc compressed instructions are the thing that costs 2 bits per 32 and was criticised as overfitted to naive compiler output. Am I thinking of something else?
- avianes 5y agoYes, but RISC-V still has a lot of encoding-space free and the benefit of C extension is huge. It's a trade-off. I don't think RISC-V is perfect or universal, but on this point they do a pretty good job compared to other ISAs
- JonChesterfield 5y ago32 bit, 32 registers, three register code. So add r0 r1 r2 spends fifteen bits on identifying which register to use then another two on the compressed ISA. That's half the encoding space gone before identifying the op. Never thought I'd want fewer registers but here we are. If the compressed extension is great in practice it might be a win. If the early criticism of overfit to gcc -O0 proves sound and in practice compilers don't emit it then it was an expensive experiment.
- avianes 5y ago
- damageboy 5y agoThank you for writing the obvious. Instruction Byte count is the wrong metric here 100%. Instruction Count (given reasonable decoding/timing constraints) is the thing to optimize for and indeed variable length encoding is very bad.
- tsmi 5y agoInstruction byte count matters quite a lot when you're buying ROM in volume. And today, the main commercial battleground for RISCV is in the microcontroller space where people care about these things.
- knorker 5y agoFor those of us without the expertise, could you elaborate on why that is? On the one hand we have byte count, with its obvious effect on cache space used. But to those of us who don't know, why is instruction count so important? There's macro-op fusion, which admittedly would burn transistors that could be used for other things. Could you elaborate why it's not sufficient? And then the fact that modern x86 does the opposite to macro-op fusion, by actually splitting up CISC instructions into micro-ops. Why is it so bad if they were more micro-ops to start with, if Intel chooses to do this?