3 ms·
> 1) you could argue that's bad code generation, which could easily be fixed in the compiler, as `a1` could be used as the result of the `add` Yes, as I said.
by mbitsnbites 3y ago
> 1) you could argue that's bad code generation, which could easily be fixed in the compiler, as `a1` could be used as the result of the `add`
Yes, as I said.
Please note that the code was generated by GCC (trunk) that has fairly mature RISC-V support (it has had RISC-V support since 2017 at least, and the fusion optimization was suggested even earlier than that). If fusion of load/store instructions is indeed an important part of the RISC-V design, why has this not been fixed in the compiler during the last 6+ years? I take it as there is no hard canonical way that register allocation should be done in these situations, which means that HW implementations must assume that such a sequence may clobber two registers.
> 2) since late 2021 RISC-V has the `sh2add` instruction
Which is great. As you say, not really a win in terms of code density, but definitely fewer degrees of freedom and only one clobbered register. Still not as good as dedicated indexed load/store instructions, though.
BTW, while I agree that an indexed load without indexed store would be better than nothing (I actually considered that for MRISC32), I think that it would pose a problem for compilers that generally assume symmetry (addressing modes supported by loads are assumed to be supported by stores too).
> 3) using a function like this to prove anything actually verges on dishonest (and erincandescent did the same thing) because no real program where anyone cares about performance will have such a function in hot code
True that GCC happily and often transforms indexing inside loops to address increments rather than index increments (even so for architectures that has support for indexed addressing).
The point of the example, though, was to highlight that when the compiler needs to do indexed addressing, it must clobber unrelated registers, which is a problem for instruction fusion (i.e. a fused instruction does not do the same thing as a corresponding architectural instruction, as the fused instruction has more side effects).
> It gives optimal price-performance.
Probably. But this circles back to the question about what segments you target with your design. I am fully convinced that RISC-V is the best ISA choice for many segments, especially the cost sensitive segments. However I'm less convinced that RISC-V (as it is) is the optimal choice for less cost sensitive high performance segments (e.g. premium products like those from Apple, gaming riggs, compute servers, and so on). It can probably do a good job there (just like x86 can do a good job, despite all its ISA bloat and shortcomings), but if the integer ISA was designed with support for three source operands from the start, it would be an even better fit for those segments (at least that is my belief).
- brucehoult 3y ago> Yes, as I said. Not at the point I was composing my reply. > If fusion of load/store instructions is indeed an important part of the RISC-V design I don't believe that it is. Certainly no RISC-V implementations that are in the hands of customers right now do any fusion and it doesn't seem to hurt their ability to match or exceed the performance of similar Arm cores (A55, A72). I have heard that some of the companies making very high performance cores (Apple M or AMD Zen class) are incorporating some fusion. Perhaps they will provide compiler patches if required. > I think that it would pose a problem for compilers that generally assume symmetry (addressing modes supported by loads are assumed to be supported by stores too). I think they would manage. There is certainly precedent e.g. MSP430 has 4 addressing modes for src operands but only 2 for dst (register, register indirect with 16 bit offset). Plus of course every ISA where its an addressing mode not a different instruction allows "immediate" on a load but not on a store. > i.e. a fused instruction does not do the same thing as a corresponding architectural instruction, as the fused instruction has more side effects That can never happen. Fusion will simply not be done in that case. > but if the integer ISA was designed with support for three source operands from the start, it would be an even better fit for those segments And I don't think the difference would be enough to measure or care about — single digit percentage on number of clock cycles needed. And meanwhile the simpler core may allow at least the same percentage to be gained as higher clock speed, or more cores on a die, or other compensating factors.
- camel-cdr 3y ago> Certainly no RISC-V implementations that are in the hands of customers right now do any fusion and it doesn't seem to hurt their ability to match or exceed the performance of similar Arm cores (A55, A72). You can play around with OpenXianShan though, they have a few fusion targets: https://github.com/OpenXiangShan/XiangShan/blob/master/src/main/scala/xiangshan/backend/decode/FusionDecoder.scala https://github.com/OpenXiangShan/XiangShan/blob/master/src/m... Most of the targets require the same destination, so it won't be able to fuse current codegen. I suppose there is still some time before compilers need to be ready, but it's not that much. > Perhaps they will provide compiler patches if required. I hope so, btw t-head seems to be still be trying to upstream XTheadVector: https://gcc.gnu.org/pipermail/gcc-patches/2024-January/642780.html https://gcc.gnu.org/pipermail/gcc-patches/2024-January/64278...