6 ms·
So how would you suggest re-writing their example in less than 6 instructions for RISC-V? X86/arm both have instructions that include the carry operation for lo
by robert_foss 5y ago
So how would you suggest re-writing their example in less than 6 instructions for RISC-V? X86/arm both have instructions that include the carry operation for long additions, and only require 2 instructions.
- vitno 5y agoAny != All. There is a difference between synthetic benchmarks and real world test cases.
- api 5y agoSo this person found a pathological case for the RISC-V instruction set?
- adrian_b 5y agoThis is not a pathological case, it is normal operation. A computer is supposed to compute, but the RISC-V ISA does not provide everything that is needed for all the kinds of computations that exist. The 2 most annoying missing features are the lack of support for multi-word operations, which are needed to compute with numbers larger than 64 bits, but also the lack of support for detecting overflow in the operations with standard-size integers. If you either want larger integers or safe computations with normal integers, the number of RISC-V instructions needed for implementation is very large compared to any other ISA. While there are people who do a lot of computations with large numbers, even the other users need such operations every day. Large number computations are needed at the establishment of any Internet connection, for the key exchange. For software developers, many compilers, e.g. gcc (which uses precisely libgmp), do computations with large numbers during compilation, for various kind of optimizations related to the handling of constants in the code, e.g. for sub-expression extraction or for operation complexity lowering. So every time when some project is compiled, libgmp or other equivalent library for large numbers might be used, like also every time when you click on a new link in a browser. So this case is not at all pathological, except in the vision of the RISC-V designers who omitted support for this case. That was a good decision for an ISA intended only for teaching or for embedded computers, but it becomes a bad decision when someone wants to use RISC-V outside those domains, e.g. for general-purpose personal computers.
- panick21_ 5y ago> A computer is supposed to compute, but the RISC-V ISA does not provide everything that is needed for all the kinds of computations that exist. This is non-sense. You can still do everything you need. Its just that in some cases the code size is a bit bigger or smaller. And the difference with compressed instruction is not nearly as big, if you add fusion the difference is marginal. So really its not a pathological case its 'its slightly worse' case and even that is hard to prove in the real world given the other benefit RISC-V brings that compensate. And we can find 'slightly worse case' in the opposite direction if we would go looking for them. If you gave 2 equal skilled teams 100M and told them to make the best possible personal computer chip, I would bet on the RISC-V team winning 90 times out of a 100.
- imtringued 5y agoThe assembly code in the email is trivial. You don't seem to understand that the carry bit dependency exists regardless of the architecture. So ultimately, just fetching more instructions is enough to achieve optimal performance. As others said, code density of RISC-V is very reasonable on average. It's not significantly worse than x86 across an entire binary.
- waterhouse 5y ago(RISC-V fan here) This is a real-world use case. GMP is a library for handling huge integers, and adding two huge integers is one of the operations it performs, and the way to do that is to add-with-carry one long sequence of word-sized integers into another. It's not synthetic; it's extremely specialized, but real.
- smoldesu 5y agoI don't think you're supposed to. The compiler handles that stuff, ideally RISC-V is just another compilation target.
- masklinn 5y agoDid you misunderstand the issue entirely? The context here is the implementation of one of the inner loops of a high-performance infinite-precision arithmetic library (GMP), in RISCV the loop has 3x the instruction count it has in competing architectures. “The compiler” is not relevant, this is by design stuff that the compiler is not supposed to touch because it’s unlikely to have the necessary understanding to get it as tight and efficient as possible.
- brucehoult 5y agoAn actual arbitrary-precision library would have a lot of loops with loops control and load and stores. Those aren't shown here. Those will dilute the effect of a few extra integer ALU instructions in RISC-V. Also, an high performance arbitrary-precision library would not fully propagate carries in every addition. Anywhere that a number of additions are being done in a row e.g. summing an array or series, or parts of a multiplication, you would want to use carry-save format for the intermediate results and fully propagate the carries only at the final step.
- jolmg 5y agoI don't even see the issue. RISC-V is supposed to be a RISC-type ISA. It's in the very name. That it takes more instructions when compared to a CISC-type ISA like x86 is completely normal. https://en.wikipedia.org/wiki/Reduced_instruction_set_computer https://en.wikipedia.org/wiki/Reduced_instruction_set_comput...
- theresistor 5y agoThe argument for RISC instructions (in high performance architectures) is that the faster decode makes up for the increase in instruction count. The problem is that a faster decode has a practical ceiling on how much faster it's going to make your processor, and it's much lower than 3x. If your workload is bottlenecked on an inner loop that got 3x larger in instruction count, no 15% improvement in decode performance is going to save you.
- jolmg 5y agoI don't know what the design goals of RISC-V were, but I would guess performance is not the key goal or at least not the only goal. It makes more sense that ease of implementation is a more important goal, if they want to make adoption easy. That's another argument for favoring RISC over CISC.
- monocasa 5y agoIf that's the case, you can always stick a uop cache in after the decoder.
- snvzz 5y agoAmount of instructions matters much less if they can be fused into more complex instructions before execution. RISC-V was designed with hindsight on fusion, thus it has more opportunities for doing it, and doing it at a lower cost. And, due to the very high code density RISC-V has, the decoder can do its job while not having to look at a huge window.
- saagarjha 5y agoOk, but where it the chip that can fuse these?
- rbanffy 5y agoI don't think there is anything preventing the processor to fuse those instructions into a single operation once they are decoded.
- lordnacho 5y agoHow does the instruction fusion work? It seems to be mentioned in the article and by a couple of other commenters.
- volta83 5y agoThe CPU executes the two (or more) dependent instructions "as if" they were one, e.g., in 1 cycle. The CPU has a frontend, which has a decoder, which is the part that "reads" the program instructions. When it "sees" certain pattern, like "instruction x to register r followed by instruction y consuming r", it can treat this "as if" it was a single instruction if the CPU has hardware for executing that single instruction (even if the ISA doesn't have a name for that instruction). This allows the people that build the CPU to choose whether this is something they want to add hardware for. If they don't, this runs in e.g. 2 cycles, but if they do then it runs in 1. A server CPU might want to pay the cost of running it in 1 cycle, but a micro controller CPU might not.
- cpeterso 5y agoDo RISC-V specs document which instruction combinations they recommend be fused? Sounds like the fused instructions are an implementation detail that must be well-documented for compiler writers to know to emit the magic instruction combinations.
- __init 5y agoIt generally goes the other way around -- programmers and compilers settle on a few idiomatic ways to do something, and new cores are built to execute those quickly. Because RISC-V is RISC, it seems likely that those few ways would be less idiomatic and more 'the only real way to do x', which would aid in the applicability of the fusions.