8 ms·
I am not quite understanding the missed improvement you are talking about. Are you saying that instead of (or in addition to) "branch if equal", "branch if les
by thethirdone 7y ago
I am not quite understanding the missed improvement you are talking about.
Are you saying that instead of (or in addition to) "branch if equal", "branch if less than", etc there should be a "set if equal", "set if less than", etc? In case you are, there is a "set if less than" instruction. The floating point extension does have a "set if equal" instruction despite there not being an integer equivalent.
Alternatively if you are complaining about 1 representing True rather than ~0, I don't think that matters much from a technical POV.
- ncmncm 7y agoThe condition codes from comparison operations have always been a problem for out-of-order execution units. Feeding the result directly to the register bank sidesteps the problem. Conditional-move instructions have also been a problem, which this also sidesteps. Branch-register-zero and -nonzero instructions complete the set. The extra step to get a comparison result into a register, and another to extend it to the whole register, is an inefficiency I would prefer to leave behind.
- jleahy 7y agoIt's a single instruction to then subtract 1 from the result of SLT/SLTI, which then gives you 0xffffffff or 0x00000000 rather than 0 or 1, which you can use as a mask. RISC-V was designed to be a good target for a C compiler, given that's what people almost always need. Compilers frequently see code line `int x = (a > b)` and SLT is perfect for that. Compilers don't do clever masking tricks to avoid branches, in general (there are a few hard-coded tricks like divide by signed constant).
- ncmncm 7y agoIn other words, two more instructions, instead of zero. That compilers fail to do clever masking tricks to avoid branches frequently results in 2x slower programs.
- avianes 7y agoNo, as jleahy said, it cost a single instruction to transform 0 to 0b00..00 and 1 to 0b11..11 - SUB X, $0, X. And your proposal - using the zero register as a flag register - breaks way more micro-optimizations (like the simple ones described in "The RISC-V Reader" book) than it allows
- bonzini 7y ago> Compilers don't do clever masking tricks to avoid branches, in general They actually do a lot of such tricks. Check out the GCC sources, for example tree-ssa-phiopt.c[1] and ifcvt.c[2] [1] https://github.com/gcc-mirror/gcc/blob/master/gcc/tree-ssa-phiopt.c https://github.com/gcc-mirror/gcc/blob/master/gcc/tree-ssa-p... [2] https://github.com/gcc-mirror/gcc/blob/master/gcc/ifcvt.c https://github.com/gcc-mirror/gcc/blob/master/gcc/ifcvt.c
- thethirdone 7y ago> The condition codes from comparison operations have always been a problem for out-of-order execution units. RISC-V doesn't have condition codes. > The extra step to get a comparison result into a register, and another to extend it to the whole register, is an inefficiency I would prefer to leave behind. Can you provide RISC-V assembly that demonstrates this problem? Common branching patterns in RISC-V are all a single instruction. The core branching comparison instructions: `a >= b` maps to `bge a, b, label` `a = b` maps to `beq a, b, label` `a != b` maps to `bne a, b, label` `a < b` maps to `blt a, b, label` Comparisons with 0 just use the x0 register and for <= and > just flip the order of the operands.
- ncmncm 7y agoI know it does not have condition codes. Instead, it has a mess of very complex compare-and-branch instructions.
- Taniwha 7y agoIt has exactly 6 = != < >= and signed < >= - hardly "a mess" or particularly "complex"
- ncmncm 7y agoThese do not seem, to you, complex instructions? I boggle.
- pjc50 7y agoPretty much all architectures have those as a natural consequence of the way 2's complement arithmetic works? Even the tiny 6502 spends eight of its 256 opcodes on eight different branch-if-comparison opcodes.
- ncmncm 7y agoThe RISC-V compare-and-branch instructions are objectively a great deal more complex than the 6502 instructions.
- jabl 7y agoYou're saying that you want an isa where a branch is done something like Cmp rx ry rz # compare rx ry, write result to rz Jmple rz, ra, imm # jump +- ra+imm if rz < 0 Which would avoid the use of a implicit condition register, as well as the short offsets that a combined compare+jump suffers from?
- ncmncm 7y agoNo. That is close to what existing chips do. r31 = (rx<=ry)? ~0:0 (or <, etc.), result in r31 implicitly. Followed by pc = r31? pc + n : pc i.e. a single conditional branch instruction, for all conditions; or rx &= r31 effectively a conditional move instruction, or rx -= r31 effectively a conditional increment. I.e., actually reduced, but more powerful. (Implicit destination followed by mov is free, because of register renaming.) Decoding the RISC-V compare-and-branch instruction is a whole project.
- Taniwha 7y ago9 lines of verilog in my decoder, 3 more in the branch unit - about the same as for, say jalr, or add immediate
- avianes 7y ago> (Implicit destination followed by mov is free, because of register renaming.) There is no need for an implicit destination register, just do it explicitly and you get the same result but without any magic and with more control. Register renaming can do the same job with an implicit or explicit register destination, there is no incidence on register dependencies. > r31 = (rx<=ry)? ~0:0 Reading your first post I know that you want to use 0 and ~0 instead of the traditional 0 and 1 as the result of condition statements. I don't think that's a particularly great idea because: - there are so many places where you need 1 instead of ~0; - translating 1 into ~0 only costs one "SUB" instruction. If this feature is required for a specific domain, it's sufficient to add an extension with this feature or - more drastically - to add a macro-operation fusion step to detect this pattern. > Decoding the RISC-V compare-and-branch instruction is a whole project. RISC-V encoding is made to keep decoding as simple as possible. Have you ever checked the RISC-V encoding? It looks a bit confusing at the beginning to a human but if you think about the implementation, it immediately becomes clear. In RISC-V ISA, we can criticize: the absence of a standard SIMD extension (with a big abstract vector extension instead) or the fact that some extensions require 3 read ports on the register bank instead of 2, which can make the implementation more complex. But it's trade-offs