4 ms·
This comes up a lot and I'm sympathetic to your plea, really (I enjoy fantasizing about a different reality where CPUs weren't just "machines to run C programs"
by FullyFunctional 5y ago
This comes up a lot and I'm sympathetic to your plea, really (I enjoy fantasizing about a different reality where CPUs weren't just "machines to run C programs"), but in computer architecture, what really matters for one application or a class of applications might not be important when viewed across millions of programs.
The fact is that integer operations and floating point are two completely different beasts, so much so that we have different benchmark suites for each.
Integer operations are critically latency sensitive and bagging on extra semantics doesn't come for free and for most code this would be a tax. The "overflow bit" represents an implicit result that would have to be threaded around (I'm assuming that you aren't asking for exceptions which literally nobody wants). For FP we do that, but the cost and latency of FP ops is already high so it doesn't hurt quite as much.
The RISC-V spec [1] (which I assume you have seen) already discusses all these trade offs:
"We did not include special instruction-set support for overflow checks on integer arithmetic operations in the base instruction set, as many overflow checks can be cheaply implemented using RISC-V branches. Overflow checking for unsigned addition requires only a single additional branch instruction after the addition:
add t0, t1, t2
bltu t0, t1, overflow
For signed addition, if one operand’s sign is known, overflow checking requires only a single branch after the addition:
addi t0, t1, +imm
blt t0, t1, overflow
This covers the common case of addition with an immediate operand.
For general signed addition, three additional instructions after the addition are required, leveraging the observation that the sum should be less than one of the operands if and only if the other operand is negative.
add t0, t1, t2
slti t3, t2, 0
slt t4, t0, t1
bne t3, t4, overflow
In RV64I, checks of 32-bit signed additions can be optimized further by comparing the results of ADD and ADDW on the operands."
I do think that it might have been worth adding an single instruction version for the last one (excluding the branch), but I'm not aware of it getting accepted.
[1] https://github.com/riscv/riscv-isa-manual https://github.com/riscv/riscv-isa-manual
- adrian_b 5y agoGenerating the overflow bit and storing it adds a completely negligible cost to a 64-bit adder, so touting this as a cost saving measure is just a lie, even if indeed this claim has always been present in the RISC-V documentation. Most real cases of overflow checking are of the last type. Tripling the number of instructions over a bad ISA that lacks overflow exceptions, like unfortunately almost all currently popular ISAs are, or quadrupling the number of instructions over a traditional ISA with overflow exceptions is a totally unacceptable cost. The claim that providing overflow exceptions for integer addition might be too expensive can be easily countered by the fact that generating exceptions on each instruction is not the only way to guarantee that overflows do not happen. It is enough to store 2 overflow flags, 1 flag with the result of the last operation and 1 sticky flag that is set by any overflow and is reset only by a special instruction. Having the sticky flag allows zero-overhead overflow checking for most arithmetic instructions, because it can be tested only once after many operations, e.g. at a function exit. The cost of implementing the 2 overflow bits is absolutely negligible, 2 gates and 2 flip-flops. Much more extra hardware is needed for decoding a few additional instructions for flag testing and clearing, but even that is a negligible cost compared with a typical complete RISC-V implementation. Not providing such a means of reliable and cheap overflow detection is just stupid and it is an example of hardware design disconnected from the software design for the same device. The early RISC theory was to select the features that need to be implemented in hardware by carefully examining the code generated by compilers for representative useful programs. The choices made for the RISC-V ISA, e.g. the omission of both the most frequently required addressing modes and of the overflow checking. proves that the ISA designers either have never applied the RISC methodology, or they have studied only examples of toy programs, which are allowed to provide erroneous results.
- ajb 5y agoThe extra expense is not the generation of the overflow bit, but the infrastructure needed to support a flags register, or for every instruction to be able to generate an exception. On a simple processor like a microcontroller this doesn't cost much, but it's severely hampers a superscalar or out of order processor, as it can't work out very easily which instructions can be run in parallel or out of order. The clean solution from a micro architectural point of view would be to have an overflow bit (or whatever flags you wanted) in every integer register. But that's an expense most don't want to pay.
- adrian_b 5y agoOne must not forget that on any non-toy CPU, any instruction may generate exceptions, e.g. invalid opcode exceptions or breakpoint exceptions. In every 4-5 instructions, one is a load or store, which may generate a multitude of exceptions. Allowing exceptions does not slow down a CPU. However they create the problem that a CPU must be able to restore the state previous to the exception, so the instruction results must not be committed to permanent storage before it becomes certain that they could not have generated an exception. Allowing overflow exceptions on all integer arithmetic instructions, would increase the number of instructions that cannot be committed yet at any given time. This would increase the size of various internal queues, so it would increase indeed the cost of a CPU. That is why I have explained that overflow exceptions can be avoided while still having zero-overhead overflow checking, by using sticky overflow flags. On a microcontroller with a target price under 50 cents, which may lack a floating-point unit, the infrastructure to support a flags register may be missing, so it may be argued that it is an additional cost, even if the truth is that the cost is negligible. Such an infrastructure existed in 8-bit CPUs with much less than 10 thousand transistors, so arguing that it is too expensive in 32-bit or 64-bit CPUs is BS. On the other hand, any CPU that includes the floating-point unit must have a status register for the FPU and means of testing and setting its flags, so that infrastructure already exists. It is enough to allocate some of the unused bits of the FPU status register to the integer overflow flags. So, no, there are absolutely no valid arguments that may justify the failure to provide means for overflow checking. I have no idea why they happened to make this choice, but the reasons are not those stated publicly. All this talk about "costs" is BS made up to justify an already taken decision. For a didactic CPU, as RISC-V was actually designed, lacking support for overflow checking or for indexed addressing is completely irrelevant. RISC-V is a perfect target for student implementation projects. The problem appears only when an ISA like RISC-V is taken outside its right domain of application and forced into industrial or general-purpose applications by managers who have no idea about its real advantages and disadvantages. After that, the design engineers must spend extra efforts into workarounds for the ISA shortcomings. Moreover, the claim that overflow checking may have any influence upon the parallel execution of instructions is incorrect. For a sticky overflow bit, the order in which it is updated by instructions does not matter. For an overflow bit that shows the last operation, the bit updates must be reordered, but that is also true for absolutely all the registers in a CPU. Even if 4 previous instructions that were executed in parallel had the same destination register, you must ensure that the result stored in the register is the result corresponding to the last instruction in program order. One more bit along hundreds of other bits does not matter.
- feanaro 5y ago> add t0, t1, t2 bltu t0, t1 How does this work? Isn't `bltu` simply a branch that is taken if `t0 < t1`? How does that detect addition overflow? EDIT: Ah, because the operands are `t1` and `t2`. `t0` is the result. Quack.
- deleted 5y ago[deleted]
- throwaway81523 5y agoYes I've seen that reasoning: they propose bloating 1 integer instruction into 4 instructions in the usual case where the operands are unknown. Ouch. In reality they expect programs to normally run without checking like they did in the 1980s. So this is more fuel for the criticism that RiscV is a 1980s design with new paint. Do GCC and Clang currently support -ftrapv for RiscV, and what happens to the code size and speed when it is enabled? Yes, IEEE FP uses sticky overflow bits and the idea is that integer operations could do the same thing. Integer overflow is one of those things like null pointer dereferences, which originally went unchecked but now really should always be checked. (C itself is also deficient in not having checkable unsigned int types).
- modeless 5y ago> I'm assuming that you aren't asking for exceptions which literally nobody wants I want exceptions. Why would they be a bad idea? Besides the fact that software doesn't utilize them today (because they're not implemented, chicken and egg problem)? IMO they would be as big a security win as many other complex features CPU designers are adding in the name of security, e.g. pointer authentication.