7 ms·
It sure would be nice if modern CPUs had two sets of integer instructions: one that wraps, and one that triggers an exception on wrap, with zero overhead in non
by andersa 3y ago
It sure would be nice if modern CPUs had two sets of integer instructions: one that wraps, and one that triggers an exception on wrap, with zero overhead in non-wrapping case.
Then we could compile all code with the latter, except for specifically marked edge cases where wrap is a desired part of the logic.
- explaininjs 3y agoLanguage-level concern imo. https://doc.rust-lang.org/nightly/core/num/struct.Wrapping.html https://doc.rust-lang.org/nightly/core/num/struct.Wrapping.h...
- andersa 3y agoIt could be, if the hardware supported it. Consider this quote from that page: "in some debug configurations overflow is detected and results in a panic" That's not good enough. We want to always detect it! Many critical bugs are caused by this in production builds too. Solving it at the language level would require inserting branches on every integer operation which is obviously not acceptable.
- tialaramex 3y ago> That's not good enough. We want to always detect it So, select the configuration where that's the behaviour? overflow-checks = true > Solving it at the language level would require inserting branches on every integer operation Yes, so that's what you have to do if you actually want this, if you won't pay for it then you can't have it.
- andersa 3y agoWell, the whole point of my post was that I would really like a hardware feature that does it without overhead. How that would work behind the scenes, I have no idea. Not a hardware engineer.
- tialaramex 3y agoIf you really want it, use it. Hardware vendors optimise stuff they see being done, they don't optimise stuff that somebody mentioned on a forum they'd kinda like but have never used because it was expensive. Maybe if you have Apple money and can buy ARM or something then you could just express such whims, but for mere mortals that's not an option. Newer CPUs in several lines clearly optimise to do Acquire-Release because while it's not the only possible concurrency ordering model, it's the one the programmers learned, so if you make that one go faster your benchmark numbers improve. Modern CPUs often have onboard AES. That's not because it won a competition, not directly, it's because everybody uses it - without hardware support they use the same encryption in software. The Intel 486DX and then the Pentium happened because it turns out that people like FPUs, and eventually enough people were buying an FPU for their x86 computer that you know, why not sell them a $100 CPU and a $100 FPU as a single device for $190, they save $10 and you keep the $80+ you saved because duh, that's the same device as the FPU nobody is making an actual separate FPU you fools. Even the zero-terminated string, which I think is a terrible idea, is sped up on "modern" hardware because the CPU vendors know C programs will use that after the 1970s.
- weinzierl 3y agoARM actually kind of has that. The register file has an overflow flag that is not cleared on subsequent operations (sticky overflow). So instead of triggering an exception, which is prohibitively expensive, you can do a series of calculations and check afterwards if an overflow happened in any of them. A bit like NaN for floating point. From what I understand the flag alone is still costly, so we will have to see if it survives.
- spookie 3y agoThat's quite interesting and reasonable
- andersa 3y agoIt doesn't matter if triggering the exception is expensive. At that point overflow has already occured, so your program state is now nonsense, and you might as well just let it crash unhandled. Much better outcome than reading memory at some mysterious offset. If just having the ability for an exception to occur during an instruction causes overhead, that would be a big problem though. Edit to add: We need to do the check on every operation. Just going through one iteration of the loop might have already corrupted some arbitrary memory, for example. Manually inserted checks on some flag bits don't scale to securing real programs.
- weinzierl 3y agoThink about it like that. If you allow an exception you essentially create many branches with all their negative consequences. With the sticky bit you combine them to one branch (that is still expensive[1]). [1] https://news.ycombinator.com/item?id=8766417 https://news.ycombinator.com/item?id=8766417
- Findecanor 3y ago> At that point overflow has already occured, so your program state is now nonsense, and you might as well just let it crash unhandled. The trick is to check and clear the flag before any instruction that would have a side-effect, that depends on the arithmetic result. IEEE 754-compliant floating point units have a similar behaviour with NaN that is a bit more versatile: an arithmetic instruction results in NaN if any operand is NaN, but an instruction with side-effect (compare, convert or store) will raise an exception if given a NaN.
- tpolzer 3y agoThat doesn't help at all if your loop variable is a 32 bit int that your compiler decided to transform away into vectorized loads from a 64 bit pointer. But that's exactly one of the transformations that get enabled by assuming undefined overflow.
- andersa 3y agoBut in that case, it's not going to be wrapping either. We'll just read beyond the end of the buffer, which a bounds check should catch. Or perhaps I'm not thinking of the specific sequence that would 1) not wrap during modifying index and 2) not hit bounds check after. It would need to be a requirement that compilers can't upcast all your ints to 64 bit ones, do all the math, and then write them back - would need specific instructions for each size.
- eklitzke 3y agoYou can compile code with -fwrapv and for most programs the overhead is minimal (the exception being that if you're writing number crunching code, the overhead is going to be huge). For my personal projects I have -fwrapv as part of the default compiler flags, and I remove the flag in opt builds. I honestly haven't caught that many bugs using it, but for the few bugs it did catch it saved me a lot of debugging time.
- solarexplorer 3y agoMIPS for example has this. It has `addu` for normal integer addition that does not trap and `add` if you want to trap on overflows.
- fulafel 3y agox86 had overflow checking support (via OF flag and INTO insn) but it got slowed down and later dropped from 64 bit mode.