7 ms·
I very much believe overflow checking should be on for release builds, and not just debug. Overflow is defined by the language as an illegal operation, but if i
by pslam 10y ago
I very much believe overflow checking should be on for release builds, and not just debug. Overflow is defined by the language as an illegal operation, but if it's only enabled by debug builds, then it's effectively defined as "somewhat" illegal. I think this gives the community the wrong impression, and a firmer stance should be taken. This is a killer feature, in my opinion.
In practice, for anything which isn't a micro-benchmark, it's a negligible performance hit. For things resembling a micro-benchmark, you can specifically use wrapped-arithmetic operations, or disable it for that module.
I see a strong correlation between people complaining about overflow checking overhead, and people who don't actually care about the safety features.
- kibwen 10y agoThe operations themselves may only impose a performance hit of 10% or less, but the optimizations that those checks inhibit can easily cascade into a 2x slowdown, which isn't acceptable if Rust wants to compete with C++. As for your correlation, I care a lot about safety, but I also recognize that the safest language in the world isn't going to pose any real-world improvement if it's not capable of competing with less-safe languages in their entrenched domains.
- adrianN 10y agoMany people complaining about overflow checking either work in high performance computing, where it really matters how fast you can add two numbers, or in embedded systems where you can't afford either the extra space for the instructions, or the extra execution time.
- Manishearth 10y ago> Overflow is defined by the language as an illegal operation Is it? What do you mean by illegal here? Overflow can't cause UB madness in Rust like it does in C; the compiler does not optimize loops assuming no overflow.
- vardump 10y agoUB can be confusing, but please remember it also means a lot of very reasonable optimization opportunities. Say in data flow analysis, when the compiler can assume "(value + 1) > value" to be always true, a whole expensive branch can be removed. If the overflow is never going to happen anyways, the savings are considerable (expensive branch eliminated) and the result is always correct.
- Manishearth 10y agoI know that UB can be useful for optimizations. Where did I assert otherwise? I'm saying that overflow is not undefined in Rust (the very first section of the blog post says this). I made this same mistake a few days ago and thought it was undefined (and implementation-defined to panic on debug), but I was wrong. Overflow is well defined in Rust.
- pcwalton 10y ago> In practice, for anything which isn't a micro-benchmark, it's a negligible performance hit. Citation needed? Every paper I've seen has indicated a non-negligible performance hit in real apps. > I see a strong correlation between people complaining about overflow checking overhead, and people who don't actually care about the safety features I'm a counterexample.
- vardump 10y agoIt would be interesting to have support for saturated types as well. u8saturated 200 + u8saturated 200 = 255. That's often desired behavior. Saturated arithmetic is supported in hardware in SSE. Although not much point if the values need to be constantly shuffled between SSE and scalar registers...
- Manishearth 10y ago.saturating_add() exists. You can create a wrapper type around this and implement the Add/Sub/Mul traits.
- vardump 10y agoDoes it have branchless codegen?
- Manishearth 10y agoNot right now (I'm surprised that's the case, there should be intrinsics for it), but that could change.
- dbaupp 10y agoIt gets optimised to be branchless if LLVM thinks that will be faster, yes.
- nickpsecurity 10y agoParent might mean this: http://danluu.com/integer-overflow/ http://danluu.com/integer-overflow/ It starts with 2x like kibwen says. Then says those ops will only be small part of workload or something. Extrapolating lead to 5%. Article goes on from there. Not sure if it applies here but it matches the claim you quoted of micro vs macro benchmark.
- zAy0LfpBZLC8mAC 10y agoI think the only sensible thing to do is to have distinct operators for integer arithmetic (where the impossibility to represent a result obviously is an error that should be caught) and wraparound arithmetic that you can use if you actually need wraparound semantics (or where you need the speed and can prove that it does give you the correct result better than the compiler can).
- Alphasite_ 10y agoSounds like what swift does, checked and unchecked arthritic operators.
- vog 10y agoWhy do people invent these strange terms? "wraparound arithmetic". "unchecked arithmetics". Really? The correct term is modular arithmetic, which is well established in maths for centuries. And it is properly backed by a really nice and almost dead-simple theory (at least from a programmer's point of view). Is this no longer taught in school? If you like to emphasize the fact that it's mostly modulo a power of two (2^n), then call it two's complement. However, that overemphasizes the binary representation, which I find most of the time more distracting than helpful. As a side note, integer division is usually not performed in modular arithmetics, but addition, subtraction and multiplication are.
- zAy0LfpBZLC8mAC 10y ago(2147483647 + 1) % m = -2147483648 What is the m?
- periodontal 10y ago2147483647 + 1 is congruent to -2147483648 modulo 4294967296.
- khedoros 10y ago> Why do people invent these strange terms? Same reason we talk about a "for loop" or a "while loop", instead of just a "conditional branch". Different words add different context to the concept that they're based on. For example, "modular arithmetic" doesn't include the context of checking for undesired leaks in Rust's arithmetic implementation. "Unchecked arithmetic" does.
- chewbacha 10y agoBut the article specifically states that you can enable it globally. Meaning that you can either do it or not do it. And sometimes you want overflow, so you can specifically allow it at a function level.
- Gibbon1 10y agoI get hammered every time I mention this, but what happens when a primate type overflows should be handled by the type system. IE, by default overflow checked and is an error. But you can declare a type with something like twos_complement. And the compiler will implement the expected behavior. Or if you want no you can disable overflow checks. This stuff really should not be global behavior. Because 90% of the time the cost of overflow checks is nil so you should do them.
- vardump 10y agoThat can easily lead to excessive branch target buffer invalidation and mess up branch prediction. It might look acceptable in microbenchmarks, with just 1-10% performance loss and be a total disaster in an actual running system. A mispredicted branch costs about 15 clock cycles. You'll have a lot of those when CPU can't do bookkeeping on all of your range check branches anymore. The range checks themselves might run fast and fine, but those other branches your program has to execute can become pathological. Those, where the default choice for branches not present in CPU branch prediction buffer and branch target buffer is wrong. 15 clock cycles is rather excessive on any modern CPU that's supposed to execute 1-4 instructions per clock cycle! In the worst pathological cases this can mean an order of magnitude (~10x) worse performance. Of course data flow analysis might mitigate need for actual range checks, if the compiler can prove the value can never be out of range (of representable values with a given type). Regardless it is important to understand worst case price can be very high, this type of behavior should not be default in any performance oriented language.
- vardump 10y agoTo add, to avoid any misunderstanding: the cost is probably within 1-10% in likely scenarios on most real world code. Pathological case would be code with a lot of taken branches mixed with range/overflow checks. For range checks, any sane compiler would generate code that doesn't branch if the value is in range. If no branch predictor entry is present, branches are assumed [1] to be not taken. These branches can still thrash BTB (kind of cache trashing) and cause taken branches that fell out of BTB to become always mispredicted, thus very expensive. [1]: Some architectures do support branch hints, even x86 did so in the past. Modern x86 CPUs simply ignore these hints.
- pcwalton 10y agoNot only that, you lose the ability to pattern match simple arithmetic to LEA and bloat your code size.
- giovannibajo1 10y agoIt's very hard for an overflow check to be mis-predicted. On x86, for static branch prediction, it's sufficient to generate code with a forward branch to be predicted as unlikely the first time. From that point, as the branch will never be taken (unless it actually triggers a panic but then who cares), it will keep being predicted as not taken. A correctly-predicted non taken conditional jump on x86 has 0 latency. In the specific case of overflow, you don't even need to generate a cmp instruction. It basically means that's close to a wash, in performance. The actual negative effect that could be measurable is disabling optimizations (like composition of multiple operations, as you need to check each one for overflow).