12 ms·
Optimizations Enabled by -ffast-Math
- andraz 5y agoThe question is how to enable all this in Rust... Right now there's no simple way to just tell Rust to be fast & loose with floats.
- kzrdude 5y agoIt must be clearly understood which of the flags are entirely safe and which need to be under 'unsafe', as a prerequisite. Blanket flags for the whole program do not fit very well with Rust, while point use of these flags is inconvenient or needs new syntax.. but there are discussions about these topics.
- atoav 5y agoAlso maybe you woild like to have more granular control over which parts of your program has the priority on speed and which part favours accuracy. Maybe this could be done with a seperate type (e.g. ff64) or a decorator (which would be useful if you want to enable this for someone elses library).
- andraz 5y agoMaybe... or maybe there would be just a flag to enable all this and be fine with consequences.
- mjburgess 5y ago..have you used Rust? > just a flag to enable all this and be fine with consequences Is probably the opposite of what the language is about.
- wongarsu 5y agoReasoning about the consequences along a couple functions in your hot path is one thing. Reasoning about the consequences in your entire codebase and all libraries is quite another.
- varajelle 5y agoRust is not really a fast & loose language where anything can be UB. At least it doesn't need the errno ones since rust don't use errno.
- pjmlp 5y agoThe whole point of ALGOL derived languages for systems programming, is not being fast & loose with anything, unless the seatbelt and helmet are on as well.
- jamesmishra 5y agoI don't know if most Rust programmers would be happy with any fast and loose features making it into the official Rust compiler. Besides, algorithms that benefit from -ffast-math can be implemented in C and with Rust bindings automatically generated. This solution isn't exactly "simple", but it could help projects keep track of the expectations of correctness between different algorithm implementations.
- vlovich123 5y agoRust should have native ways to express this stuff. You don't want the answer to be "write this part of your problem domain in C because we can't do it in Rust".
- pornel 5y agoThere are fast float intrinsics: https://doc.rust-lang.org/std/intrinsics/fn.fadd_fast.html https://doc.rust-lang.org/std/intrinsics/fn.fadd_fast.html but better support dies in endless bikeshed of: • People imagine enabling fast float by "scope", but there's no coherent way to specify that when it can mix with closures (even across crates) and math operators expand to std functions defined in a different scope. • Type-based float config could work, but any proposal of just "fastfloat32" grows into a HomerMobile of "Float<NaN=false, Inf=Maybe, NegZeroEqualsZero=OnFullMoonOnly, etc.>" • Rust doesn't want to allow UB in safe code, and LLVM will cry UB if it sees Inf or Nan, but nobody wants compiler inserting div != 0 checks. • You could wrap existing fast intrinsics in a newtype, except newtypes don't support literals, and <insert user-defined literals bikeshed here>
- ptidhomme 5y agoDo some applications actually use denormal floats ? I'm curious.
- im3w1l 5y agoAccidentally trigger that path now and then? Probably. Depends on subnormals to give a reasonable result? Probably rare. But what do I know...
- adrian_b 5y agoDenormal floats are not a purpose, they are not used intentionally. When the CPU generates denormal floats on underflow, that ensures that underflow does not matter, because the errors remain the same as at any other floating-point operation. Without denormal floats, underflow must be an exception condition that must be handled somehow by the program, because otherwise the computation errors can be much higher than expected. Enabling flush-to-zero instead is an optimization of the same kind as ignoring integer overflow. It can be harmless in a game or graphic application where a human will not care if the displayed image has some errors, but it can be catastrophic in a simulation program that is expected to provide reliable results. Providing a flush-to-zero option is a lazy solution for CPU designers, because there have been many examples in the past of CPU designs where denormal numbers were handled without penalties for the user (but of course with a slightly higher cost for the manufacturer) so there was no need for a flush-to-zero option.
- ptidhomme 5y agoThanks for this explanation. But yeah, I meant : are there applications where use of denormals vs. flush-to-zero is actually useful... ? If your variable will be nearing zero, and e.g. you use it as a divider, don't you need to handle a special case anyway ? Just like you should handle your integer overflows. I'm seeing denormals as an extension of the floating point range, but with a tradeoff that's not worth it. Maybe I got it wrong ?
- marcan_42 5y agoI agree with you; if your code "works" with denormals and "doesn't work" with flush-to-zero, then it's already broken and likely doesn't work for some inputs anyway, you just haven't run into it yet.
- rocqua 5y agoI found the following note for -ffinite-math-only and -fno-signed-zeros quite worrying: The program may behave in strange ways (such as not evaluating either the true or false part of an if-statement) if calculations produce Inf, NaN, or -0.0 when these flags are used. I always thought that -ffast-math was telling the compiler. "I do not care about floating point standards compliance, and I do not rely on it. So optimize things that break the standard". But instead it seems like this also implies a promise to the compiler. A promise that you will not produce Inf, NaN, or -0.0. Whilst especially Inf and NaN can be hard to exclude. This changes the flag from saying "don't care about standards, make it fast" to "I hereby guarantee this code meets a stricter standard" where also it becomes quite hard to actually deduce you will meet this standard. Especially if you want to actually keep your performance. Because if you need to start putting all divisions in if statements to prevent getting Inf or NaN, that is a huge performance penalty.
- praptak 5y agoNot only math optimizations are like this. Strict aliasing is a promise to the compiler to not make two pointers of different types point to the same address. You technically make that promise by writing in C (except a narrow set of allowed conversions) but most compilers on standard settings do let you get away with it.
- okl 5y ago> A promise that you will not produce Inf, NaN, or -0.0. Whilst especially Inf and NaN can be hard to exclude. Cumbersome but not that difficult: Range-check all input values.
- orangepanda 5y agoAs the compiler assumes you wont produce such values, wouldnt it also optimise those range checks away?
- nly 5y agoHe means simple stuff like ensuring a function like float unit_price (float volume, int count) { return volume / count; } is only called where count > 0
- gigatexal 5y agoI am not a C dev but this was a really fascinating blog post and how compilers optimize things.
- superjan 5y agoExcellent post I agree, but isn’t it more about how they can’t optimize? Compilers can seem magic how they optimize integer math but this is a great explanation why not to rely on the compiler if you want fast floating point code.
- mhh__ 5y agoIf you use the LLVM D compiler you can opt in or out of these individually on a per-function basis. I don't trust them globally.
- p0nce 5y ago+1 Every time I tried --fast-math made things a bit slower. It's not that valuable with LLVM.
- gpderetta 5y agowith gcc you can also use #pragma gcc optimize or __attribute__(optimize(...)) for a similar effect. It is not 100% bug free (at least it didn't use to) and often it prevents inlining a function into another having different optimization levels (so in practice its use has to be coarse grained).
- nly 5y agoThis pragma doesn't quite work for -ffast-math https://gcc.godbolt.org/z/voMK7x7hG https://gcc.godbolt.org/z/voMK7x7hG Try it with and without the pragma, and adding -ffast-math to the compiler command line. It seems that with the pragma sqrt(x) * sqrt(x) becomes sqrt(x*x), but with the command line version it is simplified to just x.
- gpderetta 5y agoThat's very interesting. The pragma does indeed do a lot of optimizations compared to no-pragma (for example it doesn't call sqrtf at all), but the last simplification is only done with the global flag set. I wonder if it is a missed optimization or if there is a reason for that. edit: well with pragma fast-math it appears it is simply assuming finite math and x>0 thus skipping the call to the sqrtf as there are not going to be any errors, basically only saving a jmp. Using pragma finite-math-only and an explicit check are enough to trigger the optimization (but passing finite-math as a command line is not). Generally it seems that the various math flags behave slightly differently in the pragma/attribute.
- 5y ago
- m4r35n357 5y agoScary stuff. I really hope the default is -fno-fast-math!!!
- ISL 5y agoFascinating to learn that something labelled 'unsafe' stands a reasonable chance of making a superior mathematical/arithmetic choice (even if it doesn't match on exactly to the unoptimized FP result). If you're taking the ratio of sine and cosine, I'll bet that most of the time you're better off with tangent....
- zarazas 5y ago1+1 is 2. Quick math
- tcpekin 5y agoWhy is this compiler optimization beneficial with -ffinite-math-only and -fno-signed-integers? From if (x > y) { do_something(); } else { do_something_else(); } to the form if (x <= y) { do_something_else(); } else { do_something(); } What happens when x or y are NaN?
- zro 5y ago> NaN is unordered: it is not equal to, greater than, or less than anything, including itself. x == x is false if the value of x is NaN [0] My read of this is that comparisons involving NaN on either side always evaluate to false. In the first one if X or Y is NaN then you'll get do_something_else, and in the second one you'll get do_something. As far as why one order would be more optimal than the other, I'm not sure. Maybe something to do with branch prediction? [0] https://www.gnu.org/software/libc/manual/html_node/Infinity-and-NaN.html https://www.gnu.org/software/libc/manual/html_node/Infinity-...
- progbits 5y agoIt is always false. In the first block do_something_else() gets executed when either is NaN, in the second it is do_something().
- SeanLuke 5y ago* x+0.0 cannot be optimized to x because that is not true when x is -0.0 Wait, what? What is x + -0.0 then? What are the special cases? The only case I can think of would be 0.0 + -0.0.
- foo92691 5y agoSeems like it might depend on the FPU rounding mode? https://en.wikipedia.org/wiki/Signed_zero#Arithmetic https://en.wikipedia.org/wiki/Signed_zero#Arithmetic