5 ms·
"-fno-math-errno" and "-fno-signed-zeros" can be turned on without any problems. I got a four times speedup on <cmath> functions with no loss in accuracy. See
by optimalsolver 5y ago
"-fno-math-errno" and "-fno-signed-zeros" can be turned on without any problems.
I got a four times speedup on <cmath> functions with no loss in accuracy.
See also "Why Standard C++ Math Functions Are Slow":
https://medium.com/@ryan.burn/why-standard-c-math-functions-are-slow-d10d02554e33 https://medium.com/@ryan.burn/why-standard-c-math-functions-...
- iamcreasy 5y agoSorry, if it's a basic question - but does recompiling C/C++ code(with/without flags) produce more efficient code most of the time? For example - let say I am using a binary that was compiled on a processor that didn't have support for SIMD. Assuming the program is capable of taking advantage of SIMD instructions, and also assuming my processor support SIMD - would it make sense to recompile the C/C++ code on my system again hoping the newer binary would run faster?
- bee_rider 5y agogcc has the ability to target different architectures (look up the -march and -mtune flags for example). Linux distributions are typically set to be compatible with a pretty wide range of devices, so they often don't take advantage of recent instructions. Compiling a big program can be a bit of a pain, though, so it is probably only worthwhile if you have a program that you use very frequently. Also compilers aren't magic, the bottleneck in the program you want to run could be various things: CPU stuff, memory bandwidth, weird memory access patterns, disk access, network access, etc. The compiler mostly just helps with the first one. Also, note that some libraries, like Intel's MKL, are able to check what processor you are using and just dispatch the appropriate code (your mileage may vary, they sometimes don't keep up with changes in AMD processors, causing great annoyance).
- optimalsolver 5y agoThe more CPU-bound your program, the more benefit you'll see from the optimization flags. If your program is constantly waiting around for input on a network channel, then it may not help as much. My currently used optimization flags are: -O3 -fno-math-errno -fno-signed-zeros -march=native -flto Only use -march=native if the program is only intended to run on your own machine. It carries out architecture-specific optimizations that make the program non-portable. Also look into profile-guided optimization, where you compile and run your program, automatically generate a statistical report, then recompile using that generated information. It can result in some dramatic speedups. https://ddmler.github.io/compiler/2018/06/29/profile-guided-optimization.html https://ddmler.github.io/compiler/2018/06/29/profile-guided-...
- iamcreasy 5y agoThank you. I want to make sure I understood it clearly... I mostly use Java, and my impression is JIT inside JVM introduces hardware specific optimization without any user intervention. But for C/C++ if dependency is included as source - I can use compiler flag to enable platform specific optimizations. But if the dependency was included in the form of a pre-compiled binary, such as a .dll or .so, I am probably not using the the most optimally compiled version of the dependency. Am I right so far?
- KMag 5y agoYes. Careful selection of compilation flags can greatly improve performance. My employer spends many millions of dollars annually running numerical simulations using double precision floating point numbers. Some years ago when we retired the last machines that didn't support SSE2, adding a flag to allow the compiler to generate SSE2 instructions had a big time and cost savings for our simulations.
- jcranmer 5y ago> Some years ago when we retired the last machines that didn't support SSE2, adding a flag to allow the compiler to generate SSE2 instructions had a big time and cost savings for our simulations. That's kind of a special case, though. Without SSE2, you're using x87 for floating-point numbers, and even using scalar floating point on x87 is going to be a fair bit slower than using scalar floating point SSE instructions. Of course, enabling SSE also allows you to vectorize floating point at all, but you'll still be seeing improvements just from scalar SSE instead of x87.
- mhh__ 5y ago3 thoughts: 1. Using SIMD can be a big win, so yes. 2. SIMD (vectorization) is not the only optimization your compiler can do, the compiler has a model of the processor so it can pick the right instructions and lay them out properly with as many tricks as they can describe generically. 3. Compilers have PGO. Use it (if you can). Compilers without PGO are a bit like an engine management unit with no sensors - all the gear, no idea. The compiler has to assume a hazy middle-of-the-road estimate of what branches will be exercised, whereas with PGO enabled your compiler can make the cold code smaller, and be more aggressive with hot code etc. etc.
- bee_rider 5y ago> all the gear, no idea I like this because it only makes sense in some accents. For example it wouldn't work in Boston where the r would only be pronounced on one of the words (idea).
- AstralStorm 5y agoThere are a few critical algorithms where fp error cancellation or simple ifs get optimized out if you disable signed zeros. Typically you would know which these are, they tend to appear in statistical machine learning which use sign or expect monotonicity near zero, or filters with coefficients that are near zero (and filter out NaNs explicitly).
- CamperBob2 5y agoWhat would be an example of a filter with coefficients near zero that would be adversely affected by the loss of signed-zero support? You're already in mortal peril if you're working with "coefficients near zero" because of denormals, another bad idea that should have been disabled by default and turned on only in the vanishingly-few applications that benefit from them.
- josefx 5y ago> "-fno-math-errno" ... can be turned on without any problems. There is even a good reason for that, math-errno is a posix requirement and completely optional in both C and C++ standards. If your code is intended to be portable it should avoid relying on this anti-feature anyway.