6 ms·
Regarding autovectorization: > The other drawback of this method is that the optimizer won’t even touch anything involving floats (f32 and f64 types). It’s not
by andyferris 11mo ago
Regarding autovectorization:
> The other drawback of this method is that the optimizer won’t even touch anything involving floats (f32 and f64 types). It’s not permitted to change any observable outputs of the program, and reordering float operations may alter the result due to precision loss. (There is a way to tell the compiler not to worry about precision loss, but it’s currently nightly-only).
Ah - this makes a lot of sense. I've had zero trouble getting excellent performance out of Julia using autovectorization (from LLVM) so I was wondering why this was such a "thing" in Rust. I wonder if that nightly feature is a per-crate setting or what?
- Arch-TK 11mo agoIt's not something you seem to be able to just enable globally. From what I gather this is what is being referenced: https://doc.rust-lang.org/std/intrinsics/index.html https://doc.rust-lang.org/std/intrinsics/index.html Specifically the *_fast intrinsics.
- ladyanita22 11mo agoIs this equivalent to --ffast-math?
- Arch-TK 11mo agoFrom what I know of -ffast-math and can read from the docs for *_fast. I am not convinced that the *_fast intrinsics do _everything_ -ffast-math allows. They seem focused around algebraic equivalence (a/b is equivalent to a*(1/b) ) and assumptions of finite math. There's a few other things that -ffast-math allows like ignoring certain errors, ignoring the existence of signed zero, ignoring signalling NaN handling, ignoring SIGFPE handling, etc...
- Sharlin 11mo agoYes, because many of the traditional "fast math" assumptions are definitely not something that should be hidden behind an attractive option like that. In particular assuming the nonexistence of NaNs is essentially never anything but a ticket to the UB land.
- bobmcnamara 11mo agoWe used to tweak our scalar product simulator code to match the SIMD arithmetic order so we could hash the outputs for tests. I wonder if it could autovec the simd-ordered code.
- vlovich123 11mo agoDoes Julia ignore the problem of floating point not being associative, commutative nor distributive? The reason it’s a thing is from LLVM and I’m not sure you can “language design” your way out of this problem as it seems intrinsic to IEEE 754.
- tomsmeding 11mo agoNitpick, but IEEE float operations are commutative (when relevant and appropriate). Associative and distributive they indeed are not.
- vlovich123 11mo agoUnless I’m having a brain fart it’s not commutative or you mean something by “relevant and appropriate” that I’m not understanding. a+b+c != c+b+a That’s why you need techniques like Kahan summation.
- wtallis 11mo ago"a+b+c" doesn't describe a unique evaluation order. You need some parentheses to disambiguate which changes are due to associativity vs commutativity. a+(b+c)=(c+b)+a should be true of floating point numbers, due to commutativity. a+(b+c)=(a+b)+c may fail due to the lack of associativity.
- adastra22 11mo agoIt is not, due to precision. Consider a=1.00000, b=-0.99999, and c=0.00000582618.
- dzaima 11mo agoFor vectorizing, that quote is only true for loops with dependencies between iterations, e.g. summing a list of numbers (..that's basically the only case where this really matters). For loops without such dependencies Rust should autovectorize just fine as with any other element type.
- galangalalgol 11mo agoYou just create f32x4 types, the wide crate does this. Then it autovectorizes just fine. But it still isn't the best idea if you are comparing values. We had a defect due to this recently.
- the__alchemist 11mo agoI suspect I am misunderstanding. If you create an f32x4 type, aren't you manually vectorizing? Auto-vectoring is magic SIMD use the compiler does in some cases. (But usually doesn't...)
- galangalalgol 11mo agoYou are manually vectorizing, but it lets the optimizer know you don't care about safe rounding behavior so it ends up using the simd instructions. And this way it is portable still vs using intrinsics. Floating point addition is the only one the optimizer isn't allowed to do, so if you just need multiplication or only use integers it all autovectorizes fine. The f32xN stuff is just a way to tell it you don't care about the rounding. There are better ways to do that that could be added, like a FastF32 type, but I don't know if llvm could support that. Edit: go to godbolt and load the rust aligned sum and play around with types. If you see addps that is the packed scalar simd instruction. The more you get packed, the higher your score! You'll need to pass some extra arguments they don't list to get avx512 sized registers vs the xmm or ymm ones. And not all the instances it uses support avx512 so sometimes you have to try a couple times.
- dzaima 11mo ago
- queuebert 11mo agoDoes Rust not have the equivalent of GCC's "-ffast-math"?
- demurgos 11mo agoNo it doesn't. A global flag is a no-go as it breaks modularity. A local opt-in through dedicated types or methods is being designed but it's not stable.
- Sharlin 11mo agoNo, because as I commented in another subthread, `-ffast-math` is: 1. dangerous assumptions hidden behind a simple, attractive-looking option [1]. It should be called -fwrong-math or -fdangerous-math or something (GCC does have the funnily named switch -funsafe-math-optimizations – what could go wrong with fun, safe math optimizations?!) 2. Translation-unit scoped, which means that dependencies not consented to "fast math" can break your code (as in UB land) or make the optimizations pointless, and your code can break your dependencies' semantics too via inlining. On the other hand, a library author must think very carefully what float opts to enable in order to be compatible with client code. Deciding how the scoping of non-IEEE float math operations should work is a very nontrivial question. The scope could be a translation unit, a module, a type, a function, a block, or every individual operation, and none of those is without issues, particularly regarding questions like inlining and interprocedural and link-time-optimization, as well as ergonomics. In other ways, it's yet another function coloring problem. There are currently-unstable "algebraic_add/mul/etc" methods for floats for letting LLVM treat those particular operations as if floats were real numbers [2]. They're the first step towards safe UB-free float optimizations, but of course those names are rather awkward to use in math-heavy code, and a wrapper type overloading the normal operators would be good to have. --- [1] See, eg. https://simonbyrne.github.io/notes/fastmath/ https://simonbyrne.github.io/notes/fastmath/ [2] In terms of associativity and such, not in eg. assuming the nonexistence of NaNs, which would be very unsafe.
- queuebert 11mo agoAs a student of floating point math idiosyncrasies, I had always thought -ffast-math should be renamed -fsloppy-math.
- Sharlin 11mo agoLLVM autovectorizes many FP operations just fine, the article was a bit strange in that respect. Problem is, there are many other cases where it's unable to do so, not because it can't but because it isn't allowed.
- exDM69 11mo agoIn my experience, compiling C with -ffast-math will tremendously improve floating point autovectorization and optimizations to SIMD (C vector extensions, which are similar to Rust std::simd) code in general. This obviously has a lot of caveats, and should only be enabled on a per function or per file basis. Unfortunately Rust does not currently have options for adjusting per-function compiler optimization parameters. This is possible in some C compilers using function attributes.
- CryZe 11mo agoThere are algebraic operations available on nightly: https://doc.rust-lang.org/nightly/std/primitive.f32.html#algebraic-operators https://doc.rust-lang.org/nightly/std/primitive.f32.html#alg...
- swiftcoder 11mo ago> I wonder if that nightly feature is a per-crate setting or what Unfortunately it's a set of functions you have to use to perform arithmetic ops if you want the autovectorizer to touch them