11 ms·
Beware of Fast-Math
- Sophira 1y agoPreviously discussed at https://news.ycombinator.com/item?id=29201473 https://news.ycombinator.com/item?id=29201473 (which the article itself links to at the end).
- anthk 1y agoOn Forth, there's the philosophy of the fixed point: https://www.forth.com/starting-forth/5-fixed-point-arithmetic/ https://www.forth.com/starting-forth/5-fixed-point-arithmeti... With 32 and 64 bit numbers, you can just scale decimals up. So, Torvalds was right. On dangerous contexts (uper-precise medical doses, FP has good reasons to exist, and I am not completely sure). Also, both Forth and Lisp internally suggest to use represented rationals before floating point numbers. Even toy lisps from https://t3x.org https://t3x.org have rationals too. In Scheme, you have both exact->inexact and inexact->exact which convert rationals to FP and viceversa. If you have a Linux/BSD distro, you may already have Guile installed as a dependency. Hence, run it and then: scheme@(guile-user)> (inexact->exact 2.5) $2 = 5/2 scheme@(guile-user)> (exact->inexact (/ 5 2)) $3 = 2.5 Thus, in Forth, I have a good set of q{+,-,*,/} operations for rational (custom coded, literal four lines) and they work great for a good 99% of the cases. As for irrational numbers, NASA used up 16 decimals, and the old 113/355 can be precise enough for a 99,99 of the pieces built in Earth. Maybe not for astronomical distances, but hey... In Scheme: scheme@(guile-user)> (exact->inexact (/ 355 113)) $5 = 3.1415929203539825 In Forth, you would just use : pi* 355 133 m*/ ; with a great precision for most of the objects being measured against.
- eqvinox 1y agoThose rational numbers fly out the window as soon as your math involves any kind of more complicated trigonometry, or even a square root…
- stassats 1y agoYou can turn them back into rationals, (rational (sqrt 2d0)) => 6369051672525773/4503599627370496 Or write your own operations that compute to the precision you want.
- anthk 1y agoMy post already covered inexact->exact: scheme@(guile-user)> (inexact->exact (sqrt 2.0)) $1 = 6369051672525773/4503599627370496 s9 Scheme fails on this as it's an irrational number, but the rest of Schemes such as STKlos, Guile, Mit Scheme, will do it right. With Forth (and even EForth if the images it's compiled with FP support), you are on your own to check (or rewrite) an fsqrt function with an arbitrary precision. Also, on trig, your parent commenter should check what CORDIC was. https://en.wikipedia.org/wiki/CORDIC https://en.wikipedia.org/wiki/CORDIC
- anthk 1y agoCheck CORDIC, please. https://en.wikipedia.org/wiki/CORDIC https://en.wikipedia.org/wiki/CORDIC Also, on sqrt functions, even a FP-enabled toy EForth under the Subleq VM (just as a toy, again, but it works) provides some sort of fsqrt functions: 2 f fsqrt f. 1.414 ok Under PFE Forth, something 'bigger': 40 set-precision ok 2e0 fsqrt f. 1.4142135623730951454746218587388284504414 ok EForth's FP precision it's tiny but good enough for very small microcontrollers. But it wasn't so far from the exponents the 80's engineers worked to create properly usable machinery/hardware and even software.
- dreamcompiler 1y agoIf you want high precision trig functions on rationals, nothing's stopping you from writing a Taylor series library for them. Or some other polynomial appromation or a lookup table or CORDIC.
- AlotOfReading 1y agoFloats are fixed point, just done in log space. The main change is that the designers dedicated a few bits to variable exponents, which introduces alignment and normalization steps before/after the operation. If you don't mix exponents, you can essentially treat it as identical to a lower precision fixed point system.
- anthk 1y agoNo, not even close. Scaling integers to mimic decimals under 32 and 64 bit can be much faster. And with 32 bit double numbers you can cover Plank numbers, so with 64 bit double numbers you can do any field.
- orlp 1y agoI helped design an API for "algebraic operations" in Rust: <https://github.com/rust-lang/rust/issues/136469 https://github.com/rust-lang/rust/issues/136469>, which are coming along nicely. These operations are 1. Localized, not a function-wide or program-wide flag. 2. Completely safe, -ffast-math includes assumptions such that there are no NaNs, and violating that is undefined behavior. So what do these algebraic operations do? Well, one by itself doesn't do much of anything compared to a regular operation. But a sequence of them is allowed to be transformed using optimizations which are algebraically justified, as-if all operations are done using real arithmetic.
- eqvinox 1y agoAre these calls going to clear the FTZ and DAZ flags in the MXCSR on x86? And FZ & FIZ in the FPCR on ARM?
- orlp 1y agoI don't believe so, no. Currently these operations only set the LLVM flags to allow reassociation, contraction, division replaced by reciprocal multiplication, and the assumption of no signed zeroes. This can be expanded in the future as LLVM offers more flags that fall within the scope of algebraically motivated optimizations.
- eqvinox 1y agoAh sorry I misunderstood and thought this API was for the other way around, i.e. forbidding "unsafe" operations. (I guess the question reverses to setting those flags) ('Naming: "algebraic" is not very descriptive of what this does since the operations themselves are algebraic.' :D)
- nextaccountic 1y ago> ('Naming: "algebraic" is not very descriptive of what this does since the operations themselves are algebraic.' :D) Okay, the floating point operations are literally algebraic (they form an algebra) but they don't follow some common algebraic properties like associativity. The linked tracking issue itself acknowledges that: > Naming: "algebraic" is not very descriptive of what this does since the operations themselves are algebraic. Also this comment https://github.com/rust-lang/rust/issues/136469#issuecomment-2727298355 https://github.com/rust-lang/rust/issues/136469#issuecomment... > > On that note I added an unresolved question for naming since algebraic isn't the most clear indicator of what is going on. > > I think it is fairly clear. The operations allow algebraically justified optimizations, as-if the arithmetic was real arithmetic. > > I don't think you're going to find a clearer name, but feel free to provide suggestions. One alternative one might consider is real_add, real_sub, etc. Then retorted here https://github.com/rust-lang/rust/issues/136469#issuecomment-2727298355 https://github.com/rust-lang/rust/issues/136469#issuecomment... > These names suggest that the operations are more accurate than normal, where really they are less accurate. One might misinterpret that these are infinite-precision operations (perhaps with rounding after a whole sequence of operations). > > The actual meaning isn't that these are real number operations, it's quite the opposite: they have best-effort precision with no strict guarantees. > > I find "algebraic" confusing for the same reason. > > How about approximate_add, approximate_sub? And the next comment > Saying "approximate" feels imperfect, as while these operations don't promise to produce the exact IEEE result on a per-operation basis, the overall result might well be more accurate algebraically. E.g.: > > (...) So there's a discussion going on about the naming
- eqvinox 1y agoI wish the Twitter links in this article weren't broken.
- rlpb 1y ago> I mean, the whole point of fast-math is trading off speed with correctness. If fast-math was to give always the correct results, it wouldn’t be fast-math, it would be the standard way of doing math. A similar warning applies to -O3. If an optimization in -O3 were to reliably always give better results, it wouldn't be in -O3; it'd be in -O2. So blindly compiling with -O3 also doesn't seem like a great idea.
- CamouflagedKiwi 1y agoThe optimisations in -O3 aren't supposed to give incorrect results. They're not in -O2 because they make a more aggressive space/speed tradeoff or increase compile times more significantly. In the same way, the optimisations in -O2 are not meant to be less correct than -O1, but they aren't in that group for similar reasons. -Ofast is the 'dangerous' one. (It includes -ffast-math).
- rlpb 1y ago> The optimisations in -O3 aren't supposed to give incorrect results. I didn't mean to imply that they result in incorrect results. > they make a more aggressive space/speed tradeoff... Right...so "better" becomes subjective, depends on the use case, so it doesn't make sense to choose -O3 blindly unless you understand the trade-offs and want that side of them for the particular builds you're doing. Things that everyone wants would be in -O2. That's all I'm saying.
- eqvinox 1y agoIt doesn't become subjective; things in -O3 can objectively be understood to produce equal or faster code for a higher build cost in the vast majority of cases, roughly averaged across platforms. (Without loss in correctness.) If you know your exact target and details about your input expectations, of course you can optimize further, which might involve turning off some things in -O3 (or even -O2). On a whole bunch of systems, -Os can be faster than -O3 due to I-cache size limits. But at-large, you can expect -O3 to be faster. Similar considerations apply for LTO and PGO. LTO is commonly default for release builds these days, it just costs a whole lot of compile time. PGO is done when possible (i.e. known majority inputs).
- zinekeller 1y ago(2021) Previous discussion: Beware of fast-math (Nov 12, 2021, https://news.ycombinator.com/item?id=29201473 https://news.ycombinator.com/item?id=29201473)
- Affric 1y agoFor non-associativity what is the best way to order operations? Is there an optimal order for precision whereby more similar values are added/multiplied first? EDIT: I am now reading Goldberg 1991 Double edit: Kahan Summation formula. Goldberg is always worth going back to.
- zokier 1y agoHerbie can optimize arbitrary floating point expressions for accuracy https://herbie.uwplse.org/ https://herbie.uwplse.org/
- Sharlin 1y ago> -funsafe-math-optimizations What's wrong with fun, safe math optimizations?! (:
- emn13 1y agoI get the feeling that the real problem here are the IEEE specs themselves. They include a huge bunch of restrictions that each individually aren't relevant to something like 99.9% of floating point code, and probably even in aggregate not a single one is relevant to a large majority of code segments out in the wild. That doesn't mean they're not important - but some of these features should have been locally opt-in, not opt out. And at the very least, standards need to evolve to support hardware realities of today. Not being able to auto-vectorize seems like a pretty critical bug given hardware trends that have been going on for decades now; on the other hand sacrificing platform-independent determinism isn't a trivial cost to pay either. I'm not familiar with the details of OpenCL and CUDA on this front - do they have some way to guarrantee a specific order-of-operations such that code always has a predictable result on all platforms and nevertheless parallelizes well on a GPU?
- Affric 1y agoHow does IEEE 754 prevent auto-vectorisation?
- Kubuxu 1y agoIIRC reordering additions can cause the result to change which makes auto-vectorisation tricky.
- kzrdude 1y agoIf you write a loop `for x in array { sum += x }` Then your program is a specification that you want to add the elements in exactly that order, one by one. Vectorization would change the order.
- stingraycharles 1y agoYup, because of the imprecision of floating points, cannot just assume that “(a + c) + (b + d)” is the same as “a + b + c + d”. It would be pretty ironic if at some point fixed point / bignum implementations end up being faster because of this.
- cycomanic 1y agoI think this article overstates the importance of the problems even for scientific software. In the scientific code I've written, noise processes are often orders of magnitude larger than what what is discussed here and I believe this applies to many (most?) simulations modelling the real world (i.e. Physics chemistry,..). At the same time enabling fast-math has often yielded a very significant (>10%) performance boost. I particularly find the discussion of - fassociative-math because I assume that most writers of some code that translates a mathetical formula to into simulations will not know which would be the most accurate order of operations and will simply codify their derivation of the equation to be simulated (which could have operations in any order). So if this switch changes your results it probably means that you should have a long hard look at the equations you're simulating and which ordering will give you the most correct results. That said I appreciate that the considerations might be quite different for libraries and in particular simulations for mathematics.
- londons_explore 1y agoIt would be nice if there was some syntax for "math order matters, this is the order I want it done in". Then all other math will be fast-math, except where annotated.
- hansvm 1y agoThe article mentioned that gcc and clang have such extensions. Having it in the language is nice though, and that's the approach Zig took.
- sfn42 1y agoI thought most languages have this? If you simply write a formula operations are ordered according to the language specifiction. If you want different ordering you use parentheses. Not sure how that interacts with this fast math thing, I don't use C
- kstrauser 1y agoThat’s a different kind of ordering. Imagine a function like Python’s `sum(list)`. In abstract, Python should be able to add those values in any order it wants. Maybe it could spawn a thread so that one process sums the first half in the list, another sums the second half at the same time, and then you return the sum of those intermediate values. You could imagine a clever `sum()` being many times faster, especially using SIMD instructions or a GPU or something. But alas, you can’t optimize like that with common IEEE-754 floats and expect to get the same answer out as when using the simple one-at-a-time addition. The result depends on what order you add the numbers together. Order them differently and you very well may get a different answer. That’s the kind of ordering we’re talking about here.
- quotemstr 1y agoAll I want for Christmas is a programming language that uses dependant typing to make floating point precision part of the type system. Catastrophic cancellation should be a compiler error if you assign the output to a float with better ulps than you get with worst case operands.
- thesuperbigfrog 1y agoAda might have what you want: https://www.jviotti.com/2017/12/05/an-introduction-to-adas-simple-numeric-types.html#real-types https://www.jviotti.com/2017/12/05/an-introduction-to-adas-s... http://www.ada-auth.org/standards/22rm/html/RM-3-5-7.html http://www.ada-auth.org/standards/22rm/html/RM-3-5-7.html http://www.ada-auth.org/standards/22rm/html/RM-A-5-3.html http://www.ada-auth.org/standards/22rm/html/RM-A-5-3.html Ada also has fixed point types: http://www.ada-auth.org/standards/22rm/html/RM-3-5-9.html http://www.ada-auth.org/standards/22rm/html/RM-3-5-9.html
- deleted 1y ago[deleted]
- bsenftner 1y ago[flagged]
- glkindlmann 1y agoForgive them. That kind of educational failure is common: floating point is just not covered in basic intro classes, or covered hurriedly. There are enough subtleties to how floating point code works that even when an elective class covers it, it is still peripheral to the focus of the class (e.g. intro computer graphics). So students can get by without really coming to terms with the ubiquity of rounding error, or the relationship between denormalized and normalized values. Understanding those fundamentals is a pre-req to understanding how corners are being cut with fast-math.
- bsenftner 1y agoThis is not an issue of floating point math, it's a basic understanding that nothing comes without a cost. If there is a "faster anything" switch, that is at the cost of something else, and it is our entire professional integrity to understand that relationship, and then investigate what is being compromised by that "optimization". Anything less, that's not engineering.
- tomhow 1y agoBe kind. Don't be snarky. Converse curiously; don't cross-examine. Edit out swipes. Please don't fulminate. https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html
- storus 1y agoThis problem is happening even on Apple MPS with PyTorch in deep learning, where fast math is used by default in many operations, leading to a garbage output. I hit it recently while training an autoregressive image generation model. Here is a discussion by folks that hit it as well: https://github.com/pytorch/pytorch/issues/84936 https://github.com/pytorch/pytorch/issues/84936
- JKCalhoun 1y ago> Even compiler developers can't agree. > This is perhaps the single most frequent cause of fast-math-related StackOverflow questions and GitHub bug reports The second line above should settle the first.
- layer8 1y agoThe first line points out that it doesn't, even if one thinks that it should. Also, note the "perhaps".
- teleforce 1y ago“Nothing brings fear to my heart more than a floating point number.” - Gerald Jay Sussman Is there any IEEE standards committee working on FP alternative for examples Unum and Posit [1],[2]. [1] Unum & Posit: https://posithub.org/about https://posithub.org/about [2] The End of Error: https://www.oreilly.com/library/view/the-end-of/9781482239867/ https://www.oreilly.com/library/view/the-end-of/978148223986...
- Q6T46nT668w6i3m 1y agoIs this sarcasm? If not, the proposed posit standard, IEEE P3109.
- kvemkon 1y agoI'm wondering, why there are still no announcements for hardware support of such approaches in CPUs.
- neepi 1y agoHP had proper deterministic decimal arithmetic since the 1970s.
- teleforce 1y agoAny link to that?
- vanderZwan 1y agoGustavson's last presentation starts with him holding an actual piece of hardware supporting posits, fwiw. https://m.youtube.com/watch?v=vzVlQhaAZtQ https://m.youtube.com/watch?v=vzVlQhaAZtQ
- datameta 1y agoLuckily outside of mission critical systems, like in demoscene coding, I can happily use "44/7" as a 2pi approximation (my beloved)
- razighter777 1y agoThe worst thing that strikes fear into me is seeing floating points used for real world currency. Dear god. So many things can go wrong. I always use unsigned integers counting number of cents. And if I gotta handle multiple currencies, then I'll use or make a wrapper class.
- knert 1y agoHow do you store negative numbers?
- psychoslave 1y agoMaybe as in accounting, one column for benefits, one for debts?
- MobiusHorizons 1y agoYou use a signed integer type, so you just store a negative number. You can think of fixed point as equivalent to ieee754 floats with a fixed exponent and a two’s complement mantissa instead of a sign bit.
- rcleveng 1y agoWrappers are good even when non dealing with multiple currencies since in many places some transactions are in fractions of cents, so depending on the usecase may need to push that decimal a few places out. I always have a wrapper class to put the logic of converting to whole currency units when and if needed, as well as when requirements change and now you need 4 digits past the decimal instead of 2, etc.
- simonw 1y agoI've been having an interesting challenge relating to this recently. I'm trying to calculate costs for LLM usage, but the amounts of money involved are so tiny. Gemini 1.5 Flash 8B is $0.0375 per million tokens! Should I be running my accounting system on units of 10 billionths of a dollar?
- 1y ago
- chuckadams 1y agoI haven't worked with C in nearly 20 years and even I remember warnings against -ffast-math. It really ought not to exist: it's just a super-flag for things like -funsafe-math-optizations, and the latter makes it really clear that it's, well, unsafe (or maybe it's actually funsafe!)
- smcameron 1y agoOne thing I did not see mentioned in the article, or in these comments (according to ctrl-f anyway) is the use of feenableexcept()[1] to track down the source of NaNs in your code. feenableexcept(FE_DIVBYZERO | FE_INVALID | FE_OVERFLOW); will cause your code to get a SIGFPE whenever a NaN crawls out from under a rock. Of course it doesn't work with fast-math enabled, but if you're unknowingly getting NaNs without fast-math enabled, you obviously need to fix those before even trying fast-math, and they can be hard to find, and feenableexcept() makes finding them a lot easier. [1] https://linux.die.net/man/3/feenableexcept https://linux.die.net/man/3/feenableexcept
- DavidVoid 1y agoYeah it's pretty useful to enable every once in a while just to see if anything complains. Be very careful with it in production code though [1]. If you're in a dll then changing the FPU exception flags is a big no-no (unless you're really really careful to restore them when your code goes out of scope). [1]: https://randomascii.wordpress.com/2016/09/16/everything-old-is-new-again-and-a-compiler-bug/ https://randomascii.wordpress.com/2016/09/16/everything-old-...
- jart 1y agoTrapping math is the enlightened way to do things. I wrote an example in the cosmo repo of how to use it. https://github.com/jart/cosmopolitan/blob/master/examples/trapping.c https://github.com/jart/cosmopolitan/blob/master/examples/tr...
- hyghjiyhu 1y agoOne thing I wonder is what happens if you have an inline function in a header that is compiled with fast math by one translation unit and without in another.
- jmb99 1y agoI haven’t checked, but my assumption is that the output of each compilation unit will be different. The one definition rule doesn’t apply here (there’s still one definition), and there shouldn’t be a conflict if the functions are inlined in their respective compilation units so the linker shouldn’t complain. Could be wrong but that’s my gut feeling.
- sholladay 1y agoCorrectness > performance, almost always. It’s easier to notice that you need more performance than to notice that you need more correctness. Though performance outliers can definitely be a hidden problem that will bite you. Make it work. Make it right. Make it fast.
- KingLancelot 1y ago[dead]
- cbarrick 1y agoThis page consistently crashes on Vivaldi for Android. Vivaldi 7.4.3691.52 Android 15; ASUS_AI2302 Build/AQ3A.240812.002
- hipgoat 1y ago[flagged]
- dirtyhippiefree 1y agoI’m stunned by the following admission: “If fast-math was to give always the correct results, it wouldn’t be fast-math” If it’s not always correct, whoever chooses to use it chooses to allow error… Sounds worse than worthless to me.
- mg794613 1y agoHaha, the neverending cycle. Stop trying. Let their story unfold. Let the pain commence. Wait 30 years and see them being frustrated trying to tell the next generation.
- boulos 1y agoI've also come around to --ffast-math considered harmful. It's useful though to help find optimization opportunities, but in the modern (AVX2+) world, I think the risks outweigh the benefits. I'm surprised by the take that FTZ is worse than reassociation. FTZ being environmental rather than per instruction is certainly unfortunate, but that's true of rounding modes generally in x86. And I would argue that most programs are unprepared to handle subnormals anyway. By contrast, reassociation definitely allows more optimization, but it also prohibits you from specifying the order precisely: > Allow re-association of operands in series of floating-point operations. This violates the ISO C and C++ language standard by possibly changing computation result. I haven't followed standards work in forever, but I imagine that the introduction of std::fma, gets people most of the benefit. That combined with something akin to volatile (if it actually worked) would probably be good enough for most people. Known, numerically sensitive code paths would be carefully written, while the rest of the code base can effectively be "meh, don't care".
- leephillips 1y agoThis part was fascinating: “The problem is how FTZ actually implemented on most hardware: it is not set per-instruction, but instead controlled by the floating point environment: more specifically, it is controlled by the floating point control register, which on most systems is set at the thread level: enabling FTZ will affect all other operations in the same thread. “GCC with -funsafe-math-optimizations enables FTZ (and its close relation, denormals-are-zero, or DAZ), even when building shared libraries. That means simply loading a shared library can change the results in completely unrelated code, which is a fun debugging experience.”