14 ms·
Someone’s Been Messing with My Subnormals
- puffoflogic 4y agoDynamic linking is the root of all kinds of evil, enough said.
- woodruffw 4y agoThe content of this post has nothing to do with the specifics of dynamic linking: it would be just as true if the wheels in question had static binaries instead.
- benreesman 4y agoEh, somewhere in the middle. Someone else put ‘-ffast-math’ in a compile line and it poisons FP math far away with no recompile? I believe it’s a necessary price in this case, but it does highlight how suboptimal it is to pay the price in other cases.
- woodruffw 4y agoIt's fair to point out that shared objects surface the problem here, but I don't know if I would lay the blame with them: the underlying problem is that a FPU control register isn't (and can't be) meaningfully isolated. Python needs to use shared objects for loadable extensions, but the contaminating code might be statically linked into that shared object. (I don't say this because I want to excuse dynamic linking, which I also generally dislike! Only that I think the problem is somewhere else in this particular case.)
- jeroenhd 4y agoWhat is the alternative here? To provide a python.so file with all possible binary Python packages statically linked into it? You'd need to update it every hour to include all the bugfixes in every native library yanked in! To recompile Python itself every time you install a package? Even with a compiler cache you'd have the Gentoo experience of waiting for ages every time you try to use the package manager. Dynamic linking solves a real problem, especially in this space. It comes with new problems of its own but so does the alternative.
- deleted 4y ago[deleted]
- deleted 4y ago[deleted]
- benreesman 4y agoAs a default (particularly an effectively mandatory default, looking at you glibc) it is indeed insane. But for something like a Python extension it’s what we’ve got. Which has the ancillary benefit of surfacing stuff like this.
- Tyr42 4y agoOh man, great job digging through all that. This is exactly the kind of content I want to see. Don't you love your fun safe math?
- black_knight 4y agoI ran Gentoo back in the good old days. The biggest draw was that after about a week of compiling my system ran a lot faster because of all the compiler optimisations one could enable because it only had to work on your CPU. I might be misremembering, but I think fastmath was one of the flags explicitly warned against in the Gentoo manual.
- bombcar 4y agoIt was and people would still use it because “hey it says fast”. The CPU flags was less interesting to me compared to being able to disable features like X.
- p_l 4y agoThere was a big warning that it might produce broken system, iirc
- jeffbee 4y agoChromeOS is sort of the successor to Gentoo. The images are built with profile-guided, link-time, and post-link optimization, and they are targeted to the specific CPU in a given Chromebook. Every other Linux leaves a large amount of performance on the table by targeting a common denominator CPU that's 20 years old and not having PGO.
- TazeTSchnitzel 4y agoApple avoid this problem with their OS by having a separate architecture slice for modern x64 (Haswell+).
- account42 4y agoglibc also supports loading different libraries based on CPU capabilities and some binary distros make use of that. Still, there are always (admittedly diminishing) returns if you can target the exact CPU you have - pretty sure no binary distro ever had variants for AMD's TBM instructions [0] but my Gentoo install made use of them (which was made absolutely clear when trying to run any of that on a Zen 2 machine lol). [0] https://en.wikipedia.org/wiki/X86_Bit_manipulation_instruction_set#TBM_(Trailing_Bit_Manipulation) https://en.wikipedia.org/wiki/X86_Bit_manipulation_instructi...
- jesse__ 4y ago10/10 yak shave. Would certainly read again
- hoseja 4y agoDefinitely a very bald yak.
- benreesman 4y agoThat’s…terrifying. This is a fantastic find: big, big respect to @moyix, this is going to save people’s ass.
- bee_rider 4y agoA decorator is a nice idea for this. I was going to suggest another package that just resets the MXCSR when imported, but I guess... hypothetically... some function might actually want the FTZ behavior.
- jcranmer 4y agoIf you want that behavior, you should explicitly enable it/disable it at the borders of the region where you want that behavior, rather than screwing over everybody for your own benefit.
- jcranmer 4y agoThe problem here is that enabling FTZ/DAZ flags involves modifying global (technically thread-local) state that is relatively expensive to do. Ideally, you'd want to twiddle these flags only for code that wants to work in this mode, but given the relative expense of this operation, it's not entirely practicable to auto-add twiddling to every function call, and doing it manually is somewhat challenging because compilers tend to support accessing the floating-point status rather poorly. Also, FTZ/DAZ aren't IEEE 754, so there's no portable function for twiddling these bits as there is for other rounding mode or exception controls. I will note that icc's -fp-model=fast and MSVC's /fp:fast correctly do not link code with crtfastmath. As a side note, this kind of thing is why I think a good title for a fast-math would be "Fast math, or how I learned to start worrying and hate floating point."
- deleted 4y ago[deleted]
- titzer 4y agoI don't think flipping these flags is expensive. Can you provide a source for that? AFAICT modern microarchitectures are going to register-rename that into the u-ops issued to the functional units, rather than flush the entire ROB.
- jcranmer 4y agohttps://www.agner.org/optimize/instruction_tables.pdf https://www.agner.org/optimize/instruction_tables.pdf, search for MXCSR (LDMXCSR and STMXCSR instructions). Keep in mind that twiddling these flags is going to require saving the MXCSR register to memory, or'ing or and'ing bits in memory, and then reading that memory back into MXCSR. And both saving and reading the MXCSR requires stalls, because floating point operations both read and write that register. So you require, minimum, 4 L1 cache hits and 2 partial pipeline flushes to twiddle a MXCSR bit. (As far as I'm aware, modern microarchitectures generally don't register-rename the floating-point status register.)
- ack_complete 4y ago
- garaetjjte 4y ago> it turns out that when you use -Ofast, -fno-fast-math does not, in fact, disable fast math. lol. lmao. What about -fno-unsafe-math-optimizations?
- moyix 4y agoNope, it still links in crtfastmath: $ gcc -Ofast -fno-unsafe-math-optimizations -fpic -shared foo.c -o foo.so $ objdump -j .text --disassemble=set_fast_math foo.so foo.so: file format elf64-x86-64 Disassembly of section .text: 0000000000001040 <set_fast_math>: 1040: f3 0f 1e fa endbr64 1044: 0f ae 5c 24 fc stmxcsr -0x4(%rsp) 1049: 81 4c 24 fc 40 80 00 orl $0x8040,-0x4(%rsp) 1050: 00 1051: 0f ae 54 24 fc ldmxcsr -0x4(%rsp) 1056: c3 retq
- Night_Thastus 4y agoOuch. Two flags that should reasonably stop this, and neither do. This feels a bit like the time I was told "No, -wAll does not in fact add all warnings".
- speeder 4y agoWait, it doesn't? O.o
- moyix 4y agoNope. clang has "-Weverything", and gcc has "-Wextra", both of which go beyond "-Wall". https://stackoverflow.com/questions/11714827/how-can-i-turn-on-literally-all-of-gccs-warnings https://stackoverflow.com/questions/11714827/how-can-i-turn-...
- account42 4y ago-Wextra is also in Clang and is definitely not everything in either compiler. -Weverything is everything but that is only really useful to discover new warning flags (in combination with lots of -Wno...) for which you are better off reading GCC/Clang's release notes and/or documentation.
- nsajko 4y ago-Ofast isn't a good name for the option, but in GCC's defense the manual is pretty clear about all this, and there's no excuse for blindly turning on compiler options - they literally change the semantics of your code.
- actually_a_dog 4y agoI agree. I think --ffast-math should actually be called --finexact-math. One would also hope that explicitly disabling an option on the command line would, you know, explicitly disable the option, but maybe that's too much to ask.
- mbauman 4y agoI don't think it should exist at all. It's such a crazy grab bag of code changes disguised as "optimizations" that it's completely impossible to reason about, even for folks that "don't care" about the exact floating point arithmetic. It has global effects like those in TFA, and even locally you no longer know if a line or two of arithmetic will become more precise (e.g., by using higher precision intermediate results), less precise, or become complete gibberish (e.g., because it thinks it can prove you're now dividing by zero and thus can just return whatever it wants).
- trelane 4y ago-fyolo-math? -fgoodenough-math? -fbroken-but-fast-math
- mbauman 4y agoI wholeheartedly disagree. -Ofast Disregard strict standards compliance. ... There's strict standards compliance and then there's the crazy grab bag of code changes that is `-ffast-math`. Further, I'd say gevent can defensibly say that -ffast-math is okay for them given what the manual says: -ffast-math ... it can result in incorrect output for programs that depend on an exact implementation of IEEE or ISO rules/specifications for math functions. It may, however, yield faster code for programs that do not require the guarantees of these specifications. This is 100% on the compiler people. For the option name, the documentation, and the behavior. https://gcc.gnu.org/onlinedocs/gcc-12.2.0/gcc/Optimize-Options.html#Optimize-Options https://gcc.gnu.org/onlinedocs/gcc-12.2.0/gcc/Optimize-Optio...
- ChrisRackauckas 4y agoThe Julia package ecosystem has a lot of safeguards against silent incorrect behavior like this. For example, if you try to add a package binary build which would use fast math flags, it will throw an error and tell you to repent: https://github.com/JuliaPackaging/BinaryBuilderBase.jl/blob/186b1350fa925ff9722d320b4c8549ec6a5735db/src/Runner.jl#L289 https://github.com/JuliaPackaging/BinaryBuilderBase.jl/blob/... In user codes you can do `@fastmath`, but it's at the semantic level so it will change `sin` to `sin_fast` but not recurse down into other people's functions, because at that point you're just asking for trouble. There's also calls to rename it `@unsafemath` in Julia, just to make it explicit. In summary, "Fastmath" is overused and many times people actually want other optimizations (automatic FMA), and people really need to stop throwing global changes around willy-nilly, and programming languages need to force people to avoid such global issues both semantically and within its package ecosystems norms.
- aidenn0 4y agoAutomatic FMA can change the result of operations, so it makes (some) sense to be bundled in with fastmath.
- eigenspace 4y agoThis isn't really as valid a comparison as you might think it is. The results of operations varying is not the problem with 'fast-math', the problem is that can negatively impact accuracy in catastrophic ways (among other things). Sure, automatic FMA can change the result, but to my knowledge it always gives a more accurate result, not a less accurate one, and the way in which the results may differ is bounded.
- adgjlsfhk1 4y agofma can give less accurate results, and can give very large differences. for example 1.5floatmax(Float64)-floatmax(Float64) is Inf, while the fma version will give .5floatmax(Float64)
- ChrisRackauckas 4y ago
- stabbles 4y agoSee also https://simonbyrne.github.io/notes/fastmath/ https://simonbyrne.github.io/notes/fastmath/ for a similar story in julia, where ffast-math is now banned for C/C++/Fortran dependencies
- mananaysiempre 4y agoFollowing the article’s links, I fail to find an actual example of anything failing to converge in flush-subnormals mode. I mean, I’m sure one could be squeezed out, but the justification given amounts to “Sterbenz’s lemma [the one that rephrases “catastrophic cancellation” as “exact differences”] fails, maybe something somewhere also will”. And my (shallow but not nonexistent) experience with numerical analysis is that proofs lump subnormals with underflow, and most of them don’t survive even intermediate underflows. (AFAIU the original Intel justification for pushing subnormals into 754 was gradual underflow, i.e. to give people at least something to look at for debugging when they’ve ran out of precision.) So, yes, it’s not exactly polite to fiddle with floating-point flag bits that are not yours, and it’s better that this not happen for reproducibility if nothing else, but I doubt it actually breaks any interesting numerics.
- moyix 4y agoThe gevent issue has an example: https://github.com/gevent/gevent/pull/1820 https://github.com/gevent/gevent/pull/1820 I haven't examined the code of scipy.stats.skellam.sf so I can't say for sure that it's not converging, but it's clearly some kind of pathological behavior.
- mananaysiempre 4y agoSo somebody tried to calculate, for integer arguments from 0 to 99 inclusive, the CDF of the difference of two Poisson variables with means 4e-6 and 1e-6? I... don’t know if it is at all reasonable to expect an answer to that question. As in, genuinely don’t know—obviously it’s an utterly rotten thing to compute, but at the same time maybe somebody got really interested in that and figured out a way to make it work. Anyhow, my spelunking was cut off by sleep, so the best I can tell that would end up in the CDFLIB[1] routine CUMCHN with X = 8e-6, PNONC = 2e-6, DF from 0 to 99. The insides don’t really look like the kind of magic that is held up by Sterbenz’s lemma and strategically arranged to take advantage of gradual underflow, so at first glance I wouldn’t trust anything subnormal-dependent that it would compute, but maybe it still is? Sleep. [1] https://people.sc.fsu.edu/~jburkardt/f_src/cdflib/cdflib.f90 https://people.sc.fsu.edu/~jburkardt/f_src/cdflib/cdflib.f90
- cesarb 4y agoAt a previous company I worked at, we had an issue with our software (Windows-based, written in a proprietary language) randomly crashing. After some debugging, we found that this happened whenever the user made some specific actions, but only if, in that session, the user had previously printed something or opened a file picker. The culprit was either a printer driver or a shell extension which, when loaded, changed the floating point control word to trap. That happened whenever the culprit DLL had been compiled by a specific compiler, which had the offending code in the startup routine it linked into every DLL it produced. Our solution was the inverse of the one presented in this article: instead of wrapping our routines to temporarily set the floating point control word to sane values, we wrapped the calls to either printing or the file picker, and reset the floating point control word to its previous (and sane) value after these calls.
- klysm 4y agoThat is one hell of a war story - I didn't realize that kind of failure was even possible, but it is truly terrifying.
- pavlov 4y agoDirect3D used to flip the x87 FPU to single precision mode by default. This produced some amazing bugs when your other C libraries reasonably assumed that a double would be at least 64 bits. (The FPU mode settings affected the thread that called Direct3D, and most programs used to be single-threaded.) It seems they changed this behavior in Direct3D 10: https://microsoft.public.win32.programmer.directx.graphics.narkive.com/WBTT4fr3/d3dcreate-fpu-preserve-effects https://microsoft.public.win32.programmer.directx.graphics.n...
- speeder 4y agoI stumbled into this bug in a rather spetacular manner. I was making a game using D3D, Lua and Chipmunk physics, and some of the behaviour of the game was being odd. So I started to try printing random stuff with Lua, eventually I just tried: print(5+5), and to my surprise my console outputted "11". I went into Lua's irc channel to talk about this, and everyone said I was nuts, that the number was too small to trigger precision issues, that I was a troll and so on. After a lot of searching I found out about this D3D bug, so I switched the game to use OpenGL instead there it was, 5+5 = 10 again! Now why fiddling with the FPU could make 5+5 become 11, I have no idea.
- TazeTSchnitzel 4y agoGlobal state is the root of so many evils! FPU rounding mode, FPU flush-to-zero mode, C locale, errno, and probably some other things should all be eliminated. The functionality should still exist but not as global flags.
- twic 4y agoI wonder if it would make sense for the processor to scope these settings to stack frames, and clear them when the appropriate ret happens. Somewhere on this page there's a discussion of dirty upper half states in Intel vector instructions, and something similar might help there.
- leni536 4y agoAt least many of those are thread-local. But not C locale, it is truly horrible.
- account42 4y agoOTOH MXCSR being threadlocal means that automatically linking crtfastmath.o makes even less sense as it only affects the inital thread that loaded the object so you still need to set FTY/DAZ manually if you really want them. And implicit locales should just die. It's sad that even newer functionality like std::format relies on them. Sadder that if it didn't then you'd probably have compiler developers pulling shit like GCC/libstdc++ does for std::from_chars [0] which defeats the entire point of that function but hey, at least they can mark that as implemented. Not like there are suitably licensed implementations of the functionality [1] available that they could use instead if they don't want to implement float parsing themselves. [0] https://github.com/gcc-mirror/gcc/blob/master/libstdc++-v3/src/c++17/floating_from_chars.cc#L329 https://github.com/gcc-mirror/gcc/blob/master/libstdc++-v3/s... [1] https://github.com/fastfloat/fast_float https://github.com/fastfloat/fast_float
- leni536 4y agoI thought that std::format avoided most of the implicit locale dependency. I know it has some for date formatting though.
- elina123 4y ago
- Const-me 4y agoThat thread-local MXCSR register is particularly entertaining in a thread pool environment, such as OpenMP. OSes carefully preserve that piece of thread state across context switches. I tend to avoid touching that value, even when it means extra instructions like roundpd for specific rounding mode, or shuffles to avoid division by 0 in the unused lanes.
- compiler-guy 4y ago-funsafe-math is neither fun nor safe.
- mrtesthah 4y agoI thought the purpose of Python was to make development simple and predictable. Needing to track down the compilation and linker flags of every single shared library reveals the fallacy of this abstraction.
- RodgerTheGreat 4y agoIf a language wishes to reap the rewards of a pre-existing ecosystem, it must pay for the warts and misfeatures of that ecosystem. Python is deeply dependent on C libraries to achieve acceptable performance, and this is the price.
- olliej 4y agoWow, I am surprised that -ffast-math triggers a mode switch in the FPU in part due to the author's library problem, but also because the documentation for clang at least[1] does not say it impacts behaviour of denormals and in fact has a separate mode switch for that, which is not explicitly called out as being implied by -ffast-math. [1] https://clang.llvm.org/docs/UsersManual.html#cmdoption-ffast-math https://clang.llvm.org/docs/UsersManual.html#cmdoption-ffast...
- raymondh 4y agoThis is a rockstar quality post. It is astonishing how much detective work was involved.
- leni536 4y agoDoes this only affect pypi, or should I now worry about shared libraries shipped with my distro as well? Debian is not crazy enough to ship shared libs compiled with -ffast-math, right? RIGHT?
- moyix 4y agoPlease don't do this to me, I don't know if I have it in me to go on ANOTHER big scrape & scan.
- JonChesterfield 4y agoIf the package build scripts from upstream have that in them, Debian packaged versions probably do too
- moyix 4y agoAll right fine, I give in. In current amd64 Debian unstable (main, contrib, and non-free), 48 of the 69,112 packages include a shared library built with -ffast-math, a total of 485 .so files. Actually it wasn't so bad :) All the current debs add up to a mere 113GB, and the .so files are just 46GB once extracted. Quick work. Here's the list: https://moyix.net/~moyix/debian_unstable_amd64_ffast_results.txt https://moyix.net/~moyix/debian_unstable_amd64_ffast_results... And here's a visualization of the reverse dependency graph: https://moyix.net/~moyix/ffast_debian.pdf https://moyix.net/~moyix/ffast_debian.pdf
- Twirrim 4y agoVarious distributions have meddled with ffast-math over the years. At one stage early on, I think Clear Linux might have used it, which was part of the reason they kept demolishing distributions on benchmarks. I don't know any that still actively use it, but holy crap folks, please don't.
- magicalhippo 4y agoDenormalized numbers is one reason why you really want to think carefully if you try to optimize code by rewriting expressions involving multiplication and division. For example, if you got "x = (a / b) * (c / d)" one might think that rewriting it as "x = (a * c) / (b * d)" will save you a division and gain you speed. It will and it might, respectively. However it will also potentially break an otherwise safe operation. If the numbers are very small, but still normal, then the product (b * d) might result in a denormalized number, and dividing by it will result in +/- infinity. However, the code might guarantee that the ratios (a / b) and (c / d) are not too small or too large, so that multiplying them is guaranteed to lead to a useful result.
- bee_rider 4y agoAnyway, since there aren't any dependencies between a, b, c, and d, I would expect the two divisions to end up basically in parallel in the pipeline. So the critical path is a division and a multiplication either way. Of course that is just a guess.
- magicalhippo 4y agoThat assumes you can do multiple divisions in parallel. Back in the good old days, a single division unit was the norm, and it still is on most microcontrollers (assuming they even have hardware floating-point division[1]). Anyone have any references on how the current state of affairs on modern AMD/Intels? [1]: ARM Cortex-M4 for example can have a hardware FPU, but where division and sqrt are optional, see https://developer.arm.com/documentation/102832/latest/ https://developer.arm.com/documentation/102832/latest/
- bee_rider 4y agoIt appears that Agner Fog's website is down at the moment, so we must conclude that the universe does not want to share this knowledge.
- Dylan16807 4y agohttps://en.wikichip.org/wiki/intel/microarchitectures/sunny_cove https://en.wikichip.org/wiki/intel/microarchitectures/sunny_... Looks like one FP divider on modern intel. Though you can pack multiple divisions into an instruction. For AMD I can find throughput numbers but not how many there are, in a brief search. I'd guess two??
- jupiterelastica 4y agoFor anyone (like me) not knowing what subnormal floats are, this StackOfverflow question and answer explain it quite nicely: https://stackoverflow.com/questions/15140847/denormalized-numbers-ieee-754-floating-point# https://stackoverflow.com/questions/15140847/denormalized-nu... Though they use the term denormalized, which AFAICT is the a synonym for subnormal. Edit: Thanks for this great blogpost :)
- mhh__ 4y agoIn D you can opt into specific float algorithms locally rather than a compiler flag. Use of fast math can really really really bite you sometimes, so just being able to opt into using fma and nothing else is awesome.
- adgjlsfhk1 4y agoYeah, Julia has this also (via a @fastmath macro). It's incredibly nice to be able to apply it safely.
- p0nce 4y agoI've measured it on real workload twice with LDC and each time "@fastMath" made things 1% slower. Of course this varies from backend to backend, but such woes would be avoided if the flag was named differently than "fast". It doesn't seem worth it to use it.
- mhh__ 4y agoI got a good speed up in [insert work project here] but that was more of an experiment since it isn't really performance constrained in the first place.
- superbatfish 4y agoIf you are willing to accept the various caveats that come with -ffast-math for your own library, then it looks like it's okay to use -ffast-math during compilation, but not during linking (because it links in crtfastmath.so, as the author points out). Your own compiled library functions will use the optimizations, but won't force the weird FPU register modes on the rest of the process.
- account42 4y agoPay close attention to your build system though. Many will pass all compiler flags to the linker driver as that is needed to make sure that they affect LTO.
- jandrese 4y ago> built with an appealing-sounding but dangerous compiler option It's -ffast-math isn't it? ... Yep. That option is a candy coated foot gun.
- dahart 4y ago> Finally, I am legally obligated to inform you that I used GNU Parallel to scan all the wheels FWIW, this is tangential to this awesome article, but if the author is here or anyone else who cares: that statement isn’t true, you are not legally obligated to mention GNU Parallel. It is nice to do though! The link even says this explicitly, and also separately mentions the citation notice is only asking for scientific paper citations, not blog citations. I love parallel, but boy has that citation notice thing caused a lot of confusion over the years.
- moyix 4y agoJust a small joke :)
- krallja 4y agoAlso, fortunately, GPLv3 is a contract, and not a law, so this is a contractual obligation, not a legal one.
- dahart 4y agoThe citation notice thing in parallel is not a contractual obligation either; the link explains that it is not required (nor allowed) by any version of the GPL.