6 ms·
Subnormal floating-point numbers are expensive on Intel processors
- deleted 16d ago[deleted]
- rf15 16d ago...Is this running extra micro code to fix some hardware bug/unreliability? How can this happen? Doesn't look like a normal design decision.
- Sharlin 16d agoSubnormal numbers have a different, basically fixed-point, representation. They exist in order to bridge the large (relatively speaking; indeed "infinite" in a sense) gap between the least positive normal number, zero, and the greatest negative normal number, caused by the usual significand-exponent representation. Most "mundane" uses of floating point have no need for subnormal numbers, and numbers that underflow could just be flushed to zero. But they’re sometimes important in scientific computing to ensure sufficient smoothness around zero, avoiding precision issues.
- bee_rider 16d agoI don’t know how useful they are in scientific computing either, really. They are less precise than normalized numbers… if flushing them makes a difference I think it is a bad algorithm smell.
- adrian_b 15d agoScientific computing can be done only in 2 ways, either with subnormals or by enabling the underflow exception and writing a suitable exception handler for it. If the use of subnormals is disabled with FTZ/DAZ that is guaranteed to generate big errors and it is completely unpredictable how big the errors will be. If a computational algorithm generates underflows at some place, there is no way to modify the algorithm so that flushing-to-zero will not make any difference (i.e. no errors). What is possible, is to modify the algorithm so that underflows will never happen. This was the traditional way of writing numeric algorithms. Because on early computers underflows would crash the program, the same as overflows, one had to improve the algorithm in order to avoid both underflows and overflows. Subnormals and infinities have been introduced in the standard precisely for lazier programmers, so that they would be able to avoid the rewriting of algorithms without the risks that underflows and overflows would generate major errors. Unfortunately, it seems that for some programmers this is still not enough, because they want simultaneously to not be bothered with rewriting the algorithms and to have the program run as fast as with an optimized algorithm. For this, the solution is very simple and it is not enabling FTZ/DAZ, which unless is done for a game might cause unpredictable financial losses for an unsuspecting customer, who expects that a computer must provide correct results. The right solution is to not buy Intel CPUs or any other kind of processors whose vendor believes that the correctness of computations does not matter. It should be noted however, that the Intel server CPUs use CPU cores that are obsolete in desktop and laptop CPUs, i.e. the tested Intel CPUs use cores like those in Meteor Lake and Raptor Lake CPUs. I do not know if the more recent Intel CPU cores, from Panther Lake/Arrow Lake S/Arrow Lake H/Lunar Lake, have retained this Intel misfeature, which has characterized the Intel CPUs for much more than a decade. If someone says that they have enabled FTZ/DAZ and they did not see any significant difference in the results of a program, that is complete B*S*T, because it is impossible to test exhaustively any program that does floating-point computations and the errors are expected to happen only for certain values, which are unlikely to be encountered during testing, but you cannot predict that those values will not be encountered in production.
- nayuki 16d agoI can think of one useful property of subnormal numbers off the top of my head. If subnormal processing is enabled, then for all finite values of `a` and `b`, `a != b` if and only if `a - b != 0`. But if subnormals are flushed to zero, then two tiny normal distinct values `a` and `b` would have a subnormal difference that is flushed to zero.
- bryanlarsen 16d agoIsn't that just a scale issue that exists with or without subnormals? If a and b are closer to zero than the smallest representable number, a and b compare as the same. With subnormals your smallest possible number is smaller than without, but it's still the same issue.
- Sharlin 16d agoBecause subnormals are fixed point, ie. have a fixed exponent, the difference of any two distinct subnormal values is nonzero like with integers.
- bryanlarsen 16d agoBut that same statement applies to normal values too, right? With normal numbers you might get the oddity of a-b -> a even if b is nonzero, but you don't get the oddity of a-b -> 0 unless the same number is represented, IIUC. A and B might not be bit identical, but they represent the same number if the difference is 0.
- jwmerrill 16d agoOne nice thing that subnormals get you is the property that if x-y == 0 then x == y. If you want to guard against division by 0, and your denominator is a difference of two terms, it’s nice to be able to check equality of those terms and know that if they are not equal, then their difference will not be 0. More generally, subnormals are needed for Sterbenz Lemma to hold everywhere: https://en.wikipedia.org/wiki/Sterbenz_lemma https://en.wikipedia.org/wiki/Sterbenz_lemma
- adrian_b 15d agoOn early computers, any underflow generated an exception that would crash the program if not handled. This was very good, because underflows completely break the assumptions about floating-point arithmetic on which numeric algorithms are based, so the errors in the final results become unpredictable. Subnormal numbers have been introduced as a means to avoid handling every underflow exception, because typically the use of subnormals eliminates the errors that would otherwise be caused by underflows. The flush-to-zero and denormals-of-zero options must be strictly forbidden for any general-purpose applications. They should be permitted only in applications where there is no doubt that regardless how big the errors will be they will not have any really harmful effect, which is true for games and perhaps for AI, but for little else. This is another great misfeature promoted by Intel, in order to win meaningless benchmarks. It would have been much better if these standard-breaking features would not have existed, because they are much more often used when they should not be used, than when they are harmless.
- aardvark179 16d agoI don’t know if any bugs contribute to this but this in the intel case but it has been very common historically for subnormal performance to be lower on many processors, and things like the Alpha required you to handle them in software if the COU fired a trap. Have a look at https://en.wikipedia.org/wiki/Subnormal_number https://en.wikipedia.org/wiki/Subnormal_number for some context.
- duped 16d agoIt's to satisfy IEEE 754 and it's been this way for decades.
- pohl 16d agoDoes that mean that the ARM processors in the writeup are not satisfying IEEE 754?
- tasty_freeze 16d agoIt probably means Apple spent the silicon to handle subnormals at full speed in hardware, rather than triggering a slow microcode handler for such numbers.
- adrian_b 15d agoAll the tested CPUs implement the standard and they implement it in the right way, except for Intel, who has chosen to save some bucks even if this decision might cause unpredictable financial losses for naive customers, who might choose to use the dangerous FTZ/DAZ options to avoid the Intel slowdown, which in turn may cause unpredictable computation errors, with even more unpredictable consequences. Some poster has linked a Mastodon thread, where Fabian Giesen explains that handling in hardware the subnormals is cheap in floating-point adders and in fused-multiply-add (FMA) execution units. Many processors do the multiplications in the FMA execution units, so there is no penalty for them to do the subnormal handling in the right way. On the other hand, some CPUs, including the Intel big cores, have some floating-point multipliers that are separate from the FMA units. The reason is that those separate multipliers can have lower latencies, typically by 1 or 2 clock cycles, which may help those CPUs to win some benchmarks, especially when running unoptimized legacy programs (in optimized programs, most multiplications are combined with additions into FMA operations). The separate multipliers are simplified in comparison with those included in the FMA units, and handling subnormals in them would be expensive, because then they would become so complex that there would be no advantage for them to be separate multipliers. Which is why Intel does not handle subnormal multiplication in hardware, but a microprogram is invoked for this.
- david-gpu 16d ago
- juancn 16d agoApparently it only happens on P-cores, recent E-cores have a fast path for subnormals.
- rustybolt 16d agoThat seems weird. They did throw extra hardware at it to speed it up for the efficient cores, but not for the performance cores?
- deleted 16d ago[deleted]
- cwzwarich 16d agoProbably different teams. Plus you can not underestimate the role that momentum plays in semiconductor engineering teams. If some respected person determined that subnormals are either Hard (tm) or not a Real Problem (tm), it will take a long time to correct this mistaken belief. I believe there have been some recent academic papers on FP implementation from the Intel E-Core team, which is a sign that they are a bit more with the times.
- gpderetta 16d agolarger SIMD ALUs, which also have been around for longer?
- Sesse__ 16d agoThey are designed by different teams and have different ways of doing math. You could just as easily say “why are the efficiency cores wasting an extra cycle for each multiplication doing fixups of an uncommon case”. Fabian Giesen explains in more detail here: https://mastodon.gamedev.place/@rygorous/117277063419144390 https://mastodon.gamedev.place/@rygorous/117277063419144390
- khuey 16d agoIf you don't _need_ subnormals MXCSR.DAZ/FTZ (which you can get gcc to set via -mdaz-ftz) will let you ignore all of this.
- bee_rider 16d agoIIRC intel’s compilers enable FTZ/DAZ, at least at higher optimization levels.
- account42 16d agoGCC does with the infamous -ffast-math as well.
- gpderetta 16d agoIIRC that has been since split out of the flag and need to be asked for separately at link time.
- jcranmer 16d agoIf you use -ffast-math when linking an executable but not a shared library, both gcc and clang will link in crtfastmath.o which has the bit of code to set the DAZ/FTZ flags.
- deleted 16d ago[deleted]
- gpderetta 16d agoright, but the issue was that if an application was linked with a library that happened to link with a fast-math shared library, it would unknowingly bring along the crt code. Now the only way to get it is to use the fast-math flag when linking the final binary, which at least is an explicit request.
- adrian_b 15d ago
- gwbas1c 16d agoI'm still trying to understand what a subnormal number is; IE, I'm looking for the TLDR so I know just enough to know if I'm using them and need to learn more. Unfortunately, the Wikipedia article, while probably being accurate, doesn't give a clear and concise answer. IE, is 0.0001 a subnormal? Or is it 0.000000000000000000001?
- ant6n 16d agoUsually IEEE floats have an implied 1 in the front. So for the standard represented numbers, there's some minimum number 1.bbbbbb.. * 2^-N. This allows 1bit more precision than is actually stored. between any two numbers, there's basically the same epsilon difference, but from the smallest number to zero it's bigger. A subnormal number breaks that convention, it just becomes 0.bbbbb... * 2^-N. As the numbers get smaller, the relative difference between the numbers gets larger. That also means their precision is smaller than the normal floats.
- account42 16d agoFloating point numbers are usually interpreted as sign * 1.mantissa * 2 ^ exponent where sign, mantissa and exponent are fixed bit width integers. The 1. before the number is normally implicit because it would be a waste of a bit to encode it when you could just use a diferent exponent to represent such a number. However with this simple scheme the number zero and a relatively large gap around it cannot be represented (relatively large to the gap between the smallest and next smalles number that can be represented). So there is a special case where for the smallest encodeable exponent the mantissa must also specify that 1. or 0. prefix. Because its a special case it needs special handling that clever silicon engineers might think is unimportant enough to handle in microcode instead of dedicated silicon. x86 has a mode to assume that all such small numbers are actually equal to zero which can then be handle without microcode fallback. Technically its even a bit more complicated because x86 has two different float implementations and for at least SSE floats you can control the denormals-are-zero and flush-(denormals)-to-zero-(when writing) modes independently. GCC -ffast-math actual enables that mode for the entire main thread. AFAIK ARM NEON always works in that mode so the Gravion and Apple benchmarks might be unfair here undless you compare with DAZ and FTZ enabled on Intel. No idea if the AMD benchmarks might have used different modes. Because the flags are global per thread you can easily have unrelated loaded libraries messing the benchmark up.
- pixelpoet 16d agoThis has been the case since a zillion years, since the Core 2 Duo days at minimum.
- account42 16d agoThe interesting part is that this seems to be Intel-specific.
- pixelpoet 16d agoThat's what I mean though- in the Core 2 Duo vs K8 / Athlon days, Intel was much slower for denormals and subnormals than AMD. I'm not sure why but I thought this was common knowledge.
- augment_me 16d agoThis is true on most hardware. Running ftz on H100s or B200 gives a free 10-20% boost for GEMM-epilogue workload
- adrian_b 15d agoThis is true only on most hardware that does not care about computational errors, i.e. which is intended mostly for games or for AI.
- mzs 16d agoAre the results compared across architectures?
- noselasd 16d agoThe article covers 5 CPUs.
- mzs 16d agoI mean the values of the computation not the runtime. I don’t know enough about ARM to say if doubles simply punt denormals to 0 for example.
- gpderetta 16d agoIEEE 754 defines bit exact results for a lot of FP operations, including denormals.
- applfanboysbgon 16d agoAnd yet floating point math in general is non-deterministic across different CPUs. IEEE 754 was not good enough, so it's a valid question.
- adrian_b 15d agoIt is non-deterministic mainly due to compilers that generate different code or when there are concurrent computations that are executed in an unpredictable order. A single sequence of floating-point operations executed on an IEEE 754 compliant CPU has always been perfectly deterministic since the first version of the standard. Who wants a deterministic computation must use a single-threaded program or a multi-threaded program where the order of execution is deterministic, and one must be careful so that changes in compiler versions or compilation options will not affect the type and order of the operations that are executed.
- cesaref 16d agoIt used to be quite normal to add a low level random signal to inputs when writing DSP code so as to avoid dropping into subnormal territory. Careful analysis of the algorithm would identify any points where this was also necessary (e.g. feedback paths when running delays). Obviously those lucky/unlucky enough to be writing 56k fixed precision code wouldn't have this concern, but other ones instead :) I think flush to zero is probably the preferred strategy these days.
- QuadrupleA 16d ago[dead]
- MotoriX 16d ago[flagged]