6 ms·
It's OK to compare floating-points for equality
- hun3 6mo agoI used floating timestamps as some kind of an identity. If there is ever a conflict, I just increase it by 1 ulp until it doesn't collide with anything. Sorry.
- mizmar 6mo agoThere is another way to compare floats for rough equality that I haven't seen much explored anywhere: bit-cast to integer, strip few least significant bits and then compare for equality. This is agnostic to magnitude, unlike epsilon which has to be tuned for range of values you expect to get a meaningful result.
- twic 6mo agoThis doesn't work. For any number of significant bits, there are pairs of numbers one machine epsilon apart which will truncate to different values.
- andyjohnson0 6mo ago> strip few least significant bits I'm unconvinced. Doesnt this just replace the need to choose a suitable epsilon with the need to choose the right number of bits to strip? With the latter affording much fewer choices for degree of "roughness" than does the former.
- chaboud 6mo agoNot quite. It's basically a combined mantissa and exponent test, so it can be thought of as functionally equivalent to scaling epsilon by a power of two (the shared exponent of the nearly equal floating point values) and then using that epsilon. I think I'll just use scaled epsilon... though I've gotten lots of performance wins out of direct bitwise trickery with floats (e.g., fast rounding with mantissa normalization and casting).
- SideQuark 6mo agoCompletely worked out at least 20 years ago: https://www.lomont.org/papers/2005/CompareFloat.pdf https://www.lomont.org/papers/2005/CompareFloat.pdf
- fn-mote 6mo agoNote for the skeptic: this cites Knuth, Volume II, writes out the IEEE edge cases, and optimizes.
- mizmar 6mo agoGreat reading, thanks. Describes how to handle ±0, works with difference to avoid truncation errors. First half of the paper is arriving at this correct snippet, second part of the paper is about optimizing it. bool DawsonCompare(float af, float bf, int maxDiff) { int ai = *reinterpret_cast<int*>(&af); int bi = *reinterpret_cast<int*>(&bf); if (ai < 0) ai = 0x80000000 - ai; if (bi < 0) bi = 0x80000000 - bi; int diff = ai - bi; if (abs(diff) < maxDiff) return true; return false; }
- StilesCrisis 6mo agoRather than stripping bits, you can just compare if the bit-casted numbers are less than N apart (choose an appropriate N that works for your data; a good starting point is 4). This breaks down across the positive/negative boundary, but honestly, that's probably a good property. -0.00001 is not all that similar to +0.00001 despite being close on the number line. It also requires that the inputs are finite (no INF/NAN), unless you are okay saying that FLT_MAX is roughly equal to infinity.
- ethan_smith 6mo agoThis is essentially ULP (units in the last place) comparison, and it's a solid approach. One gotcha: IEEE 754 floats have separate representations for +0 and -0, so values straddling zero (like 1e-45 and -1e-45) will look maximally far apart as integers even though they're nearly equal. You need to handle the sign bit specially.
- kbolino 6mo agoThere's another gotcha. Consider positive, normal x and y where ulp(y) != ulp(x). Bitwise comparison, regardless of tolerance, will consider x to be far from y, even though they might be adjacent numbers, e.g. if y = x+ulp(x) but y is a power of 2.
- ack_complete 6mo agoThis case actually works because for finite numbers of a given sign, the integer bit representations are monotonic with the value due to the placement of the exponent and mantissa fields and the implicit mantissa bit. For instance, 1.0 in IEEE float is 0x3F800000, and the next immediate representable value below it 1.0-e is 0x3F7FFFFF. Signed zero and the sign-magnitude representation is more of an issue, but can be resolved by XORing the sign bit into the mantissa and exponent fields, flipping the negative range. This places -0 adjacent to 0 which is typically enough, and can be fixed up for minimal additional cost (another subtract).
- kbolino 6mo agoI interpreted OP's "bit-cast to integer, strip few least significant bits and then compare for equality" message as suggesting this kind of comparison (Go): func equiv(x, y float32, ignoreBits int) bool { mask := uint32(0xFFFFFFFF) << ignoreBits xi, yi := math.Float32bits(x), math.Float32bits(y) return xi&mask == yi&mask } with the sensitivity controlled by ignoreBits, higher values being less sensitive. Supposing y is 1.0 and x is the predecessor of 1.0, the smallest value of ignoreBits for which equiv would return true is 24. But a worst case example is found at the very next power of 2, 2.0 (bitwise 0x40000000), whose predecessor is quite different (bitwise 0x3FFFFFFF). In this case, you'd have to set ignoreBits to 31, and thus equivalence here is no better than checking that the two numbers have the same sign.
- hansvm 6mo agoThat works well for sorting/bucketing/etc in a few places, but as a comparison it's prone to false negatives (your values are close and computed to not be close), so you're restricted to algorithms tolerant of that behavior.
- 4pkjai 6mo agoI do this to see if text in a PDF is exactly where it is in some other PDF. For my use case it works pretty well.
- jph 6mo agoI have this floating-point problem at scale and will donate $100 to the author, or to anyone here, who can improve my code the most. The Rust code in the assert_f64_eq macro is: if (a >= b && a - b < f64::EPSILON) || (a <= b && b - a < f64::EPSILON) I'm the author of the Rust assertables crate. It provides floating-point assert macros much as described in the article. https://github.com/SixArm/assertables-rust-crate/blob/main/src/assert_f64/assert_f64_eq.rs https://github.com/SixArm/assertables-rust-crate/blob/main/s... If there's a way to make it more precise and/or specific and/or faster, or create similar macros with better functionality and/or correctness, that's great. See the same directory for corresponding assert_* macros for less than, greater than, etc.
- lukax 6mo agoYou generally want both relative and absolute tolerances. Relative handles scale, absolute handles values near zero (raw EPSILON isn’t a universal threshold per IEEE 754). The usual pattern is abs(a - b) <= max(rel_tol * max(abs(a), abs(b)), abs_tol) to avoid both large-value and near-zero pitfalls.
- lukax 6mo agoSee the implementation of Python's math.isclose https://github.com/python/cpython/blob/d61fcf834d197f0113a6a507fdbecc1545d9d483/Modules/mathmodule.c#L2586 https://github.com/python/cpython/blob/d61fcf834d197f0113a6a...
- pclmulqdq 6mo agoYour assertion code here doesn't make a ton of sense. The epsilon of choice here is the distance between 1 and the next number up, and it's completely separated from the scale of the numbers in question. 1e-50 will compare equal to 2e-50, for example. I would suggest that "equals" actually is for "exactly equals" as in (a == b). In many pieces of floating point code this is the correct thing to test. Then also add a function for "within range of" so your users can specify an epsilon of interest, using the formula (abs(a - b) < eps). You may also want to support multidimensional quantities by allowing the user to specify a distance metric. You probably also want a relative version of the comparison in addition to an absolute version. Auto-computing epsilons for an equality check is really hard and depends on the usage, as well as the numerics of the code that is upstream and downstream of the comparison. I don't see how you would do it in an assertion library.
- AshamedCaptain 6mo agoOne of the goals of comparing floating points with an epsilon is precisely so that you can apply these types of accuracy increasing (or decreasing) changes to the operations, and still get similar results. Anything else is basically a nightmare to however has to maintain the code in the future. Also, good luck with e.g. checking if points are aligned to a grid or the like without introducing a concept of epsilon _somewhere_.
- demorro 6mo agoI guess I'm confused. I thought epsilon was the smallest possible value to account for accuracy drift across the range of a floating point representation, not just "1e-4". Done some reading. Thanks to the article to waking me up to this fact at least. I didn't realize that the epsilon provided by languages tends to be the one that only works around 1.0, and if you want to use episilons globally (which the article would say is generally a bad idea) you need to be more dynamic as your ranges, and potential errors, increase.
- rpdillon 6mo agoYeah, I'm not sure how widespread the knowledge is that floating point trades precision for magnitude. Its obvious if you know the implementation, but I'm not sure most folks do.
- ryandrake 6mo agoI remember having convincing a few coworkers that the number of distinct floating point values between 0.0 and 1.0 is the same as the number of values between 1.0 and infinity. They must not be teaching this properly anymore. Are there no longer courses that explain the basics of floating point representation? I was arguing that we could squeeze a tiny bit more precision out of our angle types by storing angles in radians (range: -π to π) instead of degrees (range: -180 to 180) because when storing as degrees, we were wasting a ton of floating point precision on angles between -1° and 1°.
- valicord 6mo agoWait this doesn't make sense. Yes you'd get smaller absolute error in radians, but it doesn't really help because it's different units. Relative error is the same in degrees and radians, that's the whole point of exponential representation. All you're doing is adding a fixed offset to the exponent, but it doesn't give you any more precision when converting to radians
- adrian_b 6mo ago
- vouwfietsman 6mo agoThis explanation is relatively reductive when it comes to its criticism of computational geometry. The thing with computational geometry is, that its usually someone else's geometry, i.e you have no control over its quality or intention. In other words, whether two points or planes or lines actually align or align within 1e-4 is no longer really mathematically interesting because its all about the intention of the user: does the user think these planes overlap?. This is why most geometry kernels (see open cascade) sport things like "fuzzy boolean operations" [0]) that lean into epsilons. These epsilons mask the error-prone supply chain of these meshes that arrive in your program by allowing some tolerance. Finally, the remark "There are many ways of solving this problem" is also overly reductive, everyone reading here should really understand that this is a topic that is being actively researched right now in 2026, hence there are currently no blessed solutions to this problem, otherwise this research would not be needed. Even more so, to some extent this problem is fundamentally unsolvable depending on what you mean by "solvable", because your input is inexact not all geometrical operations are topologically valid, hence an "exact" or let alone "correct along some dimension" result cannot be achieved for all (combination of) inputs. [0] https://dev.opencascade.org/content/fuzzy-boolean-operations https://dev.opencascade.org/content/fuzzy-boolean-operations
- throwup238 6mo ago> This is why most geometry kernels (see open cascade) sport things like "fuzzy boolean operations" [0]) that lean into epsilons. These epsilons mask the error-prone supply chain of these meshes that arrive in your program by allowing some tolerance. They don’t just lean into epsilons, the session context tolerance is used for almost every single point classification operation in geometric kernels and many primitives carry their own accumulating error component for downstream math. Even then the current state of the art (in production kernels) is tolerance expansion where the kernel goes through up to 7 expansion steps retrying point classification until it just gives up. Those edge cases were some of the hardest parts of working on a kernel. This is a fundamentally unsolvable problem with floating point math (I worked on both Parasolid and ACIS in the 2000s). Even the ray-box intersection example TFA gives is a long standing thorn - raytracing is one of the last fallbacks for nasty point classification problems.
- amelius 6mo agoThink about this. It's silly to use floating point numbers to represent geometry, because it gives coordinates closer to the origin more precision and in most cases the origin is just an arbitrary point.
- anonymars 6mo agoRandom aside but as I recall I think this is what made Kerbal Space Program so difficult. Very large distances and changing origins as you'd go to separate bodies, and I think the latter was basically because of this aspect of floating point. And because of the mismanagement of KSP2 they had to relearn these difficulties, because they didn't really have the experienced people work with the new developers. I only played it rather than modded it, so happy to be corrected or further enlightened, but seems like an interesting problem to have to solve. Edit: sure enough, it was actually discussed here: https://news.ycombinator.com/item?id=26938812 https://news.ycombinator.com/item?id=26938812
- adgjlsfhk1 6mo agoWhat KSP really should have done is just done their orbital math separately from their force propagation. If they had made a virtual node for each craft's center of mass, they could have made it so that the COM position was just never affected by intra-body forces and done the orbital math in super high (Float128?) precision.
- rpdillon 6mo agoFor all the players of the original Morrowind out there, you'll notice that your character movement gets extremely janky when you're well outside of Vvardenfell because the game was never designed to go that far from the origin. OpenMW fixes this (as do patches to the original Morrowind, though I haven't used those), since mods typically expand outwards from the original island, often by quite a bit.
- meheleventyone 6mo agoYeah in a lot of cases it's much better to use integers and a fixed precision as the absolute unit of position. For games it's just that the scale of most games works well with floats in the range they care about.
- desdenova 6mo agoThe problem with floating point comparison is not that it's nondeterministic, it's that what should be the same number may have different representations, often with different rounding behavior as well, so depending on the exact operations you use to arrive at it, it may not compare as equal, hence the need for the epsilon trick. If all you're comparing is the result from the same operations, you _may_ be fine using equality, but you should really know that you're never getting a number from an uncontrolled source.
- hpcgroup 6mo ago[dead]
- apitman 6mo ago> that's how maths works Wait is British "maths" a singular noun or is this a typo? I was willing to go along with it if it was plural, but I have to draw the line here.
- jonquark 6mo agoYes maths is singular, just like physics. We would say in the UK "maths is hard, physics is also hard"
- f33d5173 6mo ago"Maths" is short for "mathematics". The latter is not plural and can be substituted into this quote with no other alterations.
- 3836293648 6mo agoMaths is like physics
- adrian_b 6mo agoOriginally, maths/mathematics meant "things that are taught", like physics meant "natural things" and similarly for other such names. However, nowadays a word like physics is understood not as "natural things", but as an implicit abbreviation for "the science of natural things". Similarly for mathematics, mechanics, dynamics and so on. So such nouns are used as singular nouns, because the implicit noun "science" is singular.
- hansvm 6mo agoMy normal issue with floating-point epsilon shenanigans is that they don't usually pass the sniff test, suggesting something fundamentally wrong with the problem framing or its solution. It's a classic, so let's take vector normalization as an example. Topologically, you're ripping a hole in the space, and that's causing your issues. It manifests as NaN for length-zero vectors, weird precision issues too close to zero, etc, but no matter what you employ to try to fix it you're never going to have a good time squishing N-D space onto the surface of an N-D sphere if you need it to be continuous. Some common subroutines where I see this: 1. You want to know the average direction of a bunch of objects and thus have to normalize each vector contributing to that average. Solution 1: That's not what you want almost ever. In any of the sciences, or anything loosely approximating the real world, you want to average the un-normalized vectors 99.999% of the time. Solution 2: Maybe you really do need directions for some reason (e.g., tracking where birds are looking in a game). Then don't rely on vectors for your in-band signaling. Explicitly track direction and magnitude separately and observe the magic of never having direction-related precision errors. 2. You're doing some sort of lighting normalization and need to compute something involving areas of potentially near-degenerate triangles, dividing by those values to weight contributions appropriately. Solution: Same as above, this is kind of like an average of averages problem. It can make fuzzy, intuitive sense, but you'll get better results if you do your summing and averaging in an un-normalized space. If you really do need surface normals, store those explicitly and separate from magnitude. 3. You're doing some sort of ML voodoo to try to get better empirical results via some vague appeal to vanishing gradients or whatever. Solution: The core property you want is a somewhat strange constraint on your layer's Jacobian matrix, and outside of like two papers nobody is willing to put up with the code complexity or runtime costs, even when they recognize it as the right thing to do. Everything you're doing is a hack anyway, so make your normalization term x/(|x|+eps) with eps > 0 rather than equal to zero like normal. Choose eps much smaller than most of the vectors you're normalizing this way and much bigger than zero. Something like 1e-3, 1e-20, and 1e-150 should be fine for f16, f32, and f64. You don't have to tune because it's a pretty weak constraint on the model, and it's able to learn around it.
- darepublic 6mo agoPlus or minus eps
- dnautics 6mo ago> In reality it is a pretty deterministic (modulo compiler options, CPU flags, etc) IIRC this was not ALWAYS the case, on x86 not too long ago the CPU might choose to put your operation in an 80-bit fp register, and if due to multitasking the CPU state got evicted, it would only be able to store it in a 32-bit slot while it's waiting to be scheduled back in? It might not be the case now in a modern system if based on load patterns the software decides to schedule some math operations or another on the GPU vs the CPU, or maybe some sort of corner case where you are horizontally load balancing on two different GPUs (one AMD, one Nvidia) -- I'm speculating here.
- dunham 6mo agoI was bit by this years ago when our test cases failed on Linux, but worked on macos. pdftotext was behaving differently (deciding to merge two lines or not) on the two platforms - both were gcc and intel at the time. When I looked at it in a debugger or tried to log the values, it magically fixed itself. Eventually I learned about the 80-bit thing and that macos gcc was automatically adding a -ffloat-store to make == more predictable (they use a floats everywhere in the UI library). Since pdftotext was full of == comparisons, I ended up adding a -ffloat-store to the gcc command line and calling it a day.
- Dylan16807 6mo ago> IIRC this was not ALWAYS the case, on x86 not too long ago the CPU might choose to put your operation in an 80-bit fp register, and if due to multitasking the CPU state got evicted, it would only be able to store it in a 32-bit slot while it's waiting to be scheduled back in? I don't think the CPU was ever allowed to do that, but with your average compiler you were playing with fire. Did any actual OS mess up state like that? They could and should save the full registers. There's even a bultin instruction for this, FSAVE.
- Negitivefrags 6mo agoThis is the kind of misinformation that makes people more wary of floats than they should be. The same series of operations with the same input will always produce exactly the same floating point results. Every time. No exceptions. Hardware doesn't matter. Breed of CPU doesn't matter. Threads don't matter. Scheduling doesn't matter. IEEE floating point is a standard. Everyone follows the standard. Anything not producing indentical results for the same series of operations is *broken*. What you are referring to is the result of different compilers doing a different series of operations than each other. In particular, if you are using the x87 fp unit, MSVC will round 80-bit floating point down to 32/64 bits before doing a comparison, and GCC will not by default. Compliers doesn't even use 80-bit FP by default when compiling for 64 bit targets, so this is not a concern anymore, and hasn't been for a very long time.
- mcv 6mo agoThis is mostly about game logic, where I can understand the reliance on floating point numbers. I've also seen these epsilon comparisons in code that had nothing to do with game engines or positions in continuous space, and it has always hurt my eyes. I think if you want to work with values that might be exactly equal to other values, floating point is simply not the right choice. For money, use BigDecimal or something like that. For lots of purposes, int might be more appropriate. If you do need floating point, maybe compare whether the value is larger than the other value.
- thayne 6mo agoSo they say you should use epsilons, then their solution to the first problem is to use an epsilon. There may be some cases when you can get by without using epsilon comparison, but in many cases epsilon comparison is the right thing to do, but you need to choose a good value for it. This is especially true in cases where the number comes from some kind of input (user controls, sensor reading, etc.) or random number generation.
- GuB-42 6mo agoThe thing with floating point numbers is they are meant to work with physical quantities: distances, durations, etc... Physical quantities involve imprecision: measurement devices, tools, display devices, ADC/DACs etc... They all have some tolerances. And when you are using epsilons, the epsilon value should be chosen based on that physical value. For example, you set the epsilon to 1e-4 because that's 100 microns and you can't display 100 micron details. That's also the reason why there is not one size fits all solution. If you are working with microscopic objects, 100 microns is huge, and if you are doing a space simulation, 1 km may be negligible. Some operations involve a huge loss of precision, some don't, and sometimes you really want exact numbers and therefore you have to know your fractional powers of 2.
- waffletower 6mo agoThis is highly reductive, "they are meant to work with physical quantities", but agree that the applicability of an epsilon is entirely situational.
- beyondCritics 6mo agoIf your code may be compiled, to use the Intel x87 numerical coprocessor, an important issue is the so called "excess precision": Different values on chip can collapse after being rounded and stored to their memory locations, invalidating previous comparisons. Spilling can happen unexpectedly. Note that Intel calls the x87 "legacy"
- pclmulqdq 6mo agoNobody's code will be compiled to use x87 any more.
- beyondCritics 6mo agoThere is plenty of demand for so called "secure code", where such coder arrogance will not be tolerated. Trust me on that, I know it.
- adgjlsfhk1 6mo agomodern compilers do just have options to disable using x87 registers entirely.
- pclmulqdq 6mo agoWhat is "coder arrogance"? The expectation that the 20-year-old API (SSE) will be used over the deprecated 40-year-old API that is currently losing compiler (and silicon) support? Can you point to a running x86 machine today that does not have SSE? It is also faster and more precise to use double-double for math over x87 extended precision. There is literally no reason to compile an x87 instruction today aside from a programmer not knowing better.
- deleted 6mo ago[deleted]
- lisper 6mo agoWell, at least the author is honest about it: > The title of this post is an intentional clickbait. Unfortunately, that's where the honesty ends. > It's NOT OK to compare floating-points using epsilons. > So, are epsilons good or bad? Usually bad, but sometimes okay. So which is it? Emphatically NOT OK, or sometimes okay?
- mtklein 6mo agoMy preference in tests is a little different than just using IEEE 754 ==, _Bool equiv(float x, float y) { return (x <= y && y <= x) || (x != x && y != y); } which both handles NaNs sensibly (all NaNs are equivalent) and won't warn about using == on floats. I find it also easy to remember how to write when starting a new project.
- z_open 6mo agoWhy is that NaN handling sensible? I don't think it makes sense to say log(-1) equals log(-2). Mathematically it isn't true and your implementation would say it's true only because of limitations in IEEE754.
- xigoi 5mo agoMathematically it’s also not true that 2³⁰⁹ = +∞. You simply can’t expect correct results from floats.
- tengwar2 6mo agoCould you explain the bit about "all NaNs are equivalent"? IEEE requires that NaN ≠ NaN, but I suspect that I am just misunderstanding what you mean.
- mtklein 6mo agowhat I mean here about NaNs is that from a testing perspective, I want to be able to write a test that expects NaN in the same way that I write other expectations, and you can't do that with ==. assert(x == 7); // fine assert(y == NaN); // never true assert(y != y); // this is what you meant so this equiv() helper fixes that, assert(equiv(x, 7)); // fine assert(equiv(y, NaN)); // also fine now, as far as treating NaNs equivalently, the IEEE 754 float format has a huge number of possible representations of NaN, and if you did something like a bitwise comparison, you might think that 0x7fc00000, 0x7f800001, 0xffc00000, 0x7fc0f00d were all different and not equivalent, but they're all NaNs, and I find that when I'm looking for a NaN, I very rarely care about exactly which one I'm looking at. So checking (x!=x && y!=y) admits any two NaNs as equivalent.
- Cold_Miserable 6mo agoI've yet to encounter a need for == equality for floating point operations.
- spacechild1 6mo agoI need it all the time. A very common case is caching of (expensive) computations. Let's say you have a parameter for an audio plugin and every parameter change requires some non-trivial and possibly expensive computation (e.g. calculation of filter coefficients). To avoid wasting CPU cycles on every audio sample, you only do the (re)calculation when the parameter has actually changed: double freq = getInput(0); if (freq != mLastFreq) { calculateCoefficients(freq); mLastFreq = freq; } Also, keep in mind that certain languages, such as JS, store all numbers as double-precision floating point numbers. So every time you are writing a numeric for-loop in JS you are implicitly relying on floating point equality :)
- larrry 6mo agoSome environments only expose float as the number type, love2d being one I know. Fortunately love is built on LuaJIT, which does support integer math through the built-in ‘ffi’ library. My multiplayer arena game rounds reasonably-sized floats and compares them in the presentation layer, but uses fixed point integer math in the core rollback simulation (I believe JS is a similar story, its number can be either int or float with no way to guarantee integer-only math. I never needed to consider the difference outside of currency for webdev, so I’m less sure)
- Joker_vD 6mo agoTo quote from one of my previous comments: > > the myth about exactness is that you can't use strict equality with floating point numbers because they are somehow fuzzy. They are not. > They are though. All arithmetic operations involve rounding, so e.g. (7.0 / 1234 + 0.5) * 1234 is not equal to 7.0 + 617 (it differs in 1 ULP). On the other hand, (9.0 / 1234 + 0.5) * 1234 is equal to 9.0 + 617, so the end result is sometimes exact and sometimes is not. How can you know beforehand which one is the case in your specific case? Generally, you can't, any arithmetic operation can potentially give you 1 ULP of error, and it can (and likely, will) slowly accumulate. Also, please don't comment how nobody has a use for "f(x) = (x / 1234 + 0.5) * 1234": there are all kinds of queer computations people do in floating point, and for most of them, figuring out the exactness of the end result requires an absurd amount of applied numerical analysis, doing which would undermine most of the "just let the computer crunch the numbers" point of doing this computation on a computer.
- bananzamba 6mo agoIf you need equality, just use fixed point
- Panzerschrek 6mo agoI agree that using fixed point in many cases is a better option. But floating point is assumed to be the default choice for computations with non-integer numbers, because almost all popular programming languages have only floating-point built-in types, but not fixed point types.
- jheriko 6mo ago[dead]
- Asooka 6mo agoMy one small nitpick is that vector length is usually 2 instructions with SSE4: dpps xmm0, xmm0, 0x17 ; dot product of 3 lanes, write lane 0 sqrtss xmm0, xmm0 ret And is considerably faster than the fancy version, mainly because Intel still hasn't given us horizontal-max vector instruction! ARM is a bit better in that regard with their fancy vmaxvq_f32 and vmaxnmvq_f32...
- _moof 6mo agoSomething I've observed as someone who works in the physical sciences and used to work as a software engineer is: Very few software engineers understand that tolerances are fundamental. In the physical sciences, strict equality - of actual quantities, not variables in equations - is almost never a thing. Even though an equation might show that two theoretical quantities are exactly equal, the practical fact is that every quantity begins life as a measurement, and measurements have inherent uncertainty. Even in design work, nothing is exact. It's simply not possible. A resistor's value is subject to manufacturing tolerances, will vary with temperature, and can drift as it ages. A mechanical part will also have manufacturing tolerance, and changes size and shape with temperature, applied forces, and wear. So even if a spec sheet states an exact number, the heading or notes will tell you that this is a nominal value under specific conditions. (Those conditions are also impossible to achieve and maintain exactly for all the same reasons.) Even the voltages that represent 0 and 1 inside a computer aren't exact. Digital parts like CPUs, GPUs, RAM, etc. specify low and high thresholds, under or over which a voltage is considered a 0 or 1. Floating-point numbers have uses outside the physical sciences, so there's no one-size-fits-all approach to using them correctly. But if you are writing code that deals with physical quantities, making equality comparisons is almost always going to be wrong even if floating-point numbers had infinite precision and no rounding error. Physical quantities simply can't be used that way.
- Nevermark 6mo ago> It's OK to compare floating-points for equality Unless you use any compiler I write. As a purist, '==' for floats and doubles will be undefined. And by undefined, I mean unrecoverable. And by unrecoverable, I mean they may be able to extract most of the fragments of silicon, aluminum, and dynamic island, and "put you back together again". But you will "think different". Unless you are comparing references. References are ok.
- Panzerschrek 6mo agoIt's a good advice to avoid using floating point at all, if it's not that hard to do so. In one of my hobby projects I have written a simple map application based on OSM data with geometry processing code operation on integers only.
- unholiness 6mo agoI cringed very hard in the slerp example seeing `acos(dot(a,b))`. Clamping to [-1,1] to avoid NaNs still gives you bad answers and numerical sensitivity around small angles. acos and asin in general lose ~half your sig figs around their singularities[0]. Working around this by introducing a threshold seems like exactly same flavor of issue he's complaining about to begin with. There are perfectly good solutions for finding the angle using atan and the cross product. Calculating A x B will yield a vector which lies is along their normal with length tan(θ)(A · B) so we can straightforwardly say e.g.: `θ = atan2(norm(A ⅹ B), (A · B))` No thresholds, no branching, only ~1 bit of significance lost. As the original author says himself, when you're introducing arbitrary-feeling thresholds, it's likely you're missing a solution which would improve more than just the weirdness at those thresholds. [0] A good analysis of this is on Page 46 here, including a more optimized (but less obviously true) alternative to the formula above: https://people.eecs.berkeley.edu/~wkahan/Mindless.pdf https://people.eecs.berkeley.edu/~wkahan/Mindless.pdf