19 ms·
As a general rule, if you find yourself comparing floating point numbers, that's probably not what you want. I'm wondering what they are going to do with other
by cesaref 3y ago
As a general rule, if you find yourself comparing floating point numbers, that's probably not what you want.
I'm wondering what they are going to do with other operators, >=, <= etc
- ssivark 3y agoThis smells like the right way to think about things. Whenever thinking of real life measurements represented as real numbers, equality is always fuzzy — up to some resolution, and it would be better to make that explicit.
- nawgz 3y ago> if you find yourself comparing floating point numbers, that's probably not what you want Can you motivate this with an example for me? For instance, I think of games or location data encoded as FP; then, clearly, comparing them is a critical task to know questions like "what is closer" and so on. What am I missing?
- stefncb 3y agoThey meant using the equal operator, because floating point is inexact and can produce different representations for the same number depending on how it was obtained. Greater and less than are fine. Equal is usually implemented by seeing if a number fits inside a tight range.
- recursive 3y agoEvery finite floating point value precisely represents an exact quantity. Arithmetic isn't lossless though. This isn't a special property of floats. Decimals behave the same way.
- nawgz 3y agoI was with you until the end First point: FPs are unique. Clearly true Second point: Combing FPs isn't lossless, across most/all arithmetic. Clearly true, the root of our discussion Third point: decimal arithmetic is lossy too. Strongly disagree. The system of representation is what is lossy - floating point. Arithmetic (between two decimal numbers) is clearly not lossy in and of itself, only as implemented with computers.
- Dylan16807 3y agoDecimal is only lossless if you track an indefinitely-growing suffix to your number with a big bar over it to indicate "repeating". It's basically working with rational numbers but more painful. Decimal numbers with any limit on digits, even a hundred, are lossy. Decimal numbers as used by humans outside of computers have limits on digits, so they are lossy.
- tadfisher 3y agoLikewise, bit-limited integers are lossy as they are incomplete representations of ℤ, and are not reciprocal with division. The point being that, just like floating-point representations of ℝ, division will always be lossy in the general case unless you choose a symbolic algebra instead of numeric.
- nawgz 3y agoMaybe my point is uninteresting, but I feel you haven't really understood it. Combining two FP numbers with arithmetic leads to possible errors that you can't represent with FP on every operation. Combining two decimal numbers with non-division arithmetic never leads to that case. With division, sure, things are bad, but that's more of an exception than a rule to me. This is because decimal numbers don't really come with any inherent limit on digits, and it's a bit strange to add a clause to the other's claims before making a counterargument.
- recursive 3y ago> Combining two decimal numbers with non-division arithmetic never leads to that case. Most types are subject to overflow that has similar effects. Most of the FP error people encounter is actually division error. For example, the constant 0.1 contains an implicit division. The "tenths" place is defined by division by 10. I think that almost all perceptions of floating point lossiness come from this fact.
- Dylan16807 3y agoWell in part I didn't understand your point because you didn't mention division as an exception. But even then, multiplication causes an explosion in number length if you do it repeatedly. When you specifically talk about numbers not being in computers, I think it's fair to talk about digit limits. Most real-world use of decimals is done with less precision than the 16 digits we default to in computers. Let alone growing to 50, 100, 200, etc as you keep multiplying numbers together to perform some kind of analysis. Nobody uses decimal like that. Real-world decimal is lossy for a large swath of multiplication. I agree that if you're doing something like just adding numbers repeatedly in decimal, and those numbers have no repeating digits, then you have a nice purity of never losing precision. That's worth something. But on the other hand if you started with the same kind of numbers in floating point, let's say about 9 digits long, you could still add a million of them without losing precision. And nobody has said anything about irrational numbers as you dismissed in your other comment. So in summary: decimal division, usually lossy; decimal multiplication, usually lossy; decimal addition and subtraction, lossless but with the same kind of source numbers FP is usually lossless too
- zokier 3y agoCommon way of comparing floats is to compare the difference to some epsilon value. See Python PEP-485 for example https://peps.python.org/pep-0485/ https://peps.python.org/pep-0485/
- deleted 3y ago[deleted]
- orthoxerox 3y agoThis is specifically about equality. Comparing FPs for equality is very risky, as your numbers can differ by 0.00000000000001 without anyone noticing. Strict inequality (> and <) comparison are generally fine as long as you avoid NaNs.
- nawgz 3y agoOk, sure, so it was just for a specific case of comparison. Well, followup, if some error Epsilon can be introduced during any manipulations, when you check for strict inequalities do you check x < y + Epsilon? Does the language implicitly do it for you?
- klodolph 3y agoThe language just gives you the direct comparison. If you know numerics, then you can come up with the correct value for epsilon. But that is hard work. There are various things that the language could do for you, like use a dynamic amount of precision, or interval arithmetic where the error bounds are saved—but! Most of the time people just want the answer faster and with less memory used, which is what you get with bare floats. The people who make it their job to care about numerical accuracy can do it better than the language runtime would anyway. Most of the problems are more easily solvable at a higher level anyway. Like, imagine Clippy saying to you, “It looks like you’re inverting a matrix. Are you sure that’s a good idea?”
- HexDecOctBin 3y agoIdeally, you'll want to do distance calculations in fixed point numbers (of desired resolution). Floating point works well as a first approximation, but unless you can make sure that you are not going to end up wandering in the weeds (and ensuring that required understanding of computational numerics), you should probably replace them with fixed point once you understand the problem and the solution and know the limits involved.
- zdragnar 3y agoCaveat: In the following, I use "comparison" to mean "check for equality". Floating point numbers lose precision because binary arithmetic doesn't represent decimal in all cases ( 1/3 is an easy example). It's not hard to get into a situation where you're asking if a number is 0.0 but due to a precision error the number you have is 0.00000000000001 or whatever and it should have been zero. If you're dealing with anything where precision is paramount (i.e. money or high precision machining) you should consider using something else that matches the real work precision you are trying to model. You can, of course, use floats and test differences rather than strict equality (i.e. distance is < 0.01) and that's fine too... if you remember to do it, and account for the potential for small precision differences to propagate throughout your calculations and accumulate into larger errors.
- tobr 3y ago> Caveat: In the following, I use "comparison" to mean "check for equality". You never use the word “comparison” again in the comment!
- Someone 3y ago> Floating point numbers lose precision because binary arithmetic doesn't represent decimal in all cases. That’s not the reason. A counterexample is decimal arithmetic. That represents decimal in all cases, yet loses precision when doing calculations. Once you start doing division that applies even to arbitrarily sized decimal numbers. The correct reason is because the subset of the reals that is representable as IEEE float isn’t closed under the operations people want to perform on them, including the basic ones of addition, multiplication and division.
- jandrese 3y agoIMHO That was poorly worded. What it should have said is "if you find yourself using equals to compare floating point numbers...". With Floating Point your comparisons should always be less than or greater than. Precision artifacts make the equals unreliable, and you should always be mindful of that when dealing with them.
- BlueTemplar 3y agoIt's the hard task of trying to figure out the magnitude of the expected errors (which can accumulate), in the simplest case of a single operation you compare within epsilon distance : https://en.wikibooks.org/wiki/Floating_Point/Epsilon https://en.wikibooks.org/wiki/Floating_Point/Epsilon
- atemerev 3y agoWe don't have fast hardware decimals, except on IBM mainframes. So, if you need fast financial calculations which are too complicated for integers (e.g. gasp division), you usually use decimally-normalized doubles and do it very very carefully. This is a sad state of affairs, but nobody was fixing this in the last 30 years, so there's that.
- cubefox 3y agoAre there even any applications where one needs a) decimal precision and b) fast computation? Financial calculations don't need to be fast.
- nvy 3y agoHigh-frequency trading, perhaps. Physics simulations too.
- wtallis 3y agoPhysics simulations are a great example of where there's obviously no need to do arithmetic in decimal rather than binary floating point.
- atemerev 3y agoAlgorithmic trading, of course. High-frequency trading in particular. But any trading, in fact, when you run simulations for thousands of instruments and search among millions of parameters. Or when you implement an exchange, or route orders between exchanges according to their complicated criteria. Or when you do options pricing. Basically, everything in this field.
- deleted 3y ago[deleted]
- justeleblanc 3y agoSo something that produces nothing of value to civilization. Gotcha.
- wslh 3y agoFloating point numbers are dangerous but in 2023 we should find a way to improve the side effects. Playing with Python3: >>> +0.0==-0.0 True The C in GCC below also returns 1: #include <stdio.h> int main() { float a = +0.0; float b = -0.0; printf("%d\n", a == b); }
- Y_Y 3y agoErlang too, now and in the future will say that minuszero==zero, it's a different more specific operator that's changing. You'll get the same in python with >>> -0. is 0. False
- Kwpolska 3y agoThe `is` operator is checking object identity by comparing addresses. It is useless, although it may sometimes produce reasonable-looking results due to optimisations. Python 3.11 gives me this: >>> -0 is 0 <stdin>:1: SyntaxWarning: "is" with a literal. Did you mean "=="? True Because (a) -0 is an integer, and there is only one integer zero, (b) small numbers have only one instance in memory as an optimization.
- flatline 3y agoIt is 2023 and our tooling still encourages the same mistakes people were making 40 years ago. Can we really not have equality operators that do a comparison with a 1% tolerance or something as a sensible default equivalence instead of blatantly wrong bitwise comparison? I would even be happy with a default set of compiler warnings or errors, which I don’t believe I have ever seen.
- colejohnson66 3y agoThe problem with defining an “epsilon” (1% in your case) is that there is no value that would please everyone. For some, 1% will be fine, but others may need 0.0001%. The solution is to either use decimal floats if they suit your need (and eat the performance penalty), or to use a linter that flags float comparisons by equality.
- dzaima 3y agoAnd then there's the question of how you use the epsilon - i.e. whether 1.2e-16 and 1e-20 are "within 1%". Sometimes those are very much different sizes, but if you're comparing sin(π) (i.e. the 1.2e-16) with some other almost-zero computation, they're very much within 1%.
- Someone 3y agoThat’s bad enough, but it is not the problem with defining an “epsilon”, it’s a problem. Another one is that you would lose transitivity on equality. With a 1% epsilon, you would have 1.000 == 1.005 1.005 == 1.010 1.010 == 1.015 but 1.000 != 1.015 You also would have to be careful on how to define that 1% error range. the naive "x is equal to all numbers between 0.99x and 1.01x" would mean 1.0 would be considered equal to 0.99, but 0.99 would not be considered equal to 1.0 (1.01 times 0.99 is less than 1.0) You also lose that, if a == b and c == d it follows that a + c == b + d. The behavior around zero also will go against intuition. If you consider 0.99 and 1.0 to be equal, do you really want 1.0E-100 and -1.0E-100 be different?
- chubot 3y agoFrom a user POV, I think == and === on floats should simply be undefined in any language. It should be a compile time or runtime error. There can be a separate function `float_equals()` with explicit args that does what people want The only reason to use the same syntax == is for POLYMORPHIC code that is actually correct. But it's not going to be correct with floats, because they don't obey the same algebraic laws ... So the syntax should be different! --- I believe the Erlang compiler optimization presented as justification for this change is a good example The compiler wants to reason about code, independent of types But that reasoning about the =:= operator is wrong for the float case.
- regularfry 3y agoI've always held that the numbers themselves are perfectly precise. It's the operations that don't do what you expect. Of course, that observation may be more or less useful, depending on circumstances.
- ChainOfFools 3y agoSomewhere in the distant past a second grade me is staring at his math homework and fuming over the frustratingly ambiguous meaning of the minus sign. Is it an operator? A property? Both?
- Dylan16807 3y agoIt's only ambiguous if you think it does two things. I say it does one thing. -4 always removes 4, whether or not there is a number to the left.
- ChainOfFools 3y agoThe problem arises with the way that this stuff is taught at a very elementary level, which then needs to be thrown away and relearned long after intuition and habits have formed around the old model. As an adult my reasoning is that it does indeed only do one thing: negation. it doesn't take anything away at all, if anything it actually adds information. Negating a four doesn't take it away, it just specifies a different (inverted) mapping of a given value with respect to zero. And this is still a very naive take, as someone with no background in number theory. But it's much more useful than say, Billy has four apples, he gives three to Mary, etc.
- thr-nrg 3y agoMath notation is to mathematics as poetry is to English. The only acceptable grammar for math is s-expressions.
- Conscat 3y agoWell, idiomatic C++23 is an acceptable math notation too.
- sterlind 3y agoI'm assuming you mean equality comparison, rather than comparison in general! but is there a better algebraic structure to use when thinking of floats? like, in terms of limits, or something that tracks significant digits and uncertainty? how do formal methods handle these?
- Y_Y 3y agoThere's more than one way to skin that cat, but interval arithmetic is a good and simple model. You can take the part of the real line which maps to the single float you are thinking about, and then look at the image of that set under the whatever functions you are thinking about.
- toast0 3y agoSince they're not changing ==, I wouldn't think =< and >= would change. Note, Erlang has a no arrow looking comparison rule, so <= is not a comparison operator; actually it's used in bitstring comprehensions, the bitstring version of list comprehensions which use <- [1] [1] https://www.erlang.org/doc/reference_manual/expressions.html#bit-string-comprehensions https://www.erlang.org/doc/reference_manual/expressions.html...
- abhgh 3y agoAs a followup to this comment, in Python you have the option of using `isclose()` which is present in both `numpy`[1] and the standard `math` [2] libraries. This has been quite helpful for me in comparing small probability values. [1] https://numpy.org/doc/stable/reference/generated/numpy.isclose.html https://numpy.org/doc/stable/reference/generated/numpy.isclo... [2] https://docs.python.org/3/library/math.html#math.isclose https://docs.python.org/3/library/math.html#math.isclose
- spacechild1 3y agoIn general, this is good advice. However, there are cases where comparing floating point numbers for equality works just fine. For example, if you have a variable 'v' and want to update it to a new value 'f', you can do 'if (v == f)' to check if the variable would change. Or if you have a sentinel value -1.0, it is perfectly ok to do 'if (a == -1.0)' In general, if you look for a number 'f' in variable 'v', you can safely use equality as along as you expect 'v' to be set to 'f', i.e. the value of 'v' is not the result of a computation. This may sound trivial, but this is often omitted when discussing floating point comparison. If someone knows a scenario where the examples above would break, I would be curious to hear! (My main field is audio programming and code like this is omnipresent.)
- stouset 3y ago> For example, if you have a variable 'v' and want to update it to a new value 'f', you can do 'if (v == f)' to check if the variable would change. Isn't it strictly faster to simply assign? No branching, no branch predictions, only one instruction in all cases.
- kstenerud 3y agoIt depends on whether you need a set operation or a test-and-set operation.
- karpierz 3y agoAssuming that all you want to do is assign, then sure. If you want to log the change or do any additional conditional logic, then that wouldn't work.
- spacechild1 3y agoA typically pattern is: float freq = getParameter(FILTER_FREQUENCY); float coeff; if (freq != mFreq) { coeff = calculateCoeffFromFrequency(freq); // expensive mFreq = freq; mCoeff = coeff; } else { coeff = mCoeff; // use cached value } // use coeff
- 3y ago
- kragen 3y agothis is not about arithmetic comparison, it's about erlang's exact-equality operator