4 ms·
It is 2023 and our tooling still encourages the same mistakes people were making 40 years ago. Can we really not have equality operators that do a comparison wi
by flatline 3y ago
It is 2023 and our tooling still encourages the same mistakes people were making 40 years ago. Can we really not have equality operators that do a comparison with a 1% tolerance or something as a sensible default equivalence instead of blatantly wrong bitwise comparison?
I would even be happy with a default set of compiler warnings or errors, which I don’t believe I have ever seen.
- colejohnson66 3y agoThe problem with defining an “epsilon” (1% in your case) is that there is no value that would please everyone. For some, 1% will be fine, but others may need 0.0001%. The solution is to either use decimal floats if they suit your need (and eat the performance penalty), or to use a linter that flags float comparisons by equality.
- dzaima 3y agoAnd then there's the question of how you use the epsilon - i.e. whether 1.2e-16 and 1e-20 are "within 1%". Sometimes those are very much different sizes, but if you're comparing sin(π) (i.e. the 1.2e-16) with some other almost-zero computation, they're very much within 1%.
- Someone 3y agoThat’s bad enough, but it is not the problem with defining an “epsilon”, it’s a problem. Another one is that you would lose transitivity on equality. With a 1% epsilon, you would have 1.000 == 1.005 1.005 == 1.010 1.010 == 1.015 but 1.000 != 1.015 You also would have to be careful on how to define that 1% error range. the naive "x is equal to all numbers between 0.99x and 1.01x" would mean 1.0 would be considered equal to 0.99, but 0.99 would not be considered equal to 1.0 (1.01 times 0.99 is less than 1.0) You also lose that, if a == b and c == d it follows that a + c == b + d. The behavior around zero also will go against intuition. If you consider 0.99 and 1.0 to be equal, do you really want 1.0E-100 and -1.0E-100 be different?
- chubot 3y agoFrom a user POV, I think == and === on floats should simply be undefined in any language. It should be a compile time or runtime error. There can be a separate function `float_equals()` with explicit args that does what people want The only reason to use the same syntax == is for POLYMORPHIC code that is actually correct. But it's not going to be correct with floats, because they don't obey the same algebraic laws ... So the syntax should be different! --- I believe the Erlang compiler optimization presented as justification for this change is a good example The compiler wants to reason about code, independent of types But that reasoning about the =:= operator is wrong for the float case.
- kragen 3y agowhat people want varies, but often what people doing numerical programming want is to write their monomorphic code in infix syntax so they can more easily see bugs in it they're already familiar with rounding errors from that perspective you're suggesting taking a step backwards from fortran i toward assembly language the erlang compiler bug presented is not an example of what you're talking about because you're talking about arithmetic, and the buggy comparison operator it's using is not an arithmetic comparison operator
- chubot 3y agoFortran would be an interesting case, because say C++ is polymorphic, and so are all dynamic programming languages with more than 1 type (x == y could be a string or float comparison) But I would still say that floating point equality is a vanishingly rare operation I can't think of any real use case for it, other than maybe some (bad, incomplete) unit tests that assert 1.0 == x and assert 1.0 == y And that use case is perfectly served by float_equals(), and arguably served better if it has some options. Although probably abs(x - y) < eps is just as good, and you don't even need float_equals() --- Can you show some real use cases for floating point equality in good, production code? (honest question) It's similar to hashing floats, which Go and Python do allow. I made an honest request for examples of float hashing: https://lobste.rs/s/9e8qsh/go_1_21_may_have_clear_x_builtin#c_hbqltt https://lobste.rs/s/9e8qsh/go_1_21_may_have_clear_x_builtin#... https://old.reddit.com/r/ProgrammingLanguages/comments/10bm2w9/bitwise_equality_of_floats/j4bvrkd/ https://old.reddit.com/r/ProgrammingLanguages/comments/10bm2... I didn't get any answers that appeared realistic. Some people said you might want to create a histogram of floats -- but that's obviously better served by bucketing floats, so you're hashing integers. Another person said the same thing for quantizing to mesh. One person suggested fraud detection for made-up values, but that was clearly addressed by converting the float to a bit pattern, and hashing the bit pattern. ---- This isn't a theoretical question since we're working on this for https://www.oilshell.org https://www.oilshell.org. For that case it seems pretty clearly OK to omit === on floats, since I use awk and R for simple stats and percentages, and that's likely what people would do with a shell. Though as always I'm open to concrete counterexamples.
- Dylan16807 3y agoThe tolerance you need depends on how you calculated the number, so there's no way to have a sensible default.
- jacquesm 3y ago> Can we really not have equality operators that do a comparison with a 1% tolerance or something as a sensible default equivalence instead of blatantly wrong bitwise comparison? Then it wouldn't really be an equality comparison any more. For instance according to your rule 1.0 and 1.01 are the same number, but they clearly are not. But if you want to do this regularly you could create a special type and overload the equality operator (if your language supports it, otherwise you'll have to call the function directly) that points to 'boolean about_equal(float f1, float f2, float epsilon)' or something equivalent where epsilon is the amount that you would allow f1 and f2 to deviate from each other to still call them equal. And you could define epsilon as ((abs(f1) + abs(f2)) / 100) . The bitwise comparison isn't 'blatantly wrong', it's the expected behavior in just about every programming language. I'm not aware of any language that has a native float operator that does this but as you can see you can usually simply add one yourself if you really need it. But this need rarely comes up and usually indicates that you are doing something wrong.