4 ms·
Anecdotally, Hypothesis was very far from "just working" for me. I don't think it's really production-ready, and that seems to be by design.[0] I did have quit
by epgui 2y ago
Anecdotally, Hypothesis was very far from "just working" for me. I don't think it's really production-ready, and that seems to be by design.[0]
I did have quite a bit of prior experience with clojure.spec.alpha (with or without test.check), so despite some differences it's not like I was completely alien to the general ideas.
[0] https://github.com/HypothesisWorks/hypothesis/issues/3493 https://github.com/HypothesisWorks/hypothesis/issues/3493
- eru 2y agoThe issue seemed to be based on a misunderstanding of the implied probability distribution hypothesis uses to draw examples from?
- epgui 2y agoThe probability distribution in my case seemed to be a single result (the zero matrix) with P=1. So if that is intentional and by design, then you’re right, I don’t understand.
- eru 2y agoHypothesis doesn't try to give you a consistent probability distribution, especially not a uniform one. It has lots of heuristics to give you 'nasty' numbers in order to trigger more bugs. That's actually one big reason I prefer Hypothesis over eg Rust's proptest, which makes no attempt to give you 'nasty' numbers. Last time I checked, QuickCheck didn't make any attempts to give you floats well-known to cause trouble either (like inf, negative inf, different NANs, 0, 1, -1, the least positive float, the largest negative float, the largest finite float just before infinity etc.)
- mjaniczek 2y agoThere has been some research (can't find the paper right now but it was about F-metric or something?) suggesting uniform generators will trigger bugs more often than the edge-case-preferring ones. I'm still not too sure about whether I want to believe it though.
- eru 2y agoI would be very interested in that research! Anecdotally, I have triggered many more bugs with the edge-case-preferring generation: Basically, that was me using Rust's proptest library which does uniform generation by default, and I hacked it up to prefer edge cases. For simplicity, for eg u32 I set it up to with something like 50% probability to pick one from 0, 1, 0xFFFF_FFFF, 2, and a few other special values, or to pick uniformly at random. I vaguely remember some research that compared carefully hand-picked values vs property based testing. And the carefully hand-picked values performed slightly better at the same number of test cases. But if you modesty increased the number of test cases generated for property based testing, that swamped the advantage of careful hand picking. Similarly, I suspect that any disadvantage that might exist for the scheme I outlined above compared to uniform sampling would go away, if you increased the number of test cases slightly. At least as long as you give a decent chunk of probability weight in my scheme to uniform sampling.
- epgui 2y ago> Hypothesis doesn't try to give you a consistent probability distribution, especially not a uniform one. This is totally understood and expected. But a “probability distribution” of exactly ONE value with P=1 is not much of a distribution. I don’t understand this response, because it doesn’t address the issue at all.
- mjaniczek 2y agoI've tried to replicate your issue in the Elm testing library I'm a maintainer of, which uses some of the core ideas from Hypothesis. It was able to generate values just fine ([1]), so your problems might have been just some issue of the Hypothesis _implementation_ instead of something wrong with the general algorithm? Note that your usage of Hypothesis was "weird" in two aspects: 1. You're using the test runner for generation of data. That's sometimes useful, particularly when you need to find a specific example passing some specific non-trivial property (eg. how people solve the wolf, goat, cabbage puzzle with PBT - "computer, I'm declaring it's impossible to solve the puzzle, prove me wrong!"), but your particular example had just "assert True" inside, so you were just generating data. There might be libraries better suited to the task. Elm has packages elm/random (for generation of data) and elm-explorations/test (for PBT and testing in general). Maybe Clojure or Python have a faker library that doesn't concern itself with shrinking, which would be a better tool for your task. The below gist [1] has an example of using a generator library instead of a testing library as well. 2. You were generating then filtering (for "at least one non-0 values across a column"). That's usually worse than constructively generating the data you want. An example alternative to the filtering approach might be to generate data, and then if a column has all zeroes, generate a number from range 1..MAXINT to replace the first row's column with. Anyways, "all zeroes" shouldn't happen all that often so it's probably not a huge concern that you filtered. [1]: https://gist.github.com/Janiczek/c71050aa89fb0b19d73e683f9a3cbb99 https://gist.github.com/Janiczek/c71050aa89fb0b19d73e683f9a3...
- epgui 2y agoThe code provided was not real test code. The only reason `assert True` is there is to highlight the fact that the issue was with generation, and not any other part of the test. > Anyways, "all zeroes" shouldn't happen all that often so it's probably not a huge concern that you filtered. This is really the point. Like I said, Clojure’s spec (with or without test.check) has no trouble.