5 ms·
The python package Hypothesis[0] already does a great job bringing property-based testing to the people! I've used it and it's extremely powerful. [0]: https:/
by tomnicholas1 2y ago
The python package Hypothesis[0] already does a great job bringing property-based testing to the people! I've used it and it's extremely powerful.
[0]: https://github.com/HypothesisWorks/hypothesis https://github.com/HypothesisWorks/hypothesis
- epgui 2y agoI have used Python's `hypothesis` as well, and I wish it were better. We had to rip it out at work as we were running into too many issues. I have also used Haskell's `QuickCheck` and Clojure's `spec` / `test.check` and have had a great experience with these. In my experience they "just work". Conversely, if you're trying to generate non-trivial datasets, you will likely run into situations where your specification is correct but Hypothesis' implementation fails to generate data, or takes an unreasonable amount of time to generate data. Example: Generate a 100x25 array of numeric values, where the only condition is that they must not all be zero simultaneously. [1] [1] https://github.com/HypothesisWorks/hypothesis/issues/3493 https://github.com/HypothesisWorks/hypothesis/issues/3493
- mrcsd 2y agoCare to expand upon the issues you were running into with hypothesis? I'm genuinely curious as I may soon be evaluating whether to use it in a professional context.
- rtpg 2y agoI understand your pain in some sense, but on another I feel like people with a decent amount of hypothesis experience "know" how the generator works and would understand that you basically _never_ want to use `filter` if you can avoid it, instead relying on unfalsifiable generation. Silly idea for your generator would to generate an array, and if it's zero... draw a random index and a random non-zero number and add it into the array. Leads to some weird non-convexity properties but is a workable hack. In your own example you turned off the "data too slow" issue, probably because building up a dataframe (all to just do a column sum!) is actually kind of costly at large numbers! Your complaint is probably actually meant for the pandas extras (or pandas itself) rather than the concept of hypothesis.
- epgui 2y agoNo, I ran into the same issues with basic data structures. The dataframe wasn’t necessary, it just matched the expected input of some function I wanted to test.
- rtpg 2y agoI took your case, I got way better perf just generating a list of numbers and then reshaping it into a dataframe. But! Even though it doesn't even get that much slower at a certain number of rows it just starts hanging! Like at 49 rows everything is still fine and at 50 it no longer wants to work. It's very bizarre and I'll see if I can debug it. But I think your test case isn't indicative of some fundamental issue with Hypothesis rather than some sort of bug.
- tybug 2y agoThat kind of behavior can happen at the threshold of Hypothesis' internal limit on entropy - though if you're not hitting HealthCheck.data_too_large then this seems unlikely. Let me know if you have a reproducer, I'd be curious to take a look.
- deleted 2y ago[deleted]
- epgui 2y ago> Even though it doesn't even get that much slower at a certain number of rows it just starts hanging Yes, this brings back memories. I've definitely seen this kind of behaviour as well, in many different, not-particularly-exotic, situations. I am absolutely convinced the issue I raised on the github project was a bug or a defect, despite the maintainers not taking it seriously. I find QuickCheck and Clojure spec/test.check much more straightforward to use. I just never ran into this sort of thing with these other tools.
- chriswarbo 2y agoAs the comments on your linked issue point out: (a) Filtering is a last resort and is best avoided. As an example, the Gen type in Haskell's falsify package can't be filtered, since it's a bad idea. As another example, ScalaCheck's Gen type can be filtered, but they also allow "retries" (by default, up to 10,000 times), because filtering is very wasteful. (b) If you're going to filter, scope it to be as small as possible (e.g. one comment points out that you're discarding and regenerating entire dataframes, when the filter only depends on one particular column) (c) Have some vague awareness of how your generators will shrink, to avoid infinite loops. In your case, shrinking will make it more likely to fail your filter; and the "smallest" dataframe (all zeros) will definitely fail.
- skybrian 2y agoIt’s still very weird that the generator can’t avoid an all-zeros array. Figuring out the root cause might find something interesting.
- dllthomas 2y agoNot weighing in on any particular tech (hypothesis or otherwise), but intrigued by your example... My initial impulse is to pick a random cell which must not be zero, generate a random number for each other cell and a random non-zerp number for that one. I'm not immediately decided on whether it's uniformly distributed.
- eslaught 2y agoI would pick the number of non-zeros first, assert that it's non-zero, then continue filling in the values themselves. And probably not with a uniform distribution. Any algorithm that cares about the number of non-zeros could have non-trivial interactions with their arrangement and count, so picking something that generates non-trivial sparsity (and doesn't just make the array look like white noise) is going to have the best chance of exposing interesting behavior. The tricky part is thinking through how to generate "interesting" patterns, which admittedly I haven't put enough thought into.
- dllthomas 2y agoAh, yeah, generating for property testing probably doesn't want a uniform distribution. What patterns are interesting will surely depend on what we're doing with the array.
- jgalt212 2y agoIndeed, the world is not IID. As such, test cases should not be a uniformly distributed sample of some mathematical distribution.
- dllthomas 2y ago> Indeed, the world is not IID. Right! And even if it were, in the sense that that's what we should expect as real world input, it wouldn't generally be the best distribution for finding bugs.
- chriswarbo 2y agoAs far as I'm aware, Hypothesis is fundamentally based around the idea of "generators are parsers of randomness" discussed in this paper; i.e. a Hypothesis "strategy" is essentially a function from bytestrings to values. To generate random values, those strategies are run on a random bytestring; to shrink a previous value, the bytestring that lead to that value is shrunk. Haskell's "falsify" package takes a similar approach, but uses a tree of random values. This has the advantage that composite generators can run each of their parts against a different sub-tree, and hence they can be shrunk independently without interfering.