10 ms·
“Expect tests” make test-writing feel like a REPL session
- CGamesPlay 4y agoSnapshot testing is great, and I wish more test frameworks included first-class support for them. This means that they can auto update with a flag, and can be stored either in the source inline or in an external file (both modes have different use cases). Note that doc tests can also be a form of this, e.g. in Python's. "Expect tests" seems like a bad name, since that covers all tests.
- chii 4y agoi find that snapshot testing gets overused in javascript - and mistakes can creep in easily, and if the snapshot is big, and in a separate file, code review can miss it. I much prefer property based testing over expectation based testing. You have to explicitly think about what properties hold true about the thing you're writing. For example, fib(N+1) = fib(N) + fib(N), so this property can be tested for all N; primitive generators can easily generate the data, and good composition framework can easily generate complex data from primitive data. Of course, you have to have a property you can specify easily. Otherwise, it'd be exactly the same as expectation based testing.
- IanCal 4y agoEvery single time I've introduced property based testing, even as a simple example, I've discovered a bug in either the code or the spec. I've found a bug in a Haskell program about fib generation - your test would work (if fixed for the subtractions) but incorrectly as there was an overflow in the addition. A basic property of "fib(n+1) > fib(n)" for n>1 finds this. I like this type of testing as it asks you to more generally consider what guarantees your code is making about its operation. Edit - your example is a good one and necessary, I just wanted to add a bit extra as I really like property based testing
- sesm 4y agoSnapshot testing works well for component systems, especially with storybook. There is a service called Chromatic that lets you diff component changes visually using storybook output.
- tantalor 4y ago> update with a flag Yes this is right level of automation, not whatever this article is going on about with the editor integration. Yuck.
- thedufer 4y agoThe open source use pattern for expect tests in OCaml (via dune) is exactly as you describe (see https://dune.readthedocs.io/en/stable/tests.html https://dune.readthedocs.io/en/stable/tests.html) - you run the tests with `--auto-promote` to tell it to update. The editor integration is a very simple keybinding on top of more generic tooling.
- ElliotH 4y agoI wonder if this has the same downsides as golden and screenshot type tests, where you end up over-asserting resulting in tests that break for unrelated changes? Obviously that’s a risk for hand written tests too but it’s easier (today… who knows what copilot like systems will offer soon!) for a human to reason about what’s relevant.
- avgcorrection 4y agoQ: Why does this test assert the value X? A: The value X was revealed to me by ChatGPT.
- potatoyogurt 4y agoYes, that is definitely a downside for these tests. The worst is when the text of some exception is printed and it includes line numbers. It does still require some discipline to think about what you're printing and avoid output that will be very noisy. This problem is mitigated quite a bit by the ease of accepting changes when these tests fail for obviously nonsense reasons though (just hit a couple buttons in an emacs buffer).
- vdm 4y agoA similar approach with pytest and pdb https://simonwillison.net/2020/Feb/11/cheating-at-unit-tests-pytest-black/ https://simonwillison.net/2020/Feb/11/cheating-at-unit-tests... This does get me writing tests sooner.
- foobarbecue 4y agoI do this with pytest-regtest's --regtest-reset command.
- avgcorrection 4y ago> I think you’re supposed to write some nonsense, like assert fibonacci(15) == 8, then when the test says “WRONG! Expected 8, got 610”, you’re supposed to copy and paste the 610 from your terminal buffer into your editor. > This is insane! The sane approach is presumably to either expand the call tree and verify all the unique subsolutions. Or to do every step with a calculator if you can’t expand the call tree. > The %expect block starts out blank precisely because you don’t know what to expect. You let the computer figure it out for you. In our setup, you don’t just get a build failure telling you that you want 610 instead of a blank string. You get a diff showing you the exact change you’d need to make to your file to make this test pass; and with a keybinding you can “accept” that diff. The Emacs buffer you’re in will literally be overwritten in place with the new contents [1]: Oh okay. The non-insane approach is to do the first thing but Emacs copies the result on your behalf.
- eru 4y agoWell, the non-insane thing is to do property-based testing. Instead of testing only a handful of examples.
- c-cube 4y agoThey also do that, the post refers to their Quickcheck library. But how do you property test the Fibonacci function ? There isn't much to say about it...
- avgcorrection 4y agoYou use the naive implementation as a test oracle, limit `n` to something small (through the property tester), and use the test oracle on your efficient implementation. Unit testing elegant functions has no value. (fib is often used as an example. But you asked how to test it.)
- davidgrenier 4y agoIn combinatorics the adjusted fibonacci numbers start with 1 instead and is more commonly used as it aligns with many other results. One might want to document in the code, via a test, which sequence is of interest. This is just an example of course but elegant functions might need to be tested.
- evrimoztamur 4y agoHere's an older post from 2015 (also from Jane Street) explaining the same process https://blog.janestreet.com/testing-with-expectations/ https://blog.janestreet.com/testing-with-expectations/, but at the infancy of the method. It looks like they heavily polished it! I like the approach, and I was indeed copy-pasting the result from my console...
- scotty79 4y agoDoesn't this approach make you update results of failing tests wholesale and possibly miss where a new result of some test is actually wrong? https://docs.rs/expect-test/latest/expect_test/ https://docs.rs/expect-test/latest/expect_test/
- eru 4y agoAt Google the nickname for these kinds of tests was 'change detector tests'.
- mannykannot 4y agoIf you are saying this approach would tend to produce a lot of change-detector tests, then that is an issue, but I think scotty79 is making a different point: this approach would seem to make it easy to overlook any regressions that the latest change has created.
- eru 4y agoYes, and that's exactly the issue with change detector tests.
- imajoredinecon 4y agoYeah, the OP's counterargument is that you can filter down what goes into the test output. But at that point it seems not too different qualitatively from the traditional bottom-up approach where you just write assertions yourself, except that the framework does the job of populating the assertions' expected values.
- mabbo 4y ago> But think: everything in those describe blocks had to be written by hand. It also had to be thought about by the developer. Someone had to say "I want the code to do this under these conditions". If your tests can be autogenerated then they aren't verifying expected behaviour, they're just locking in your implementation such that it can't change later. They are saying "hey look everyone, I got my coverage metric to 100% (despite any bugs I may have)."
- codetrotter 4y agoOne of the projects at a place where I have worked was set up so that when you ran the tests it automatically and silently updated the values that were expected. Completely bonkers because the first time I was contributing to the project I prepared the tests first and then started the implementation, and then while I was working on it I ran the tests which at this point should fail because I hadn’t finished writing the code but instead all tests passed. Because helpfully the test setup overwrote the expected values that I had prepared in my new tests, with the bad data. Yeah great, very helpful >:( Oh yeah and the whole test setup was also way too tied to the implementation rather than verifying behaviour. Complete trash the whole thing.
- travisjungroth 4y agoWould it do this just the first time? It’s still bad it was doing this silently, but it’s pretty common to test web APIs in a similar way manually. Make a request, check the response you get back looks right (important step) and then save it as the expected value. Edit: or after reading the article, like in the article.
- codetrotter 4y agoIt did this every time, not just the first time.
- travisjungroth 4y agoWell, you know what they say: Expect the unexpected!
- wmanley 4y agoWas discussed recently here: https://news.ycombinator.com/item?id=34350749 https://news.ycombinator.com/item?id=34350749
- bmitc 4y agoI don’t really understand this. How is this different from just writing the code and just assuming that you got it correct, and then locking in a potentially wrong implementation? > What does fibonacci(15) equal? If you already know, terrific—but what are you meant to do if you don’t? > I think you’re supposed to write some nonsense, like assert fibonacci(15) == 8, then when the test says “WRONG! Expected 8, got 610”, you’re supposed to copy and paste the 610 from your terminal buffer into your editor. Who does that? How do you know 610 is correct? That’s just assuming your implementation is right from the get go. For such a function, I’d independently calculate it, using some method I trust (maybe Wolfram Alpha). I’d do this for a handful of examples, trying to cover base and extreme cases. And then I’d do property testing if I really wanted good coverage. Further, this expect test library seems to just smoothen the experience of copying what the function returns into a test. This whole “expect test” business seems to rely on the developer looking at what the function returns for a given input, evaluating if it’s correct or not and then locking that in as “this is what this function is supposed to do”. That seems backwards and no different from how one implements functions in the first place, so I don’t know what is actually being tested. The entire point of testing is saying “this is what this function should do” and not “this is what the function did and thus that’s what it should always do”.
- angio 4y agoYou're supposed to use it as a repl, so you start with a test for `fib(1) = 1`, then `fib(2)` and so on. Once you're confident of your implementation, you use quickcheck to test general properties of the system. Similarly if you find a bug in the live system, you add a test for that and the initial output will be wrong. Then you fix your code until it prints the correct value and commit that so any regression will be caught.
- gleb 4y agoSimilar idea in Elixir, where the library itself handles the interactive bits: https://github.com/assert-value/assert_value_elixir https://github.com/assert-value/assert_value_elixir
- arcturus17 4y agoIs there anything like this in Python or C#? I have worked with OCaml extensively in coursework, but there’s no chance I’ll be using it in prod any time soon and I’d love toying with this approach in my working languages.
- wildcow 4y agoIn C# https://theramis.github.io/Snapper/#/pages/quickstart https://theramis.github.io/Snapper/#/pages/quickstart. I actually think there are other as well. But this is what i am trying to scratch to see if I can use it somehow.
- drothlis 4y agohttps://approvaltests.com/ https://approvaltests.com/
- anuragsoni 4y agoFor python there is https://github.com/ezyang/expecttest https://github.com/ezyang/expecttest which is modeled after the OCaml expect test library.
- theptip 4y agoI tend to think that tests should be carefully crafted for readability just like normal code. The “content of a REPL” is unlikely to be well-thought out enough to preserve meaningful invariants while remaining supple in the direction of likely changes. Perhaps in the hands of very good engineers this tool is net positive, but I shudder at giving junior engineers a tool that encourages less structure in tests. A good set of fixture/helper functions should let you write really short and expressive tests (or tabular parametrized tests, if you prefer) which seems to me to resolve most of the pain points the author is complaining about. One big advantage I do see with this approach is it seems to be a very compact rendering of a table of outputs; in Python+pytest+PyCharm if I run a 10-example parametrized test, I have to click through to see each failure individually. Perhaps there is a UX learning here that just rendering the raw errors into the code beside the test matrix could help visualize results faster. As an aside, I have recently been enjoying the “write an ascii representation as your test assert” mode of testing, it can give a different way of intuiting what is going on.
- adrianmonk 4y agoI think this would suffer from the same problem as partial self-driving cars: it's human nature for vigilance to falter if it doesn't feel like you're the sole/primary one in control. Of course, you can say "I won't let myself do that", but working against human nature is not a formula for success. If my back hurts, I can tell myself I'm just going to go lie down on the bed for 10 minutes but not take a nap, but then 30 minutes later I wake up feeling groggy.
- CJefferson 4y agoI work with a language where all test are expect tests ( GAP ). The biggest problem is you can basically never change how built in types are printed, as you'll break all tests in every program. For example, someone wanted to improve how plurals are printed, but that would break every test.
- BoppreH 4y agoSome years ago I wrote a Python function, "replace_me"[1], that edits the caller's source code. You can use it for code generation, inserting comments, generating fixed random seeds, etc. And one more use case I found was exactly what TFA describes, but even easier: import replace_me replace_me.test(1+1) Once executed, it evaluates the argument and becomes an assertion: import replace_me replace_me.test(1+1, 2) I never actually used it for anything important, but it comes back to my mind once in a while. [1]: https://github.com/boppreh/replace_me https://github.com/boppreh/replace_me