6 ms·
Paper author here. Scoring coverage for every pass is an interesting idea and something I'd like to look into. I'm not sure it would change the result as much
by lmmi 11y ago
Paper author here. Scoring coverage for every pass is an interesting idea and something I'd like to look into. I'm not sure it would change the result as much as you think, though. The basic finding of the paper is that coverage is a complicated way of measuring the size of the suite. Counting the number of times each line is hit will have the same problem, I think: writing more tests increases that score but also increases the number of bugs found, causing a spurious correlation.
My hunch is that the quality of the oracle matters more than the coverage score. You can write tests that cover all the code without actually checking anything; the tests will only catch bugs if you're carefully comparing expected and actual results. Maybe a simple metric like "number of asserts" would be useful -- except, of course, that will also be correlated with the size of the suite... It's a tough problem.
The point about the title is fair. I erred on the side of clickbait when I wrote the paper and regret it a bit. On the other hand, it worked. :)
- jacques_chester 11y agoI see your point about coverage volume being a better-fitted proxy for suite size. It'd also be a proxy for path coverage. Still, it'd be fun to count the horse's teeth anyhow. Would you say that coverage's worth as a negative metric still seems meaningful, at least as a heuristic? I imagine that's covered in other literature. Where I work we don't really fuss too much about coverage. We TDD, so in practice our coverage hovers around the high 90s as a matter of course. When I am writing a test I often manually mutate the code and test once it goes green as a quick validation that the test does what I think it does. One last question -- did you classify tests? Feature, integration and unit tests should show quite different curves. Especially heavily mockist style unit tests.
- lmmi 11y agoI'd definitely say that the absence of coverage is a problem. My view is that coverage is necessary but not sufficient for good testing. We talked a bit about classifying tests but didn't do it in the end because it's surprisingly hard to do. I do know of one paper that looked at different kinds of tests, called "The Effect of Code Coverage on Fault Detection under Different Testing Profiles": http://goo.gl/nnxgwE http://goo.gl/nnxgwE. The authors found differences between tests for error cases vs. tests for normal operation and between functional tests vs. random tests. IIRC, they had undergrads do a term project that had to pass 1200 tests before the final submission, and the professors themselves wrote the tests, so categorization was a bit easier.
- jacques_chester 11y agoAs a rough way to automatically classify tests, you can look for tools like capybara, selenium or htmlunit for feature and mocking libraries for unit. Mind you, there's as many taxonomies for tests as there are tests. To be honest I expect one way to classify them is by working backwards from coverage -- feature tests should have low volume but wide distribution across a codebase (perhaps that's another metric -- density?). Unit tests would be narrow but deep on a particular module.
- lmmi 11y agoThose are good ideas, thanks!
- jacques_chester 11y agoThrow me on as the tenth or eleventh coauthor and I'll buy beer to sweeten the deal. If you need a corpus of code from a highly doctrinaire TDD shop, or if you think we can help, let me know: jchester@pivotal.io.