4 ms·
I am not surprised by the findings of the study - not at all. In discussions about software metrics, I'm always trying to make the point that you can only use
by struppi 11y ago
I am not surprised by the findings of the study - not at all.
In discussions about software metrics, I'm always trying to make the point that you can only use most metrics from within a team to gain insights, not as an external measure about how good/bad the code (or the team) is. In other words, if a team thinks they have a problem, they can use metrics to gain insights and explore possible solutions. But as soon as someone says "[foo metric] needs to be at least [value]", you have already lost - the cheating and gaming begins. Even if the agreement on [value] comes from within the team.
Back to the topic :) I am not surprised by the findings - Higher test coverage does not mean that everything is fine. But very low test coverage indicates that there might be hidden problems here or there. This is how I like to use test coverage and how I try to teach it.
But it is great that we have now empirical data here: From now on, I can point others to this study when we are discussing whether the build server should reject commits with less than [value] coverage.
- jacques_chester 11y agoA brief skim of the study shows that the headline is a bit misleading. The correlation between coverage and bug yield is weak when suite size is controlled for. Another way of looking at this is that large suites improve bug yield. They tend to have high coverage as a by-product, because a large test suite will hammer each unit from various directions. My suspicion is that scoring coverage for every pass over an LOC or a pathway would change the result substantially. Nothing in this paper is an argument that test coverage is useless; rather, it should be seen as a secondary metric that rapidly loses indicative value as it approaches 100%. Number of test cases seems to be the independent variable here.
- marcosdumay 11y agoIn other words, writing another test for the hot section of your code seems to improve quality on average about as much as writing a test for some rarely run code that you didn't bother to test until now. If anything, I got surprised coverage is that important. But as the article (and you) says, it's only important on lower values, so not that surprising.
- jacques_chester 11y ago> In other words, writing another test for the hot section of your code seems to improve quality on average about as much as writing a test for some rarely run code that you didn't bother to test until now. I'm not sure. I've given it some more thought. I think the problem with distinguishing between the two is that "coverage" is defined, so to speak, as an area. Once an LOC is executed by any test, its value ticks over once. It never gets counted again. If instead each LOC execution was summed up independently, each LOC gets a "height" -- the number of times it was exercised in a suite. Then the overall suite gets test coverage volume, rather than test coverage area. Once you track per-LOC volumes, you could start trying to tease out the differences between introducing a test at all and adding another test to critical code. And, I suspect, the correlation between coverage volume and bug yield would be much more robust than area when controlled for number of cases. Because area when controlling for number of cases is ... a rough proxy for test coverage volume. Edit: though with an interesting difference. Coverage area converges to 100%, but coverage volume is effectively unbounded. This would hopefully upset the abuse of coverage as a management metric instead of as a code smell.
- marcosdumay 11y agoMy thoughts go that way too, but I expected some volumes to be more important than others. And because of it, I expected the article to measure a negative correlation. Hence my surprise. But reading the paper, the study is about random samples of the complete test suite, and the results are far too weak to get any conclusion about it, finding negative correlations some times, and weak positive other times. Anyway, they got a strong correlation, with good p-value (not very worthy on this context, but it's all we get) for your hypothesis about coverage volume. It's just my hypothesis about some code being more important than other that is inconclusive.
- jacques_chester 11y ago> It's just my hypothesis about some code being more important than other that is inconclusive. There's a lot of literature going back a long way showing that defects tend to cluster. A leading indicator might be defects found in a module divided by coverage volume. On the other hand, defects tend to cluster in the hard stuff, not at random.
- lmmi 11y agoPaper author here. Scoring coverage for every pass is an interesting idea and something I'd like to look into. I'm not sure it would change the result as much as you think, though. The basic finding of the paper is that coverage is a complicated way of measuring the size of the suite. Counting the number of times each line is hit will have the same problem, I think: writing more tests increases that score but also increases the number of bugs found, causing a spurious correlation. My hunch is that the quality of the oracle matters more than the coverage score. You can write tests that cover all the code without actually checking anything; the tests will only catch bugs if you're carefully comparing expected and actual results. Maybe a simple metric like "number of asserts" would be useful -- except, of course, that will also be correlated with the size of the suite... It's a tough problem. The point about the title is fair. I erred on the side of clickbait when I wrote the paper and regret it a bit. On the other hand, it worked. :)
- jacques_chester 11y agoI see your point about coverage volume being a better-fitted proxy for suite size. It'd also be a proxy for path coverage. Still, it'd be fun to count the horse's teeth anyhow. Would you say that coverage's worth as a negative metric still seems meaningful, at least as a heuristic? I imagine that's covered in other literature. Where I work we don't really fuss too much about coverage. We TDD, so in practice our coverage hovers around the high 90s as a matter of course. When I am writing a test I often manually mutate the code and test once it goes green as a quick validation that the test does what I think it does. One last question -- did you classify tests? Feature, integration and unit tests should show quite different curves. Especially heavily mockist style unit tests.
- lmmi 11y agoI'd definitely say that the absence of coverage is a problem. My view is that coverage is necessary but not sufficient for good testing. We talked a bit about classifying tests but didn't do it in the end because it's surprisingly hard to do. I do know of one paper that looked at different kinds of tests, called "The Effect of Code Coverage on Fault Detection under Different Testing Profiles": http://goo.gl/nnxgwE http://goo.gl/nnxgwE. The authors found differences between tests for error cases vs. tests for normal operation and between functional tests vs. random tests. IIRC, they had undergrads do a term project that had to pass 1200 tests before the final submission, and the professors themselves wrote the tests, so categorization was a bit easier.
- ZeroGravitas 11y agohttps://en.wikipedia.org/wiki/Goodhart%27s_law https://en.wikipedia.org/wiki/Goodhart%27s_law
- weavie 11y agoI would imagine coverage is a more important metric for dynamic languages than static. Just running through code paths to ensure it doesn't crash due to misspellings etc.. is not really something a compiled language would need as the compiler pretty much handles that.
- icebraining 11y agoA linter catches most of those in dynamic languages; many already do cross-file type inference.