6 ms·
I have a few objections to the evidence presented. Keeping in mind that I don't disagree with the statement. Test coverage is an objective metric, and test Eff
by GhotiFish 11y ago
I have a few objections to the evidence presented.
Keeping in mind that I don't disagree with the statement. Test coverage is an objective metric, and test Effectiveness is a... what is it again? How many bugs you'll find with it? Obviously the two are separate concepts.
This is my first objection: The paper seems to say that Mutation Testing is test effectiveness, but Mutation testing is merely another metric. They cite other papers, but papers that attempt to demonstrate that this metric is correlated with test "effectiveness".
Metric against metric, is that meaningful?
She presents graphs of results that seem to demonstrate linear correlation between test suite size/test suite coverage/and mutation testing (called "effectiveness") This is addressed later, "isn't this what we would expect?" yah! And the explanation for why it's unexpected sailed fully over my head. (I admit it! I'm dumb.)
finally, many test suites are generated on the presumption of achieving code coverage, is this test valid without also having test suites that were made without that goal in mind? Could such a suite exist?
so summary of my objections
* is mutation testing a meaningful measure of effectiveness?
* can you measure one metric against another another, get a linear relationship, and conclude any meaningful differences?
* does the presence of code coverage as a target spoil the conclusion?
I'd love to hear input on this.
- lmmi 11y agoPaper author here. We measured "effectiveness" as the mutation score. We showed in a separate paper that the mutation score is a good way to measure a test suite's ability to detect real faults. Of course, using the mutation score instead of the real fault detection rate does add a layer of indirection, but it's a lot more practical for a large empirical study (i.e., can be automated), so there's a bit of a tradeoff there. (The other paper is also on my website if you want to read it: http://www.linozemtseva.com/research/2014/fse/mutant_validity/ http://www.linozemtseva.com/research/2014/fse/mutant_validit...) What the graphs show is that effectiveness rises with suite size, which is expected: more tests catch more bugs. We also see that effectiveness rises with coverage, which seems intuitive: you can't catch bugs in code you never run. But when you graph effectiveness against coverage for suites that are all the same size, the correlation drops significantly, in some cases to 0. Here's an analogy: the number of PhD students that graduate in the US is highly correlated with the amount of profit generated by arcades in the US, but there's no causal relationship between them. They probably both depend on the size of the population. Similarly, coverage and effectiveness are correlated because they both depend on the size of the suite, but we can't say that there's a causal relationship. In other words, saying that a suite will catch a lot of bugs because it has high coverage is like saying that a lot of PhD students will graduate because arcades turned a good profit this year. Your third point is a really good question. It's true that developers don't write tests randomly, so our method of making new suites by picking random test cases isn't quite realistic. What impact that would have on the results, I'm not sure. That's something I'd like to look into in the future.
- GhotiFish 11y agoThank you very much for the response! This paper is pretty interesting for sure. It's just it doesn't have what I would consider good evidence that coverage isn't correlated with effectiveness. (or... I don't understand the section that explains it <:( ) What I would consider good evidence would be to demonstrate that different test suites of the same size and different coverage achieve the same effectiveness. * test suite A has a size of 300 sloc, and a coverage of 5%, effectiveness of ~8% * test suite B has a size of 300 sloc, and a coverage of 12%, effectiveness of ~8% That to me would be evidence of the stated conclusion, but I don't see where this is demonstrated (or where this is demonstrated, I don't understand). I do see where it is stated! Though... now that I have read through the paper a bit more thoroughly. On the subject of normalized effectiveness. Suppose we are comparing suite A, with 50% coverage, to suite B, with 60% coverage. Suite B will almost certainly have a higher raw effectiveness measurement, since it covers more code and will therefore almost certainly kill more mutants. However, if suite A kills 80% of the mutants that it covers, while suite B kills only 70% of the mutants that it covers, suite A is in some sense a better suite." I don't believe a majority of people would agree with this. To me, this says coverage is positively correlated with effectiveness, and that suite B is doing more with less. Maybe that's a philosophical stand point? By normalizing, have you eliminated the premise? Anyway, thank you for your time and work!
- lmmi 11y agoWhat you're looking for is in Figure 3 (admittedly a bit hard to read because I had to squish it into the paper). Each panel in that figure shows the results for test suites of a fixed size. For example, the top left panel shows the results for suites with three test cases for Apache POI. If you draw a horizontal line through the graph, all of the test suites that fall on that line have the same effectiveness score even though they have different coverage levels. I gave a talk about this paper at GTAC this year and Google posts all the videos on YouTube, so if you have time and you're still curious the explanation in the talk might help (and the graphs are much easier to read). https://www.youtube.com/watch?v=sAfROROGujU https://www.youtube.com/watch?v=sAfROROGujU The normalization was a point of contention with the peer reviewers as well, so in the end I tried both the normalized and unnormalized metrics and found similar results with both. The other tables and figures are available on my site if you want to look at them. I'm not sure I understand what you mean when you say suite B is doing more with less, though. In the example, I was trying to say that suite B covers more code, so it will kill more mutants. Maybe suite A kills 20 mutants and suite B kills 25 mutants, just to have some numbers to talk about. But if B covers 50 mutants, and is only killing 25 of them, while A covers 25 mutants and kills 20 of the 25, it seems like suite A is doing a better job of testing the code it covers. Or to put it another way, suite B is broad but shallow while suite A is focused but deep. B isn't necessarily a bad suite, but I wouldn't say it's doing more with less, just that it has a different focus. Maybe I'm misunderstanding your point, though. Another way of thinking about it is that the raw mutation score measures breadth: B is better than A because 25 > 20. The normalized score measures depth: A is better than B because 80% (20/25) > 50% (25/50).