3 ms·
Thank you very much for the response! This paper is pretty interesting for sure. It's just it doesn't have what I would consider good evidence that coverage is
by GhotiFish 11y ago
Thank you very much for the response! This paper is pretty interesting for sure.
It's just it doesn't have what I would consider good evidence that coverage isn't correlated with effectiveness. (or... I don't understand the section that explains it <:( )
What I would consider good evidence would be to demonstrate that different test suites of the same size and different coverage achieve the same effectiveness.
* test suite A has a size of 300 sloc, and a coverage of 5%, effectiveness of ~8%
* test suite B has a size of 300 sloc, and a coverage of 12%, effectiveness of ~8%
That to me would be evidence of the stated conclusion, but I don't see where this is demonstrated (or where this is demonstrated, I don't understand). I do see where it is stated! Though... now that I have read through the paper a bit more thoroughly.
On the subject of normalized effectiveness.
Suppose
we are comparing suite A, with 50% coverage, to suite B, with
60% coverage. Suite B will almost certainly have a higher
raw effectiveness measurement, since it covers more code and
will therefore almost certainly kill more mutants. However,
if suite A kills 80% of the mutants that it covers, while suite
B kills only 70% of the mutants that it covers, suite A is
in some sense a better suite."
I don't believe a majority of people would agree with this. To me, this says coverage is positively correlated with effectiveness, and that suite B is doing more with less. Maybe that's a philosophical stand point? By normalizing, have you eliminated the premise?
Anyway, thank you for your time and work!
- lmmi 11y agoWhat you're looking for is in Figure 3 (admittedly a bit hard to read because I had to squish it into the paper). Each panel in that figure shows the results for test suites of a fixed size. For example, the top left panel shows the results for suites with three test cases for Apache POI. If you draw a horizontal line through the graph, all of the test suites that fall on that line have the same effectiveness score even though they have different coverage levels. I gave a talk about this paper at GTAC this year and Google posts all the videos on YouTube, so if you have time and you're still curious the explanation in the talk might help (and the graphs are much easier to read). https://www.youtube.com/watch?v=sAfROROGujU https://www.youtube.com/watch?v=sAfROROGujU The normalization was a point of contention with the peer reviewers as well, so in the end I tried both the normalized and unnormalized metrics and found similar results with both. The other tables and figures are available on my site if you want to look at them. I'm not sure I understand what you mean when you say suite B is doing more with less, though. In the example, I was trying to say that suite B covers more code, so it will kill more mutants. Maybe suite A kills 20 mutants and suite B kills 25 mutants, just to have some numbers to talk about. But if B covers 50 mutants, and is only killing 25 of them, while A covers 25 mutants and kills 20 of the 25, it seems like suite A is doing a better job of testing the code it covers. Or to put it another way, suite B is broad but shallow while suite A is focused but deep. B isn't necessarily a bad suite, but I wouldn't say it's doing more with less, just that it has a different focus. Maybe I'm misunderstanding your point, though. Another way of thinking about it is that the raw mutation score measures breadth: B is better than A because 25 > 20. The normalized score measures depth: A is better than B because 80% (20/25) > 50% (25/50).