5 ms·
What you're looking for is in Figure 3 (admittedly a bit hard to read because I had to squish it into the paper). Each panel in that figure shows the results f
by lmmi 11y ago
What you're looking for is in Figure 3 (admittedly a bit hard to read because I had to squish it into the paper). Each panel in that figure shows the results for test suites of a fixed size. For example, the top left panel shows the results for suites with three test cases for Apache POI. If you draw a horizontal line through the graph, all of the test suites that fall on that line have the same effectiveness score even though they have different coverage levels. I gave a talk about this paper at GTAC this year and Google posts all the videos on YouTube, so if you have time and you're still curious the explanation in the talk might help (and the graphs are much easier to read). https://www.youtube.com/watch?v=sAfROROGujU https://www.youtube.com/watch?v=sAfROROGujU
The normalization was a point of contention with the peer reviewers as well, so in the end I tried both the normalized and unnormalized metrics and found similar results with both. The other tables and figures are available on my site if you want to look at them.
I'm not sure I understand what you mean when you say suite B is doing more with less, though. In the example, I was trying to say that suite B covers more code, so it will kill more mutants. Maybe suite A kills 20 mutants and suite B kills 25 mutants, just to have some numbers to talk about. But if B covers 50 mutants, and is only killing 25 of them, while A covers 25 mutants and kills 20 of the 25, it seems like suite A is doing a better job of testing the code it covers. Or to put it another way, suite B is broad but shallow while suite A is focused but deep. B isn't necessarily a bad suite, but I wouldn't say it's doing more with less, just that it has a different focus. Maybe I'm misunderstanding your point, though.
Another way of thinking about it is that the raw mutation score measures breadth: B is better than A because 25 > 20. The normalized score measures depth: A is better than B because 80% (20/25) > 50% (25/50).