4 ms·
One random thing I've been impressed with (that I know is public) is mutant testing: https://testing.googleblog.com/2021/04/mutation-testing.html https://testin
by rwiggins 4y ago
One random thing I've been impressed with (that I know is public) is mutant testing: https://testing.googleblog.com/2021/04/mutation-testing.html https://testing.googleblog.com/2021/04/mutation-testing.html , https://research.google/pubs/pub46584/ https://research.google/pubs/pub46584/
To pull an example from the paper, if your line of code says
if (a == b || b == 1)
you might get a comment that says something like
> "Changing this line to
if (a != b || b == 1)
> does not cause any test to fail."
Page 4 lists other mutations, like replacing logical ANDs/ORs with just `true` or `false`, switching arithmetic plus to minus, etc.
This isn't a larger system design context thing, though, just a testing one.
Disclosure: I work at Google.
- technion 4y agoThere's a pretty good Ruby gem I've used for this before: https://github.com/mbj/mutant https://github.com/mbj/mutant
- YZF 4y agoThis is for test coverage? I get it technically but not clear to me if this adds value. Covering all the internal states/code paths in tests seems very hard and potentially of diminishing returns? Isn't time better invested in other things? I'm also guessing this ties into Google's tooling of being able to tell which line of code is hit by which tests and running those as part of the mutation test? EDIT: The blog has some discussion about the correlation between some of these mutation failures and real bugs. I remain a little suspicious to be honest. Also "I also looked into the developer behavior changes after using mutation testing on a project for longer periods of time, and discovered that projects that use mutation testing get more tests over time, as developers get exposed to more and more mutants. Not only do developers write more test cases, but those test cases are more effective in killing mutants: less and less mutants get reported over time. I noticed this from personal experience too: when writing unit tests, I would see where I cut some corners in the tests, and anticipated the mutant. Now I just add the missing test cases, rather than facing a mutant in my Code review, and I rarely see mutants these days, as I’ve learned to anticipate and preempt them." Now this one is expected, if you build automated tooling that comments on your code review you expect people to try and avoid that. However there's still the question of what's the quality improvement and what's the productivity impact.
- gravypod 4y agoProduction impact: if you can replace: if (a == 10) With if (true) Then you're not testing the behavior of your application. If in production `a != 10` and you've never tested that it might be a problem. Mutation testing can also do stuff like removing entire statements, etc.
- YZF 4y agoThe question (well, one question) is why would we suspect the lack of a test indicates a bug? For most non-trivial software the possible state-space is enormous and we generally don't/can't test all of it. So "not testing the (full) behaviour of your application is the default for any test strategy", if we could we wouldn't have bugs... Last I checked most software (including Google's) has plenty of bugs. The next question would be let's say I spend my time writing the tests to resolve this (could be a lot of work) is that time better spent vs. other things I could be doing? (i.e. what's the ROI) Even ignoring that is there data to support that the quality of software where mutation testing was added improved measurably (e.g. less bugs files against the deployed product, better uptime, etc?) Is this method better than just looking at code coverage? Possibly none of the tests enter the if statement at all? EDIT: where I'm coming from is that it's not a given this is an improvement to the software development process. It seems like there was some experimentation around the validity of this method which is good but like a lot of software studies somewhat limited. It also seems there's a lot of heuristics based on user feedback, which is also good I guess, but presumably also somewhat biased. The related paper has a lot of details including: "Since the ultimate goal is not just to write tests for mutants, but to prevent real bugs, we investigated a dataset of high-priority bugs and analyzed mutants before and after the fix with an experimental rig of our mutation testing system. In 70% of cases, a bug is coupled with a mutant that, had it been reported during code review, could have prevented the introduction of that bug." Which should imply(???) that 70% of "high priority" bugs can be eliminated during the review process by using this sort of mutation testing. Seeing data to that effect would be cool (i.e. after the fact) and if it's real that'd be pretty incredible and we should all be doing that.
- 4y ago