Lesson 12 / 25
Are the Tests Strong Enough? Mutation Checks
Tests that cannot catch a wrong fix cannot verify a right one.
Mutate the code, see if tests notice
A patch "passing the tests" means little if the tests are weak. Mutation testing makes small deliberate changes (mutants) to the code, such as flipping a comparison or changing a constant, and checks whether the tests fail. Surviving mutants reveal untested behaviour, often exactly the boundary a fix touched. Tools such as mutmut or cosmic-ray (Python), Stryker (JavaScript) and PIT (Java) automate it; running it on the changed lines keeps it affordable.
Weak versus strong tests against four mutants, run
I ran this with Python 3 (standard library) and, where it uses git, real git in a throwaway temporary repository. Candidate patches are written by hand to stand in for model output. With only one test (a large order) the suite kills 2 of 4 mutants; changing >= to > or 50 to 49 survives, so a wrong boundary fix would still pass. Adding boundary tests at 50, 49 and 30 kills all 4.
def discount(total, code):
if code == "FLAT50" and total >= 50:
return total - 50
return total
mutants = {
">= -> >": lambda t, c: t - 50 if c == "FLAT50" and t > 50 else t,
">= -> <=": lambda t, c: t - 50 if c == "FLAT50" and t <= 50 else t,
"50 -> 49": lambda t, c: t - 50 if c == "FLAT50" and t >= 49 else t,
"- -> +": lambda t, c: t + 50 if c == "FLAT50" and t >= 50 else t,
}
weak_tests = [((120, "FLAT50"), 70)]
strong_tests = weak_tests + [((50, "FLAT50"), 0), ((49, "FLAT50"), 49), ((30, "FLAT50"), 30)]
for name, tests in [("weak suite", weak_tests), ("strong suite", strong_tests)]:
assert all(discount(*a) == e for a, e in tests)
killed = [m for m, f in mutants.items() if any(f(*a) != e for a, e in tests)]
print(f"{name:<12} kills {len(killed)}/{len(mutants)} mutants; survivors: {[m for m in mutants if m not in killed]}")
Output:
weak suite kills 2/4 mutants; survivors: ['>= -> >', '50 -> 49'] strong suite kills 4/4 mutants; survivors: []
Require boundary tests with fixes
For any comparison or limit change, require tests just below, at and above the boundary.
Quick check: What does a surviving mutant tell you?
- The code is perfect
- The tests did not detect a change in behaviour, so they are too weak there
- The mutant is the correct fix
- Tests should be deleted
Answer
The tests did not detect a change in behaviour, so they are too weak there — Strengthen tests where mutants survive.