# Are the Tests Strong Enough? Mutation Checks — Safe Autonomous Code Fixing

Source: https://www.skillbyai.com/en/safe-autonomous-code-fixing/v-mutation

> Tests that cannot catch a wrong fix cannot verify a right one.

## Mutate the code, see if tests notice

A patch "passing the tests" means little if the tests are weak. **Mutation testing** makes small deliberate changes (mutants) to the code, such as flipping a comparison or changing a constant, and checks whether the tests fail. Surviving mutants reveal untested behaviour, often exactly the boundary a fix touched. Tools such as mutmut or cosmic-ray (Python), Stryker (JavaScript) and PIT (Java) automate it; running it on the changed lines keeps it affordable.

## Weak versus strong tests against four mutants, run

I ran this with Python 3 (standard library) and, where it uses git, real git in a throwaway temporary repository. Candidate patches are written by hand to stand in for model output. With only one test (a large order) the suite kills 2 of 4 mutants; changing >= to > or 50 to 49 survives, so a wrong boundary fix would still pass. Adding boundary tests at 50, 49 and 30 kills all 4.

```python
def discount(total, code):
    if code == "FLAT50" and total >= 50:
        return total - 50
    return total
mutants = {
    ">= -> >":   lambda t, c: t - 50 if c == "FLAT50" and t > 50 else t,
    ">= -> <=":  lambda t, c: t - 50 if c == "FLAT50" and t <= 50 else t,
    "50 -> 49":  lambda t, c: t - 50 if c == "FLAT50" and t >= 49 else t,
    "- -> +":    lambda t, c: t + 50 if c == "FLAT50" and t >= 50 else t,
}
weak_tests = [((120, "FLAT50"), 70)]
strong_tests = weak_tests + [((50, "FLAT50"), 0), ((49, "FLAT50"), 49), ((30, "FLAT50"), 30)]
for name, tests in [("weak suite", weak_tests), ("strong suite", strong_tests)]:
    assert all(discount(*a) == e for a, e in tests)
    killed = [m for m, f in mutants.items() if any(f(*a) != e for a, e in tests)]
    print(f"{name:<12} kills {len(killed)}/{len(mutants)} mutants; survivors: {[m for m in mutants if m not in killed]}")
```

Output:

```
weak suite   kills 2/4 mutants; survivors: ['>= -> >', '50 -> 49']
strong suite kills 4/4 mutants; survivors: []
```

## Require boundary tests with fixes

For any comparison or limit change, require tests just below, at and above the boundary.

**Quiz:** What does a surviving mutant tell you?

- [ ] The code is perfect
- [x] The tests did not detect a change in behaviour, so they are too weak there
- [ ] The mutant is the correct fix
- [ ] Tests should be deleted

*Answer:* The tests did not detect a change in behaviour, so they are too weak there. Strengthen tests where mutants survive.
