Lesson 2 / 25
What Can Go Wrong
Plausible patches are not correct patches.
Failure modes of automated patches
An automated fixer is optimising to make a check pass, which invites shortcuts: deleting or weakening tests, special-casing the exact test input, catching and hiding exceptions, broad refactors that slip in unrelated changes, editing CI configuration, adding dependencies, leaking secrets into code, or "fixing" a flaky test that was never a code bug. Agents that read issues and logs can also be steered by prompt injection in that text. Each risk needs a specific control, not just a careful prompt.
Risks and their controls
Map each risk to a mechanical check.
risk control
tests deleted / skipped / weakened test-tampering detector, protected test paths
special-casing the test input extra hidden tests, mutation checks, review
unrelated changes diff size limits, file allowlists
CI / infra / lockfile edits protected paths
secrets in code secret scanning on every patch
flaky test "fixed" with code changes rerun-based flakiness detection
prompt injection via issues / logs treat text as data, least-privilege tokens, sandboxAssume the optimiser will find shortcuts
If a check can be passed by cheating, eventually a patch will cheat; close the loophole in the pipeline.
Quick check: Why might an automated fixer delete a failing test?
- Deleting the test also makes the check pass, unless the pipeline forbids it
- Tests are always wrong
- Models cannot read tests
- Deleting tests improves coverage
Answer
Deleting the test also makes the check pass, unless the pipeline forbids it — Guard against gaming the check.