# Measuring an Auto-Fix System — Safe Autonomous Code Fixing

Source: https://www.skillbyai.com/en/safe-autonomous-code-fixing/d-metrics

> Trust is earned with numbers.

## Merge rate, revert rate, review time

Track per task type and repository: **attempts**, how many produced a patch, how many were **blocked by gates** (and which gate), **merged** after review, **reverted** after merge, review time, and time from issue to fix. A high merge rate with a high revert rate means the gates or tests are too weak; many gate blocks may mean the fixer needs better context. Use these numbers to widen or narrow the scope of automation.

## Summarising an attempt log, run

I ran this with Python 3 (standard library) and, where it uses git, real git in a throwaway temporary repository. Candidate patches are written by hand to stand in for model output. Across eight example attempts, four were merged (50%), one of those was later reverted (25% revert rate), two were blocked by automated gates, and human review took 7.2 minutes on average.

```python
log = [  # one row per automated fix attempt
    {"result": "merged", "reverted": False, "review_min": 6}, {"result": "merged", "reverted": True, "review_min": 4},
    {"result": "rejected_gate", "reverted": False, "review_min": 0}, {"result": "merged", "reverted": False, "review_min": 9},
    {"result": "rejected_review", "reverted": False, "review_min": 12}, {"result": "no_patch", "reverted": False, "review_min": 0},
    {"result": "merged", "reverted": False, "review_min": 5}, {"result": "rejected_gate", "reverted": False, "review_min": 0},
]
n = len(log); merged = [r for r in log if r["result"] == "merged"]
print(f"attempts {n} | merged {len(merged)} ({len(merged) / n:.0%})")
print(f"revert rate of merged fixes: {sum(r['reverted'] for r in merged) / len(merged):.0%}")
print(f"blocked by automated gates: {sum(r['result'] == 'rejected_gate' for r in log)}")
reviewed = [r for r in log if r["review_min"]]
print(f"average human review time: {sum(r['review_min'] for r in reviewed) / len(reviewed):.1f} min")
```

Output:

```
attempts 8 | merged 4 (50%)
revert rate of merged fixes: 25%
blocked by automated gates: 2
average human review time: 7.2 min
```

## Review reverted fixes together

Each revert is a lesson: add a test, gate or scope rule so the same class of mistake is blocked next time.

**Quiz:** What does a high revert rate on merged bot fixes suggest?

- [ ] Reviews are too slow
- [ ] The bot is perfect
- [x] Gates or tests are not catching bad fixes
- [ ] There are too few bugs

*Answer:* Gates or tests are not catching bad fixes. Strengthen verification before widening scope.
