Lesson 4 / 25
Reproducing With a Failing Test
Red before green.
Turn the report into a test
Convert the bug report into an automated test that fails for the right reason before any fix: the boundary order, the bad input, the wrong output. This test becomes the contract for the fix and protects against regressions later. Keep it minimal and deterministic (fixed inputs, no network, no time dependence). If the bug cannot be reproduced, gather more information rather than letting an agent guess.
Prove the bug, then find it
A failing test proves the bug exists; traces and bisect point to where it lives.
From failing test to passing fix, run
I ran this with Python 3 (standard library) and, where it uses git, real git in a throwaway temporary repository. Candidate patches are written by hand to stand in for model output. A discount function is meant to apply a flat 50 off orders of 50 or more. The new boundary test fails against the original code (> 50) and passes after changing it to >= 50; the existing test keeps passing.
import os, subprocess, sys, tempfile, textwrap
d = tempfile.mkdtemp()
open(os.path.join(d, "pricing.py"), "w").write(textwrap.dedent("""
def discount(total, code):
if code == "SAVE10":
return total * 0.9
if code == "FLAT50" and total > 50: # bug report: a 50.00 order should qualify
return total - 50
return total
"""))
open(os.path.join(d, "test_pricing.py"), "w").write(textwrap.dedent("""
import unittest
from pricing import discount
class T(unittest.TestCase):
def test_save10(self): self.assertAlmostEqual(discount(100, "SAVE10"), 90)
def test_flat50_boundary(self): self.assertEqual(discount(50, "FLAT50"), 0) # reproduces the report
"""))
def run():
r = subprocess.run([sys.executable, "-m", "unittest", "-q"], cwd=d, capture_output=True, text=True)
return r.returncode, r.stderr.strip().splitlines()[-1]
print("before fix:", run())
src = open(os.path.join(d, "pricing.py")).read().replace("total > 50", "total >= 50")
open(os.path.join(d, "pricing.py"), "w").write(src)
print("after fix :", run())
Output:
before fix: (1, 'FAILED (failures=1)') after fix : (0, 'OK')
Commit the test first
Commit the failing test separately from the fix so reviewers can see it fail on the original code.
Quick check: Why must the new test fail before the fix?
- Tests should always fail
- It proves the test actually captures the bug
- It makes CI faster
- Passing tests cannot be committed
Answer
It proves the test actually captures the bug — A test that never failed proves nothing.