SkillByAIOpen interactive version →

Lesson 10 / 25

Isolated Verification of Each Candidate

Fresh copy, full suite, clear verdict.

Every attempt starts clean

Verify each candidate in a fresh, isolated copy of the repository (a temporary directory, git worktree or container), so one attempt cannot affect another or the developer's workspace. Apply the patch, run the reproduction test and the full test suite, plus linters and type checks, and record the verdict with logs. Reject patches that touch tests or protected files before even running them. Only candidates that pass every gate move on.

Trust evidence, not plausibility

Every candidate is verified in isolation with tests that are strong enough to catch wrong fixes.

Figure 4.1 — Isolation, flakiness, mutation and static checks.

Three candidate patches through the gates, run

I ran this with Python 3 (standard library) and, where it uses git, real git in a throwaway temporary repository. Candidate patches are written by hand to stand in for model output. Each candidate is applied to its own temporary copy. A: deleting the boundary test is rejected before running because it edits tests. B: removing the total check makes the small-order test fail. C: changing > 50 to >= 50 passes all three tests and is accepted.

import os, shutil, subprocess, sys, tempfile, textwrap
base = tempfile.mkdtemp()
open(os.path.join(base, "pricing.py"), "w").write(textwrap.dedent("""
    def discount(total, code):
        if code == "FLAT50" and total > 50:
            return total - 50
        return total
"""))
os.makedirs(os.path.join(base, "tests"))
open(os.path.join(base, "tests", "test_pricing.py"), "w").write(textwrap.dedent("""
    import unittest, sys, os
    sys.path.insert(0, os.path.dirname(os.path.dirname(__file__)))
    from pricing import discount
    class T(unittest.TestCase):
        def test_boundary(self): self.assertEqual(discount(50, "FLAT50"), 0)
        def test_large(self): self.assertEqual(discount(120, "FLAT50"), 70)
        def test_small(self): self.assertEqual(discount(30, "FLAT50"), 30)
"""))
candidates = {   # what a model might propose; each is (file, old, new)
    "A: delete the boundary test": ("tests/test_pricing.py", 'def test_boundary(self): self.assertEqual(discount(50, "FLAT50"), 0)', "pass"),
    "B: drop the total check": ("pricing.py", 'code == "FLAT50" and total > 50', 'code == "FLAT50"'),
    "C: use >= 50": ("pricing.py", "total > 50", "total >= 50"),
}
def evaluate(name, change):
    work = tempfile.mkdtemp(); shutil.copytree(base, work, dirs_exist_ok=True)   # isolated copy per attempt
    path, old, new = change
    if path.startswith("tests/"):
        return "rejected: patch edits tests"
    p = os.path.join(work, path); text = open(p).read()
    open(p, "w").write(text.replace(old, new))
    r = subprocess.run([sys.executable, "-m", "unittest", "discover", "-s", "tests", "-q"], cwd=work, capture_output=True, text=True)
    return "ACCEPT: all tests pass" if r.returncode == 0 else "rejected: " + r.stderr.strip().splitlines()[-1]
for name, change in candidates.items():
    print(f"{name:<28} {evaluate(name, change)}")

Output:

A: delete the boundary test  rejected: patch edits tests
B: drop the total check      rejected: FAILED (failures=1)
C: use >= 50                 ACCEPT: all tests pass

Run the full suite, not just the new test

A fix that passes its own test but breaks another feature is a regression; candidate B shows how.

Quick check: Why verify each candidate in a fresh copy?

  • So attempts cannot interfere with each other or with the developer's workspace
  • It makes tests pass automatically
  • Git requires it
  • It removes the need for tests
Answer

So attempts cannot interfere with each other or with the developer's workspace — Isolation keeps verdicts trustworthy.