# Regression Gates in CI — Prompt Quality Scoring

Source: https://www.skillbyai.com/en/prompt-quality-scoring/t-gate

> Improvements must not break critical cases.

## Case-level comparison

A new prompt can raise the average while breaking important cases. Compare at the **case level**: list newly passing and newly failing cases, and mark **critical** cases (safety, legal wording, key customers, past incidents) that must never regress. The CI gate blocks the change if any critical case fails, or if overall scores drop beyond a tolerance, and prints the details for review. Run it on every prompt change, just like unit tests.

## A case-level regression gate, run

I ran this with Python 3 (scipy 1.18.1 where imported) on example or seeded simulated data, not results from a real product. The candidate keeps the pass count at 8 of 10, fixing t03 and t07 but breaking t04 and t09. Both broken cases are marked critical, so the gate blocks the change even though the average did not drop.

```python
baseline = {"t01": 1, "t02": 1, "t03": 0, "t04": 1, "t05": 1, "t06": 1, "t07": 0, "t08": 1, "t09": 1, "t10": 1}
candidate = {"t01": 1, "t02": 1, "t03": 1, "t04": 0, "t05": 1, "t06": 1, "t07": 1, "t08": 1, "t09": 0, "t10": 1}
critical = {"t04", "t09"}            # cases that must never regress (e.g. safety, legal wording)
fixed = [t for t in baseline if not baseline[t] and candidate[t]]
broken = [t for t in baseline if baseline[t] and not candidate[t]]
print(f"pass rate {sum(baseline.values())}/10 -> {sum(candidate.values())}/10")
print("newly passing:", fixed, "| newly failing:", broken)
blocked = sorted(set(broken) & critical)
print("gate:", f"BLOCK (critical regressions {blocked})" if blocked else "pass")
```

Output:

```
pass rate 8/10 -> 8/10
newly passing: ['t03', 't07'] | newly failing: ['t04', 't09']
gate: BLOCK (critical regressions ['t04', 't09'])
```

## Show diffs of outputs

For each newly failing case, show the old and new outputs side by side so reviewers see what changed.

**Quiz:** Why can the average stay the same while the gate blocks a change?

- [ ] Averages are always wrong
- [x] Critical cases regressed even though others improved
- [ ] The gate ignores test cases
- [ ] Gates only check spelling

*Answer:* Critical cases regressed even though others improved. Look at which cases changed, not only how many.
