# Paired Significance Tests — Prompt Quality Scoring

Source: https://www.skillbyai.com/en/prompt-quality-scoring/v-mcnemar

> Focus on the cases where versions disagree.

## McNemar's test for pass/fail

When two prompt versions are scored pass/fail on the same test cases, only the **discordant** cases matter: those where one passes and the other fails. **McNemar's test** asks whether these discordant cases are split more unevenly than chance. It is more sensitive than comparing two pass rates as independent samples, because it uses the pairing. For continuous scores, use a paired t-test, Wilcoxon test or paired bootstrap.

## McNemar on 120 shared test cases, run

I ran this with Python 3 (scipy 1.18.1 where imported) on example or seeded simulated data, not results from a real product. Prompt A passes 72.5% and prompt B 83.3% of 120 cases. Of the 25 cases where they disagree, B wins 19 and A 6; the exact McNemar p-value is 0.0146, so B is very likely better.

```python
from scipy.stats import binomtest
# per test case: did prompt A pass? did prompt B pass?  (same 120 cases)
both, a_only, b_only, neither = 81, 6, 19, 14
p = binomtest(b_only, a_only + b_only, 0.5).pvalue
print(f"A pass rate {(both + a_only) / 120:.3f} | B pass rate {(both + b_only) / 120:.3f}")
print(f"discordant cases: A-only {a_only}, B-only {b_only}")
print(f"exact McNemar test p-value {p:.4f} -> {'B is better' if p < 0.05 and b_only > a_only else 'no clear difference'}")
```

Output:

```
A pass rate 0.725 | B pass rate 0.833
discordant cases: A-only 6, B-only 19
exact McNemar test p-value 0.0146 -> B is better
```

## Read the discordant cases

Beyond the p-value, read the cases where A passes and B fails; they show what the new prompt broke.

**Quiz:** Which test cases drive McNemar's test?

- [x] The cases where exactly one version passes
- [ ] Cases where both pass
- [ ] Cases where both fail
- [ ] Only the first ten cases

*Answer:* The cases where exactly one version passes. Concordant cases carry no information about the difference.
