पाठ 18 / 25
Paired Significance Tests
Focus on the cases where versions disagree.
McNemar's test for pass/fail
When two prompt versions are scored pass/fail on the same test cases, only the discordant cases matter: those where one passes and the other fails. McNemar's test asks whether these discordant cases are split more unevenly than chance. It is more sensitive than comparing two pass rates as independent samples, because it uses the pairing. For continuous scores, use a paired t-test, Wilcoxon test or paired bootstrap.
McNemar on 120 shared test cases, run
I ran this with Python 3 (scipy 1.18.1 where imported) on example or seeded simulated data, not results from a real product. Prompt A passes 72.5% and prompt B 83.3% of 120 cases. Of the 25 cases where they disagree, B wins 19 and A 6; the exact McNemar p-value is 0.0146, so B is very likely better.
from scipy.stats import binomtest
# per test case: did prompt A pass? did prompt B pass? (same 120 cases)
both, a_only, b_only, neither = 81, 6, 19, 14
p = binomtest(b_only, a_only + b_only, 0.5).pvalue
print(f"A pass rate {(both + a_only) / 120:.3f} | B pass rate {(both + b_only) / 120:.3f}")
print(f"discordant cases: A-only {a_only}, B-only {b_only}")
print(f"exact McNemar test p-value {p:.4f} -> {'B is better' if p < 0.05 and b_only > a_only else 'no clear difference'}")
Output:
A pass rate 0.725 | B pass rate 0.833 discordant cases: A-only 6, B-only 19 exact McNemar test p-value 0.0146 -> B is better
Read the discordant cases
Beyond the p-value, read the cases where A passes and B fails; they show what the new prompt broke.
त्वरित जाँच: Which test cases drive McNemar's test?
- The cases where exactly one version passes
- Cases where both pass
- Cases where both fail
- Only the first ten cases
Answer
The cases where exactly one version passes — Concordant cases carry no information about the difference.