Lesson 25 / 25
A Prompt Scoring Checklist
Before trusting any prompt score.
Ten questions
Is there a versioned test set with slices, hard cases, declines and a held-out set? Is the prompt linted? Are format and constraint checks automated with per-check pass rates? Is there an anchored rubric with must-pass criteria? Are human grades blind and double-graded on a sample? Is the LLM judge calibrated (weighted kappa, generosity, recall on failures) and checked for length and position bias? Are comparisons paired and significance-tested, with variance measured? Is the choice on the quality-cost frontier? Does a case-level regression gate protect critical cases?
The checklist
Use it before declaring a prompt "better".
[ ] versioned test set: typical, slices, hard, adversarial, decline + held-out
[ ] prompt lint: task, format, constraints, example, missing info, delimiters
[ ] deterministic checks with per-check pass rates
[ ] anchored rubric; must-pass criteria reported first
[ ] blind human grading; 10% double-graded; agreement measured
[ ] judge calibrated: weighted kappa, generosity, recall on failures
[ ] length bias and position bias checked
[ ] paired comparison + significance; A/A noise floor known
[ ] option chosen on the quality-cost frontier
[ ] case-level regression gate with critical cases in CIRe-score after model updates
A provider model update can change every score; re-run the full pipeline before trusting old numbers.
Quick check: Which item belongs on a prompt scoring checklist?
- Optimising only the average score
- Trusting a single run on five inputs
- Calibrating the LLM judge against human grades
- Grading with the version names visible
Answer
Calibrating the LLM judge against human grades — Calibrated, paired, gated.