Lesson 25 / 25

A Prompt Scoring Checklist

Before trusting any prompt score.

Ten questions

Is there a versioned test set with slices, hard cases, declines and a held-out set? Is the prompt linted? Are format and constraint checks automated with per-check pass rates? Is there an anchored rubric with must-pass criteria? Are human grades blind and double-graded on a sample? Is the LLM judge calibrated (weighted kappa, generosity, recall on failures) and checked for length and position bias? Are comparisons paired and significance-tested, with variance measured? Is the choice on the quality-cost frontier? Does a case-level regression gate protect critical cases?

The checklist

Use it before declaring a prompt "better".

[ ] versioned test set: typical, slices, hard, adversarial, decline + held-out
[ ] prompt lint: task, format, constraints, example, missing info, delimiters
[ ] deterministic checks with per-check pass rates
[ ] anchored rubric; must-pass criteria reported first
[ ] blind human grading; 10% double-graded; agreement measured
[ ] judge calibrated: weighted kappa, generosity, recall on failures
[ ] length bias and position bias checked
[ ] paired comparison + significance; A/A noise floor known
[ ] option chosen on the quality-cost frontier
[ ] case-level regression gate with critical cases in CI

Re-score after model updates

A provider model update can change every score; re-run the full pipeline before trusting old numbers.

Quick check: Which item belongs on a prompt scoring checklist?

  • Optimising only the average score
  • Trusting a single run on five inputs
  • Calibrating the LLM judge against human grades
  • Grading with the version names visible
Answer

Calibrating the LLM judge against human grades — Calibrated, paired, gated.