# A Prompt Scoring Checklist — Prompt Quality Scoring

Source: https://www.skillbyai.com/en/prompt-quality-scoring/x-check

> Before trusting any prompt score.

## Ten questions

Is there a versioned test set with slices, hard cases, declines and a held-out set? Is the prompt linted? Are format and constraint checks automated with per-check pass rates? Is there an anchored rubric with must-pass criteria? Are human grades blind and double-graded on a sample? Is the LLM judge calibrated (weighted kappa, generosity, recall on failures) and checked for length and position bias? Are comparisons paired and significance-tested, with variance measured? Is the choice on the quality-cost frontier? Does a case-level regression gate protect critical cases?

## The checklist

Use it before declaring a prompt "better".

```text
[ ] versioned test set: typical, slices, hard, adversarial, decline + held-out
[ ] prompt lint: task, format, constraints, example, missing info, delimiters
[ ] deterministic checks with per-check pass rates
[ ] anchored rubric; must-pass criteria reported first
[ ] blind human grading; 10% double-graded; agreement measured
[ ] judge calibrated: weighted kappa, generosity, recall on failures
[ ] length bias and position bias checked
[ ] paired comparison + significance; A/A noise floor known
[ ] option chosen on the quality-cost frontier
[ ] case-level regression gate with critical cases in CI
```

## Re-score after model updates

A provider model update can change every score; re-run the full pipeline before trusting old numbers.

**Quiz:** Which item belongs on a prompt scoring checklist?

- [ ] Optimising only the average score
- [ ] Trusting a single run on five inputs
- [x] Calibrating the LLM judge against human grades
- [ ] Grading with the version names visible

*Answer:* Calibrating the LLM judge against human grades. Calibrated, paired, gated.
