Lesson 2 / 25
What to Score: Prompt, Output, System
Three levels of quality.
Each level answers a different question
You can score the prompt text itself (is it clear, complete and unambiguous? cheap static checks); the outputs it produces on test inputs (are they correct, well-formatted, faithful, safe?); and the system in use (do users succeed, how often do they edit or complain, what does it cost?). Output scoring is the core, because users experience outputs, but static checks catch obvious problems early and system metrics confirm real-world value.
Levels of scoring
From cheapest to most realistic.
level examples of scores cost when
prompt text lint checklist, vague words, conflicts ~free every edit
outputs format checks, rubric scores, judge ratings moderate every edit + release
system task success, edits, thumbs, escalations, cost high after launch, A/B testsUse all three, at different speeds
Lint on every keystroke, score outputs on every commit, watch system metrics weekly.
Quick check: Which level of scoring reflects what users actually experience most directly?
- Scoring the outputs (and system metrics in use)
- Counting words in the prompt
- The prompt file name
- The number of prompt versions
Answer
Scoring the outputs (and system metrics in use) — Users see outputs, not prompts.