Lesson 2 / 25

What to Score: Prompt, Output, System

Three levels of quality.

Each level answers a different question

You can score the prompt text itself (is it clear, complete and unambiguous? cheap static checks); the outputs it produces on test inputs (are they correct, well-formatted, faithful, safe?); and the system in use (do users succeed, how often do they edit or complain, what does it cost?). Output scoring is the core, because users experience outputs, but static checks catch obvious problems early and system metrics confirm real-world value.

Levels of scoring

From cheapest to most realistic.

level        examples of scores                              cost      when
prompt text  lint checklist, vague words, conflicts           ~free     every edit
outputs      format checks, rubric scores, judge ratings       moderate  every edit + release
system       task success, edits, thumbs, escalations, cost    high      after launch, A/B tests

Use all three, at different speeds

Lint on every keystroke, score outputs on every commit, watch system metrics weekly.

Quick check: Which level of scoring reflects what users actually experience most directly?

  • Scoring the outputs (and system metrics in use)
  • Counting words in the prompt
  • The prompt file name
  • The number of prompt versions
Answer

Scoring the outputs (and system metrics in use) — Users see outputs, not prompts.