Lesson 3 / 25

A Scoring Pipeline

Inputs, runs, checks, judgements, report.

Repeatable from end to end

A scoring pipeline: (1) a versioned test set of realistic inputs with expectations; (2) run the prompt on each input, sometimes several times; (3) apply deterministic checks (format, required fields, length); (4) apply rubric grading by humans or calibrated LLM judges; (5) aggregate into per-criterion, per-slice and overall scores with uncertainty; (6) compare with the baseline version and flag regressions. Store everything so any score can be traced back to the exact outputs.

The pipeline

Each stage writes its results.

test set (versioned) -> run prompt vN (k samples per input, fixed settings)
 -> deterministic checks: JSON valid? keys? allowed values? length?
 -> rubric grading: humans on a sample + calibrated judge on all
 -> aggregate: per criterion, per slice, overall, with intervals
 -> compare vs baseline: improved / regressed cases, significance
 -> report + gate (block critical regressions)

Fix model settings when scoring

Use the same model version, temperature and parameters for both versions you compare, or you measure the settings, not the prompt.

Quick check: Why store raw outputs, not only final scores?

  • It is required by Python
  • Raw outputs are smaller
  • Scores cannot be computed without deleting outputs
  • So any score can be traced back and checked
Answer

So any score can be traced back and checked — Traceability makes scores trustworthy.