# A Scoring Pipeline — Prompt Quality Scoring

Source: https://www.skillbyai.com/en/prompt-quality-scoring/i-pipeline

> Inputs, runs, checks, judgements, report.

## Repeatable from end to end

A scoring pipeline: (1) a versioned **test set** of realistic inputs with expectations; (2) **run** the prompt on each input, sometimes several times; (3) apply **deterministic checks** (format, required fields, length); (4) apply **rubric** grading by humans or calibrated **LLM judges**; (5) **aggregate** into per-criterion, per-slice and overall scores with uncertainty; (6) **compare** with the baseline version and flag regressions. Store everything so any score can be traced back to the exact outputs.

## The pipeline

Each stage writes its results.

```text
test set (versioned) -> run prompt vN (k samples per input, fixed settings)
 -> deterministic checks: JSON valid? keys? allowed values? length?
 -> rubric grading: humans on a sample + calibrated judge on all
 -> aggregate: per criterion, per slice, overall, with intervals
 -> compare vs baseline: improved / regressed cases, significance
 -> report + gate (block critical regressions)
```

## Fix model settings when scoring

Use the same model version, temperature and parameters for both versions you compare, or you measure the settings, not the prompt.

**Quiz:** Why store raw outputs, not only final scores?

- [ ] It is required by Python
- [ ] Raw outputs are smaller
- [ ] Scores cannot be computed without deleting outputs
- [x] So any score can be traced back and checked

*Answer:* So any score can be traced back and checked. Traceability makes scores trustworthy.
