# Building Test Sets and a Scoring Dashboard — Prompt Quality Scoring

Source: https://www.skillbyai.com/en/prompt-quality-scoring/x-set

> Representative inputs, visible results.

## Coverage, slices, history

A good test set covers common inputs, important **slices** (languages, customer types, input lengths), known hard cases, adversarial inputs and cases where the right answer is to decline. Start with 50 to 200 cases and grow it from production failures. A dashboard shows, per prompt version: must-pass rate, rubric scores per criterion, per-slice results with sample sizes, output length, cost and latency, plus links to example outputs. History over versions makes trends and regressions visible.

## Set up, apply, review

Build the test set and dashboard, see the method in a case study, and use a checklist.

![Three ideas: test sets and dashboards, case study, checklist.](assets/figures/prompt-quality-scoring/section-8-map.svg) — Figure 8.1 — Test sets, case study and checklist.

## Test-set composition

Aim for coverage, not just volume.

```text
share   category
50%     typical production inputs (sampled, anonymised)
15%     important slices underrepresented in traffic (languages, segments)
15%     known hard cases and past failures
10%     adversarial / injection / unsafe requests
10%     cases that should decline or say "unknown"
+ held-out set (separate, used only at release)
```

## Anonymise production samples

Remove names, contact details and identifiers before adding real inputs to shared test sets.

**Quiz:** Why include cases where the correct behaviour is to decline?

- [x] To check the prompt does not invent answers when it should not answer
- [ ] They make scores higher automatically
- [ ] Declining is always wrong
- [ ] They are easier to grade

*Answer:* To check the prompt does not invent answers when it should not answer. Knowing when not to answer is part of quality.
