पाठ 23 / 25
Building Test Sets and a Scoring Dashboard
Representative inputs, visible results.
Coverage, slices, history
A good test set covers common inputs, important slices (languages, customer types, input lengths), known hard cases, adversarial inputs and cases where the right answer is to decline. Start with 50 to 200 cases and grow it from production failures. A dashboard shows, per prompt version: must-pass rate, rubric scores per criterion, per-slice results with sample sizes, output length, cost and latency, plus links to example outputs. History over versions makes trends and regressions visible.
Set up, apply, review
Build the test set and dashboard, see the method in a case study, and use a checklist.
Test-set composition
Aim for coverage, not just volume.
share category
50% typical production inputs (sampled, anonymised)
15% important slices underrepresented in traffic (languages, segments)
15% known hard cases and past failures
10% adversarial / injection / unsafe requests
10% cases that should decline or say "unknown"
+ held-out set (separate, used only at release)Anonymise production samples
Remove names, contact details and identifiers before adding real inputs to shared test sets.
त्वरित जाँच: Why include cases where the correct behaviour is to decline?
- To check the prompt does not invent answers when it should not answer
- They make scores higher automatically
- Declining is always wrong
- They are easier to grade
Answer
To check the prompt does not invent answers when it should not answer — Knowing when not to answer is part of quality.