# Designing a Rubric — Prompt Quality Scoring

Source: https://www.skillbyai.com/en/prompt-quality-scoring/r-design

> Criteria, scales and anchors.

## Specific criteria with anchored scales

A **rubric** lists the criteria that define a good output for this task: faithfulness to the source, addressing the user's request, correct format, concision, tone, safety. For each criterion define a small scale (pass/fail or 1 to 3 is easier to apply consistently than 1 to 10) with **anchors**: a written description and example of each level. Keep criteria independent where possible, so one problem is not counted several times.

## Define good before you grade

A rubric breaks quality into criteria with weights and must-pass rules, applied consistently.

![Three ideas: rubric design, must-pass criteria, human grading.](assets/figures/prompt-quality-scoring/section-4-map.svg) — Figure 4.1 — Rubrics, must-pass criteria and human grading.

## A rubric with anchors

Each level described in words.

```text
criterion: faithful to source (pass/fail, MUST PASS)
  pass: every claim is supported by the ticket text
  fail: any claim not in the ticket (invented order numbers, promises)
criterion: addresses the ask (0 / 0.5 / 1)
  1: states what the customer wants clearly
  0.5: mentions it vaguely
  0: missing or wrong
criterion: concise (0 / 0.5 / 1)
  1: <= 60 words, no repetition   0.5: 60-100 words   0: > 100 words
```

## Pilot the rubric

Have two people grade 20 outputs, discuss disagreements and refine anchors before scaling up.

**Quiz:** Why prefer small scales with anchors over 1-10 ratings?

- [x] They are applied more consistently by different graders
- [ ] They are always more lenient
- [ ] Larger scales are illegal
- [ ] Anchors make grading slower without benefit

*Answer:* They are applied more consistently by different graders. Consistency matters more than resolution.
