# Designing a Judge Prompt — Prompt Quality Scoring

Source: https://www.skillbyai.com/en/prompt-quality-scoring/j-design

> Narrow questions, evidence, structured verdicts.

## Grade one criterion at a time

An **LLM judge** applies your rubric automatically. Reliable judges grade **one criterion per call** (or a small set), receive the **evidence** they need (input, output, source context, reference answer if any), use the same **anchored scale** as human graders, explain briefly, and return a **structured** verdict. Use a strong model, keep the judge prompt versioned, and avoid letting the same model grade its own outputs without checks.

## Scale grading without losing trust

Models can grade outputs at scale, but only after calibration and bias checks.

![Four ideas: judge design, calibration, length bias, position bias.](assets/figures/prompt-quality-scoring/section-5-map.svg) — Figure 5.1 — Design, calibration, length bias and position bias.

## A faithfulness judge prompt

One criterion, evidence included, structured output.

```text
You grade ONE criterion: is every claim in the SUMMARY supported by the TICKET?
Pass = all claims supported. Fail = any claim missing from or contradicting the ticket.

<ticket>{ticket}</ticket>
<summary>{summary}</summary>

List each claim with "supported" or "unsupported", then return JSON:
{"verdict": "pass" | "fail", "unsupported_claims": [...]}
```

## Ask for evidence

Requiring the judge to list supported and unsupported claims makes its verdicts checkable and more accurate.

**Quiz:** Which judge design tends to be more reliable?

- [ ] A single 1-10 score for "overall quality" with no rubric
- [x] One criterion per call with evidence and a structured verdict
- [ ] No access to the source text
- [ ] Free-form essays as verdicts

*Answer:* One criterion per call with evidence and a structured verdict. Narrow and evidence-based.
