Lesson 13 / 25
Designing a Judge Prompt
Narrow questions, evidence, structured verdicts.
Grade one criterion at a time
An LLM judge applies your rubric automatically. Reliable judges grade one criterion per call (or a small set), receive the evidence they need (input, output, source context, reference answer if any), use the same anchored scale as human graders, explain briefly, and return a structured verdict. Use a strong model, keep the judge prompt versioned, and avoid letting the same model grade its own outputs without checks.
Scale grading without losing trust
Models can grade outputs at scale, but only after calibration and bias checks.
A faithfulness judge prompt
One criterion, evidence included, structured output.
You grade ONE criterion: is every claim in the SUMMARY supported by the TICKET?
Pass = all claims supported. Fail = any claim missing from or contradicting the ticket.
<ticket>{ticket}</ticket>
<summary>{summary}</summary>
List each claim with "supported" or "unsupported", then return JSON:
{"verdict": "pass" | "fail", "unsupported_claims": [...]}Ask for evidence
Requiring the judge to list supported and unsupported claims makes its verdicts checkable and more accurate.
Quick check: Which judge design tends to be more reliable?
- A single 1-10 score for "overall quality" with no rubric
- One criterion per call with evidence and a structured verdict
- No access to the source text
- Free-form essays as verdicts
Answer
One criterion per call with evidence and a structured verdict — Narrow and evidence-based.