Lesson 17 / 25

Designing a Judge Prompt

Narrow questions, explicit rubrics, structured verdicts.

One criterion per call

An LLM judge grades outputs with a prompt. Reliable judges ask one narrow question at a time (is this claim supported by this context? does the answer address the question?), use an explicit rubric with examples for each score, receive the evidence they need (question, context, reference answer if any), and return a structured verdict with a short reason. Prefer binary or 3-level scales over 1 to 10. Use a strong model for judging, keep the judge prompt versioned, and never let the system under test grade itself without checks.

Rubrics, calibration, bias

Model-based grading scales evaluation, but only after you check it against humans.

Three ideas: rubric, calibration, bias.
Figure 6.1 — Rubric, calibration and bias.

A faithfulness judge prompt

One claim per call, structured output. Not run here.

You check whether a CLAIM is supported by the CONTEXT.
Supported = the context states it or directly implies it.
Not supported = absent from the context or contradicted by it.

<context>{context}</context>
<claim>{claim}</claim>

Return JSON: {"verdict": "supported" | "not_supported", "evidence": "<quote or empty>"}

Ask for the quote

Requiring the supporting quote makes verdicts checkable and reduces lazy agreement.

Quick check: Which judge design is more reliable?

  • Letting the generator grade itself unchecked
  • A single 1-10 overall quality score with no rubric
  • One narrow criterion per call with a rubric and structured verdict
  • No evidence given to the judge
Answer

One narrow criterion per call with a rubric and structured verdict — Narrow, evidence-based questions give consistent grades.