पाठ 13 / 25

Designing a Judge Prompt

Narrow questions, evidence, structured verdicts.

Grade one criterion at a time

An LLM judge applies your rubric automatically. Reliable judges grade one criterion per call (or a small set), receive the evidence they need (input, output, source context, reference answer if any), use the same anchored scale as human graders, explain briefly, and return a structured verdict. Use a strong model, keep the judge prompt versioned, and avoid letting the same model grade its own outputs without checks.

Scale grading without losing trust

Models can grade outputs at scale, but only after calibration and bias checks.

Four ideas: judge design, calibration, length bias, position bias.
Figure 5.1 — Design, calibration, length bias and position bias.

A faithfulness judge prompt

One criterion, evidence included, structured output.

You grade ONE criterion: is every claim in the SUMMARY supported by the TICKET?
Pass = all claims supported. Fail = any claim missing from or contradicting the ticket.

<ticket>{ticket}</ticket>
<summary>{summary}</summary>

List each claim with "supported" or "unsupported", then return JSON:
{"verdict": "pass" | "fail", "unsupported_claims": [...]}

Ask for evidence

Requiring the judge to list supported and unsupported claims makes its verdicts checkable and more accurate.

त्वरित जाँच: Which judge design tends to be more reliable?

  • A single 1-10 score for "overall quality" with no rubric
  • One criterion per call with evidence and a structured verdict
  • No access to the source text
  • Free-form essays as verdicts
Answer

One criterion per call with evidence and a structured verdict — Narrow and evidence-based.