# Designing a Judge Prompt — RAG Retrieval & Evaluation

Source: https://www.skillbyai.com/en/rag-evaluation/j-rubric

> Narrow questions, explicit rubrics, structured verdicts.

## One criterion per call

An **LLM judge** grades outputs with a prompt. Reliable judges ask **one narrow question** at a time (is this claim supported by this context? does the answer address the question?), use an explicit **rubric** with examples for each score, receive the **evidence** they need (question, context, reference answer if any), and return a **structured verdict** with a short reason. Prefer binary or 3-level scales over 1 to 10. Use a strong model for judging, keep the judge prompt versioned, and never let the system under test grade itself without checks.

## Rubrics, calibration, bias

Model-based grading scales evaluation, but only after you check it against humans.

![Three ideas: rubric, calibration, bias.](assets/figures/rag-evaluation/section-6-map.svg) — Figure 6.1 — Rubric, calibration and bias.

## A faithfulness judge prompt

One claim per call, structured output. Not run here.

```text
You check whether a CLAIM is supported by the CONTEXT.
Supported = the context states it or directly implies it.
Not supported = absent from the context or contradicted by it.

<context>{context}</context>
<claim>{claim}</claim>

Return JSON: {"verdict": "supported" | "not_supported", "evidence": "<quote or empty>"}
```

## Ask for the quote

Requiring the supporting quote makes verdicts checkable and reduces lazy agreement.

**Quiz:** Which judge design is more reliable?

- [ ] Letting the generator grade itself unchecked
- [ ] A single 1-10 overall quality score with no rubric
- [x] One narrow criterion per call with a rubric and structured verdict
- [ ] No evidence given to the judge

*Answer:* One narrow criterion per call with a rubric and structured verdict. Narrow, evidence-based questions give consistent grades.
