# What Counts as Relevant — RAG Retrieval & Evaluation

Source: https://www.skillbyai.com/en/rag-evaluation/w-relevance

> Write a definition before anyone labels anything.

## Binary or graded, by written rules

Every retrieval metric depends on **relevance judgements**: for each query, which passages are relevant. Decide the rules first: is a passage relevant if it **contains the answer**, if it is **useful context**, or only if it is **sufficient on its own**? Use **binary** labels (relevant or not) for simple metrics, or **graded** labels (for example 0 to 3) when some passages are clearly better than others. Write examples of each grade, have two people label a sample, and measure agreement; vague definitions make metrics noisy and arguments endless.

## A graded relevance rubric

Adapt the wording to your domain.

```text
3  perfect   contains the complete answer to the question
2  useful    contains part of the answer or a necessary condition
1  marginal  on topic but does not help answer
0  not relevant

Rule: judge the passage alone, without other passages.
Rule: an outdated version of the policy is 0, even if on topic.
```

## Label outdated content as not relevant

Make freshness part of the definition, or the metric will reward retrieving superseded documents.

**Quiz:** Why write relevance rules before labelling?

- [x] Without shared rules, labels disagree and metrics become noisy
- [ ] Rules make the retriever faster
- [ ] Metrics do not use labels
- [ ] Rules replace the evaluation set

*Answer:* Without shared rules, labels disagree and metrics become noisy. Metrics are only as consistent as the labels.
