# Building an Evaluation Set — RAG Retrieval & Evaluation

Source: https://www.skillbyai.com/en/rag-evaluation/w-set

> Realistic questions, known answers, labelled passages.

## Sources of good questions

An evaluation set is a list of **questions**, the **relevant passage ids** (or graded labels), and ideally a **reference answer**. Good sources: real user queries from logs (anonymised), support tickets, subject-matter experts writing questions, and **synthetic questions** generated by a model from passages. Synthetic questions are cheap but tend to copy the passage wording, which makes retrieval look easier than reality; review and rewrite a sample and keep real queries in the mix. Include hard cases: typos, vague questions, multi-part questions, questions whose answer is absent, and every important user segment. Start with 50 to 100 and grow it from production failures.

## One evaluation record

Store sets as versioned JSONL.

```text
{"id": "q-0142",
 "question": "how long do refunds take for laptops",
 "relevant": {"policy-refund#3": 3, "faq-refunds#1": 2},
 "reference_answer": "Within 14 days of receiving the returned laptop.",
 "slice": "billing", "source": "support-ticket", "answerable": true}
```

## Add every production failure

When a user reports a bad answer, add the question to the set with labels; your set then grows where the system is weak.

**Quiz:** What is a known weakness of model-generated synthetic questions?

- [ ] They require no review because they are perfect
- [ ] They are always wrong
- [ ] They cannot be stored
- [x] They often reuse passage wording, making retrieval look easier than with real queries

*Answer:* They often reuse passage wording, making retrieval look easier than with real queries. Mix in and prefer real queries.
