पाठ 3 / 25
Building an Evaluation Set
Realistic questions, known answers, labelled passages.
Sources of good questions
An evaluation set is a list of questions, the relevant passage ids (or graded labels), and ideally a reference answer. Good sources: real user queries from logs (anonymised), support tickets, subject-matter experts writing questions, and synthetic questions generated by a model from passages. Synthetic questions are cheap but tend to copy the passage wording, which makes retrieval look easier than reality; review and rewrite a sample and keep real queries in the mix. Include hard cases: typos, vague questions, multi-part questions, questions whose answer is absent, and every important user segment. Start with 50 to 100 and grow it from production failures.
One evaluation record
Store sets as versioned JSONL.
{"id": "q-0142",
"question": "how long do refunds take for laptops",
"relevant": {"policy-refund#3": 3, "faq-refunds#1": 2},
"reference_answer": "Within 14 days of receiving the returned laptop.",
"slice": "billing", "source": "support-ticket", "answerable": true}Add every production failure
When a user reports a bad answer, add the question to the set with labels; your set then grows where the system is weak.
त्वरित जाँच: What is a known weakness of model-generated synthetic questions?
- They require no review because they are perfect
- They are always wrong
- They cannot be stored
- They often reuse passage wording, making retrieval look easier than with real queries
Answer
They often reuse passage wording, making retrieval look easier than with real queries — Mix in and prefer real queries.