पाठ 3 / 25

Building an Evaluation Set

Realistic questions, known answers, labelled passages.

Sources of good questions

An evaluation set is a list of questions, the relevant passage ids (or graded labels), and ideally a reference answer. Good sources: real user queries from logs (anonymised), support tickets, subject-matter experts writing questions, and synthetic questions generated by a model from passages. Synthetic questions are cheap but tend to copy the passage wording, which makes retrieval look easier than reality; review and rewrite a sample and keep real queries in the mix. Include hard cases: typos, vague questions, multi-part questions, questions whose answer is absent, and every important user segment. Start with 50 to 100 and grow it from production failures.

One evaluation record

Store sets as versioned JSONL.

{"id": "q-0142",
 "question": "how long do refunds take for laptops",
 "relevant": {"policy-refund#3": 3, "faq-refunds#1": 2},
 "reference_answer": "Within 14 days of receiving the returned laptop.",
 "slice": "billing", "source": "support-ticket", "answerable": true}

Add every production failure

When a user reports a bad answer, add the question to the set with labels; your set then grows where the system is weak.

त्वरित जाँच: What is a known weakness of model-generated synthetic questions?

  • They require no review because they are perfect
  • They are always wrong
  • They cannot be stored
  • They often reuse passage wording, making retrieval look easier than with real queries
Answer

They often reuse passage wording, making retrieval look easier than with real queries — Mix in and prefer real queries.