पाठ 25 / 25
A RAG Evaluation Checklist
Use it before trusting any RAG metric.
Ten questions
Is there a versioned evaluation set with real queries, hard cases and unanswerable questions? Are relevance rules written and labels checked for agreement? Is retrieval measured separately from answers, at the k actually sent? Are metrics reported with n, intervals and per slice? Are comparisons paired? Are chunking, retriever and stage changes evaluated one at a time? Are answers checked for correctness, faithfulness and citations? Are LLM judges calibrated against humans and checked for bias? Is there a release gate? Are online signals and drift monitored and failures fed back into the set?
The checklist
Use it in design and release reviews.
[ ] versioned eval set: real queries, hard + unanswerable cases, slices
[ ] written relevance rules; labeller agreement measured
[ ] retrieval metrics at the k actually sent (recall, MRR / nDCG)
[ ] n, intervals and per-slice results reported
[ ] paired comparisons for A vs B
[ ] one-variable experiments for retriever / chunking / reranker
[ ] answer metrics: correctness, faithfulness, citations
[ ] LLM judges calibrated (kappa, failure recall), order-swapped
[ ] release gate with guardrails and slice floors
[ ] online signals, drift report, failures added back to the setStart small, but start
Fifty labelled questions and recall@5 today beat a perfect evaluation framework next quarter.
त्वरित जाँच: Which item belongs on a RAG evaluation checklist?
- Reporting averages without sample size
- Only reading the final answer by eye
- Measuring retrieval separately at the k actually sent to the model
- Letting the generator grade itself unchecked
Answer
Measuring retrieval separately at the k actually sent to the model — Separate, quantified, repeatable measurement.