RAG Retrieval & Evaluation
Measure a RAG system instead of eyeballing it: evaluation sets, retrieval metrics, statistics, retriever and chunking experiments, answer faithfulness, calibrated LLM judges and release gates, with every snippet run.
What you'll learn
- Build a labelled evaluation set of realistic questions with relevance judgements.
- Compute and interpret recall@k, precision@k, hit rate, MRR and nDCG for a retriever.
- Decide whether a difference between two systems is real using sample sizes, paired bootstrap intervals and per-slice results.
- Compare retrievers, chunk sizes and rerankers with controlled experiments.
- Evaluate generated answers for correctness, faithfulness and citation quality, including calibrated LLM judges.
- Run evaluation as a release gate and monitor retrieval quality in production.
Syllabus
Why Evaluate Retrieval Separately
Retrieval Metrics
- Recall@k, Precision@k, Hit Rate and MRR
- Graded Relevance and nDCG
- Choosing k and the Primary Metric
- Context Precision and Noise
Is the Difference Real?
Experiments on Retrieval Design
- Comparing Retrievers on Your Queries
- Chunk Size Experiments
- Evaluating Rerankers, Hybrid Search and Query Rewriting
Evaluating Generated Answers
- Reference-Based Metrics: Exact Match and Token F1
- Faithfulness: Is Every Claim Supported?
- Citation Precision and Recall