RAG Retrieval & Evaluation

Measure a RAG system instead of eyeballing it: evaluation sets, retrieval metrics, statistics, retriever and chunking experiments, answer faithfulness, calibrated LLM judges and release gates, with every snippet run.

Start course →

What you'll learn

  • Build a labelled evaluation set of realistic questions with relevance judgements.
  • Compute and interpret recall@k, precision@k, hit rate, MRR and nDCG for a retriever.
  • Decide whether a difference between two systems is real using sample sizes, paired bootstrap intervals and per-slice results.
  • Compare retrievers, chunk sizes and rerankers with controlled experiments.
  • Evaluate generated answers for correctness, faithfulness and citation quality, including calibrated LLM judges.
  • Run evaluation as a release gate and monitor retrieval quality in production.

Syllabus

Why Evaluate Retrieval Separately

  1. Retrieval Errors Versus Generation Errors
  2. What Counts as Relevant
  3. Building an Evaluation Set

Retrieval Metrics

  1. Recall@k, Precision@k, Hit Rate and MRR
  2. Graded Relevance and nDCG
  3. Choosing k and the Primary Metric
  4. Context Precision and Noise

Is the Difference Real?

  1. How Many Queries Do You Need?
  2. Paired Comparisons and the Bootstrap
  3. Slices: Averages Hide Weak Spots

Experiments on Retrieval Design

  1. Comparing Retrievers on Your Queries
  2. Chunk Size Experiments
  3. Evaluating Rerankers, Hybrid Search and Query Rewriting

Evaluating Generated Answers

  1. Reference-Based Metrics: Exact Match and Token F1
  2. Faithfulness: Is Every Claim Supported?
  3. Citation Precision and Recall

LLM Judges You Can Trust

  1. Designing a Judge Prompt
  2. Calibrating a Judge Against Humans
  3. Judge Biases and Mitigations

Evaluation in the Release Cycle and in Production

  1. Evaluation as a Release Gate
  2. Online Signals and A/B Tests
  3. Drift and Index Freshness

Putting Evaluation Into Practice

  1. Failure Analysis With a Taxonomy
  2. Case Study: Evaluating a Policy Assistant
  3. A RAG Evaluation Checklist