SkillByAIOpen interactive version →

RAG Retrieval & Evaluation

Measure a RAG system instead of eyeballing it: evaluation sets, retrieval metrics, statistics, retriever and chunking experiments, answer faithfulness, calibrated LLM judges and release gates, with every snippet run.

Start course →

What you'll learn

Syllabus

Why Evaluate Retrieval Separately

  1. Retrieval Errors Versus Generation Errors
  2. What Counts as Relevant
  3. Building an Evaluation Set

Retrieval Metrics

  1. Recall@k, Precision@k, Hit Rate and MRR
  2. Graded Relevance and nDCG
  3. Choosing k and the Primary Metric
  4. Context Precision and Noise

Is the Difference Real?

  1. How Many Queries Do You Need?
  2. Paired Comparisons and the Bootstrap
  3. Slices: Averages Hide Weak Spots

Experiments on Retrieval Design

  1. Comparing Retrievers on Your Queries
  2. Chunk Size Experiments
  3. Evaluating Rerankers, Hybrid Search and Query Rewriting

Evaluating Generated Answers

  1. Reference-Based Metrics: Exact Match and Token F1
  2. Faithfulness: Is Every Claim Supported?
  3. Citation Precision and Recall

LLM Judges You Can Trust

  1. Designing a Judge Prompt
  2. Calibrating a Judge Against Humans
  3. Judge Biases and Mitigations

Evaluation in the Release Cycle and in Production

  1. Evaluation as a Release Gate
  2. Online Signals and A/B Tests
  3. Drift and Index Freshness

Putting Evaluation Into Practice

  1. Failure Analysis With a Taxonomy
  2. Case Study: Evaluating a Policy Assistant
  3. A RAG Evaluation Checklist