# Case Study: Evaluating a Policy Assistant — RAG Retrieval & Evaluation

Source: https://www.skillbyai.com/en/rag-evaluation/r-case

> The whole loop on one realistic scenario.

## From a hunch to a measured fix

A hypothetical team hears complaints that their policy assistant gives outdated answers. They build a 300-question set from support tickets and expert-written questions, labelled with graded relevance and a reference answer, tagged by product and language. Baseline: recall@5 0.78, faithfulness 0.90, and a Hindi slice at 0.52. Failure analysis shows outdated versions and vocabulary mismatch as the biggest causes. They add a current-version metadata filter and hybrid search, evaluate with a paired bootstrap, confirm the improvement on each slice, add a release gate, and monitor not-found rates. The numbers here are illustrative; the method is the point.

## The evaluation loop

Repeat this cycle for every change.

```text
collect questions (logs, tickets, experts) -> label relevance + references
 -> baseline metrics overall + per slice (with n and intervals)
 -> failure analysis -> pick biggest cause
 -> change one component -> paired comparison vs baseline
 -> release gate -> deploy -> online signals + drift report
 -> add new failures to the set -> repeat
```

## Show the slice table to stakeholders

Per-slice results make priorities obvious and stop debates over a single average.

**Quiz:** In the case study, what came right after measuring the baseline?

- [ ] Deleting the Hindi queries
- [ ] Switching to the newest model immediately
- [x] Failure analysis to find the biggest cause
- [ ] Skipping evaluation

*Answer:* Failure analysis to find the biggest cause. Diagnose before changing.
