पाठ 24 / 25
Case Study: Evaluating a Policy Assistant
The whole loop on one realistic scenario.
From a hunch to a measured fix
A hypothetical team hears complaints that their policy assistant gives outdated answers. They build a 300-question set from support tickets and expert-written questions, labelled with graded relevance and a reference answer, tagged by product and language. Baseline: recall@5 0.78, faithfulness 0.90, and a Hindi slice at 0.52. Failure analysis shows outdated versions and vocabulary mismatch as the biggest causes. They add a current-version metadata filter and hybrid search, evaluate with a paired bootstrap, confirm the improvement on each slice, add a release gate, and monitor not-found rates. The numbers here are illustrative; the method is the point.
The evaluation loop
Repeat this cycle for every change.
collect questions (logs, tickets, experts) -> label relevance + references
-> baseline metrics overall + per slice (with n and intervals)
-> failure analysis -> pick biggest cause
-> change one component -> paired comparison vs baseline
-> release gate -> deploy -> online signals + drift report
-> add new failures to the set -> repeatShow the slice table to stakeholders
Per-slice results make priorities obvious and stop debates over a single average.
त्वरित जाँच: In the case study, what came right after measuring the baseline?
- Deleting the Hindi queries
- Switching to the newest model immediately
- Failure analysis to find the biggest cause
- Skipping evaluation
Answer
Failure analysis to find the biggest cause — Diagnose before changing.