पाठ 24 / 25

Case Study: Evaluating a Policy Assistant

The whole loop on one realistic scenario.

From a hunch to a measured fix

A hypothetical team hears complaints that their policy assistant gives outdated answers. They build a 300-question set from support tickets and expert-written questions, labelled with graded relevance and a reference answer, tagged by product and language. Baseline: recall@5 0.78, faithfulness 0.90, and a Hindi slice at 0.52. Failure analysis shows outdated versions and vocabulary mismatch as the biggest causes. They add a current-version metadata filter and hybrid search, evaluate with a paired bootstrap, confirm the improvement on each slice, add a release gate, and monitor not-found rates. The numbers here are illustrative; the method is the point.

The evaluation loop

Repeat this cycle for every change.

collect questions (logs, tickets, experts) -> label relevance + references
 -> baseline metrics overall + per slice (with n and intervals)
 -> failure analysis -> pick biggest cause
 -> change one component -> paired comparison vs baseline
 -> release gate -> deploy -> online signals + drift report
 -> add new failures to the set -> repeat

Show the slice table to stakeholders

Per-slice results make priorities obvious and stop debates over a single average.

त्वरित जाँच: In the case study, what came right after measuring the baseline?

  • Deleting the Hindi queries
  • Switching to the newest model immediately
  • Failure analysis to find the biggest cause
  • Skipping evaluation
Answer

Failure analysis to find the biggest cause — Diagnose before changing.