Lesson 24 / 25
Case Study: AI Summaries for Support Agents
The plan applied end to end.
From plan to 100%
A hypothetical company launches AI ticket summaries for support agents. The plan sets success (handle time down 10%) and guardrails (thumbs-down at most 5%, zero personal-data leaks). Offline evaluation passes overall, but the Hindi slice is below its floor, so Hindi tickets are excluded at first. Red-teaming finds a personal-data leak in summaries, fixed with redaction and retested. Dogfooding finds latency problems, fixed with a shorter prompt. The ramp pauses once at 5% when thumbs-down rises after a prompt change, which is rolled back by config version. Hindi joins later after improvements. The numbers and events are illustrative.
The timeline
Gates, pauses and decisions.
week 1 plan signed; eval: overall 0.868 ok, Hindi 0.762 below floor -> exclude Hindi
week 2 red-team: PII leak in summaries -> add redaction; retest 0/40 -> pass
week 3 dogfood: p95 6.1 s > 5 s budget -> shorter prompt -> 3.9 s
week 5 beta 14 days: useful rating 83%; support briefed
week 7 1% -> 5%: thumbs-down 4.1% vs 3.1% after prompt v8 -> roll back to v7, pause
week 8 v9 passes eval + canary -> 25% -> 50% -> 100% (non-Hindi)
week 12 Hindi slice 0.84 after fixes -> Hindi cohort ramp beginsPauses are success, not failure
A ramp that pauses on a guardrail did its job; record it as a caught regression.
Quick check: In the case study, how was the thumbs-down spike at 5% handled?
- The metric was removed
- The ramp jumped to 100%
- The config was rolled back to the previous prompt version and the ramp paused
- Users were told to stop giving feedback
Answer
The config was rolled back to the previous prompt version and the ramp paused — Versioned config makes rollback quick.