SkillByAIOpen interactive version →

Lesson 24 / 25

Case Study: AI Summaries for Support Agents

The plan applied end to end.

From plan to 100%

A hypothetical company launches AI ticket summaries for support agents. The plan sets success (handle time down 10%) and guardrails (thumbs-down at most 5%, zero personal-data leaks). Offline evaluation passes overall, but the Hindi slice is below its floor, so Hindi tickets are excluded at first. Red-teaming finds a personal-data leak in summaries, fixed with redaction and retested. Dogfooding finds latency problems, fixed with a shorter prompt. The ramp pauses once at 5% when thumbs-down rises after a prompt change, which is rolled back by config version. Hindi joins later after improvements. The numbers and events are illustrative.

The timeline

Gates, pauses and decisions.

week 1   plan signed; eval: overall 0.868 ok, Hindi 0.762 below floor -> exclude Hindi
week 2   red-team: PII leak in summaries -> add redaction; retest 0/40 -> pass
week 3   dogfood: p95 6.1 s > 5 s budget -> shorter prompt -> 3.9 s
week 5   beta 14 days: useful rating 83%; support briefed
week 7   1% -> 5%: thumbs-down 4.1% vs 3.1% after prompt v8 -> roll back to v7, pause
week 8   v9 passes eval + canary -> 25% -> 50% -> 100% (non-Hindi)
week 12  Hindi slice 0.84 after fixes -> Hindi cohort ramp begins

Pauses are success, not failure

A ramp that pauses on a guardrail did its job; record it as a caught regression.

Quick check: In the case study, how was the thumbs-down spike at 5% handled?

  • The metric was removed
  • The ramp jumped to 100%
  • The config was rolled back to the previous prompt version and the ramp paused
  • Users were told to stop giving feedback
Answer

The config was rolled back to the previous prompt version and the ramp paused — Versioned config makes rollback quick.