पाठ 24 / 25

Case Study: A Support Assistant Over Policy Documents

Apply the techniques end to end.

From overflowing prompts to a lean context

A hypothetical team runs a support assistant over 3,000 policy pages. Symptoms: occasional context-length errors, slow answers, and citations of outdated policies. Their fixes, in order: a token budget with a trimming priority; retrieval filtered to current policy versions and the user's region; a relevance floor and de-duplication; best chunks first and the question last with a short rules reminder; history as a summary plus four recent turns; stable content first for caching; JSON output with citations validated in code; and a not-found path. They test each step with an ablation on 200 real questions. This scenario is illustrative; your own measurements decide which steps matter.

The assembled pipeline

Each stage maps to a topic in this course.

request
 -> load pinned state + summary + last 4 turns
 -> retrieve (ACL + current versions + region) -> min score -> dedupe -> compress
 -> budget check (reserve output, trim docs, then history)
 -> render template: [system][examples][history][documents best-first][question][reminder]
 -> call model (cache-friendly prefix, max output tokens)
 -> validate JSON + citations -> retry once -> fallback / NOT_FOUND
 -> log versions, tokens, cache hits, stop reason

Fix the biggest failure first

Sort failures from the evaluation by cause (retrieval miss, wrong version, format error) and fix the most common cause before polishing wording.

त्वरित जाँच: In the case study, what fixed citations of outdated policies?

  • Sending all 3,000 pages
  • Raising the temperature
  • Filtering retrieval to current policy versions
  • Removing the question
Answer

Filtering retrieval to current policy versions — Metadata filtering beats hoping the model notices dates.