Lesson 23 / 27
Case Study: Support Ticket Routing
Follow the ladder for a real-shaped problem.
From a 3,500-token prompt to a decision
A team routes 200,000 support tickets a month into 12 categories. Their few-shot prompt is 3,500 tokens and a strong model scores well but costs too much and is sometimes slow. Following the ladder: they build a 600-ticket evaluation set labelled by two agents, tighten the prompt and try prompt caching (cheaper, still above budget), then check the learning curve: accuracy is still rising with data. They have 8,000 reviewed historical tickets, so they fine-tune a small model with a short prompt, compare it with the best prompted baseline on the held-out set, check the old categories for forgetting, and keep the prompted model as fallback. This story is a hypothetical illustration, not a measured result; their own numbers would decide.
Case study, platform choice and lifecycle
Walk a realistic decision end to end, then plan for the model's whole life.
The plan as a checklist
Order matters more than any single step.
1. 600-ticket eval set, two labellers, agreement measured
2. best prompt + caching baseline: quality, cost, latency recorded
3. learning curve on historical tickets: still rising -> training may help
4. clean + dedupe + group-split 8,000 tickets; scan for personal data
5. fine-tune a small model (LoRA first), short prompt, early stopping
6. compare with step 2 on the held-out set, with intervals
7. check old categories and refusal behaviour for regressions
8. ship behind a flag, prompted model as fallback, monitor driftDecide the bar in advance
Write the required gain (for example, quality parity at half the cost) before training, so the result cannot be spun afterwards.
Quick check: In the case study, what justified fine-tuning?
- Cheaper prompting was still over budget and the learning curve was still rising with data
- Fine-tuning was fashionable
- They had no evaluation set
- Retrieval was impossible
Answer
Cheaper prompting was still over budget and the learning curve was still rising with data — Evidence from the ladder, not enthusiasm, justified the step.