# Case Study: Support Ticket Routing — Fine-tuning vs Prompting

Source: https://www.skillbyai.com/en/fine-tuning/pr-case

> Follow the ladder for a real-shaped problem.

## From a 3,500-token prompt to a decision

A team routes 200,000 support tickets a month into 12 categories. Their few-shot prompt is 3,500 tokens and a strong model scores well but costs too much and is sometimes slow. Following the ladder: they build a 600-ticket evaluation set labelled by two agents, tighten the prompt and try prompt caching (cheaper, still above budget), then check the learning curve: accuracy is still rising with data. They have 8,000 reviewed historical tickets, so they fine-tune a small model with a short prompt, compare it with the best prompted baseline on the held-out set, check the old categories for forgetting, and keep the prompted model as fallback. This story is a **hypothetical illustration**, not a measured result; their own numbers would decide.

## Case study, platform choice and lifecycle

Walk a realistic decision end to end, then plan for the model's whole life.

![Three ideas: case, platform, lifecycle.](assets/figures/fine-tuning/section-7-map.svg) — Figure 7.1 — Case, platform and lifecycle.

## The plan as a checklist

Order matters more than any single step.

```text
1. 600-ticket eval set, two labellers, agreement measured
2. best prompt + caching baseline: quality, cost, latency recorded
3. learning curve on historical tickets: still rising -> training may help
4. clean + dedupe + group-split 8,000 tickets; scan for personal data
5. fine-tune a small model (LoRA first), short prompt, early stopping
6. compare with step 2 on the held-out set, with intervals
7. check old categories and refusal behaviour for regressions
8. ship behind a flag, prompted model as fallback, monitor drift
```

## Decide the bar in advance

Write the required gain (for example, quality parity at half the cost) before training, so the result cannot be spun afterwards.

**Quiz:** In the case study, what justified fine-tuning?

- [x] Cheaper prompting was still over budget and the learning curve was still rising with data
- [ ] Fine-tuning was fashionable
- [ ] They had no evaluation set
- [ ] Retrieval was impossible

*Answer:* Cheaper prompting was still over budget and the learning curve was still rising with data. Evidence from the ladder, not enthusiasm, justified the step.
