# The Cost of Long Prompts and Prompt Caching — Fine-tuning vs Prompting

Source: https://www.skillbyai.com/en/fine-tuning/p-long

> Know what repeated examples cost, and what caching changes.

## Pay for the examples on every call

Long few-shot prompts are paid for **on every request** and add latency. Two things soften this. **Prompt caching** (offered by several providers) bills a repeated, unchanging prefix at a reduced rate and speeds it up, so put stable content first. And **dynamic few-shot** picks only the **most relevant examples per request** with similarity search. Either may remove most of the pressure to fine-tune just to save tokens, so price the long prompt with these before deciding, and re-check when prices change.

## A lunch box versus a recipe

Carrying the full recipe to every kitchen is a long prompt; training the cook so they already know it is fine-tuning; caching is leaving the recipe pinned on the wall.

## Stable first, variable last

Order the prompt as instructions, fixed examples, then the user input, so the cacheable prefix is as long as possible.

**Quiz:** Which can cut the cost of a long fixed few-shot prompt without training?

- [x] Prompt caching and selecting only the most relevant examples per request
- [ ] Making the examples longer
- [ ] Removing the evaluation set
- [ ] Doubling the temperature

*Answer:* Prompt caching and selecting only the most relevant examples per request. Cheaper prompting often removes the cost argument for fine-tuning.
