Lesson 7 / 27

The Cost of Long Prompts and Prompt Caching

Know what repeated examples cost, and what caching changes.

Pay for the examples on every call

Long few-shot prompts are paid for on every request and add latency. Two things soften this. Prompt caching (offered by several providers) bills a repeated, unchanging prefix at a reduced rate and speeds it up, so put stable content first. And dynamic few-shot picks only the most relevant examples per request with similarity search. Either may remove most of the pressure to fine-tune just to save tokens, so price the long prompt with these before deciding, and re-check when prices change.

A lunch box versus a recipe

Carrying the full recipe to every kitchen is a long prompt; training the cook so they already know it is fine-tuning; caching is leaving the recipe pinned on the wall.

Stable first, variable last

Order the prompt as instructions, fixed examples, then the user input, so the cacheable prefix is as long as possible.

Quick check: Which can cut the cost of a long fixed few-shot prompt without training?

  • Prompt caching and selecting only the most relevant examples per request
  • Making the examples longer
  • Removing the evaluation set
  • Doubling the temperature
Answer

Prompt caching and selecting only the most relevant examples per request — Cheaper prompting often removes the cost argument for fine-tuning.