पाठ 17 / 25

Structuring Prompts for Caching

Stable content first, variable content last.

Prefix caching rewards stable prefixes

Several providers offer prompt caching: a repeated, identical prefix of the prompt is processed once and reused, which lowers cost and latency on later calls. Caching works on the exact prefix, so a single changing value near the top (a timestamp, a user name, a request ID) breaks it for everything after. Put system rules, tool definitions and fixed examples first, then per-request content such as documents, the date and the question. Check your provider's rules for minimum length, how long a cache entry lives and how cache hits are reported.

Fast, parseable, reproducible

Structure the prompt so it can be cached, the output can be checked, and every change can be traced.

Three ideas: cache, contract, version.
Figure 6.1 — Cache, contract and version.

How a timestamp at the top breaks the shared prefix, run

I ran this with plain Python 3 (standard library only); the data is made-up example data. Two prompts with a timestamp first share only 10 of 2,462 characters; moving the timestamp after the stable instructions and examples makes them share 2,450 characters.

import os
def common_prefix(a, b):
    return len(os.path.commonprefix([a, b]))
system = "You are a support assistant. Rules: ... " * 40
examples = "Example 1 ... Example 2 ... " * 30
bad_1 = "Time: 10:01\n" + system + examples + "Q: refund?"
bad_2 = "Time: 10:02\n" + system + examples + "Q: shipping?"
good_1 = system + examples + "Time: 10:01\nQ: refund?"
good_2 = system + examples + "Time: 10:02\nQ: shipping?"
print("dynamic first : shared prefix", common_prefix(bad_1, bad_2), "of", len(bad_1), "chars")
print("stable first  : shared prefix", common_prefix(good_1, good_2), "of", len(good_1), "chars")

Output:

dynamic first : shared prefix 10 of 2462 chars
stable first  : shared prefix 2450 of 2462 chars

Watch the cache-hit metric

Log the cached-token count the API reports; a sudden drop usually means someone added dynamic text to the top of the prompt.

त्वरित जाँच: Where should a per-request timestamp go to keep caching effective?

  • Inside every example
  • At the very top of the prompt
  • After the stable instructions and examples
  • In the tool definitions
Answer

After the stable instructions and examples — Caching reuses the exact identical prefix.