Lesson 4 / 27

The Economics: Break-Even Volume

Compare a long prompt on a base model with a short prompt on a tuned one.

Training is a fixed cost; the saving is per call

Fine-tuning has a one-off cost (data preparation and review, training, evaluation, engineering time) and maybe a different per-token price for the tuned model. The benefit is per call: shorter prompt, cheaper model, fewer retries, lower latency. So there is a break-even volume. Below it, prompting wins on cost; above it, tuning can pay back. Count hidden costs too: maintaining the dataset, re-training when the base model is retired, and evaluating every time. The prices below are invented placeholders.

A break-even calculation, run

I ran this with plain Python 3 (standard library only), using example numbers. With invented prices, a 3,500-token few-shot prompt costs $0.01275 per call; a tuned model needing only 400 prompt tokens (assumed 1.5x per-token price) costs $0.00517, saving $0.00757. A one-off $800 investment breaks even near 105,600 calls.

# Break-even between a long few-shot prompt and a fine-tuned model that needs a short prompt. All prices are examples.
price_in, price_out = 3.0, 15.0          # base model, USD per million tokens
ft_mult = 1.5                            # assume the fine-tuned model costs 1.5x per token to serve
long_prompt, short_prompt, out_tokens = 3500, 400, 150
train_cost = 800.0                       # one-off cost of preparing data + training + evaluating (example)

per_call_base = (long_prompt * price_in + out_tokens * price_out) / 1e6
per_call_ft = ((short_prompt * price_in + out_tokens * price_out) * ft_mult) / 1e6
saving = per_call_base - per_call_ft
print(f"base model + long prompt: ${per_call_base:.5f} per call")
print(f"fine-tuned + short prompt: ${per_call_ft:.5f} per call  (saves ${saving:.5f})")
print(f"break-even volume: {train_cost / saving:,.0f} calls")
for calls in (50_000, 200_000, 1_000_000):
    print(f"  {calls:>9,} calls: net {calls * saving - train_cost:>+10,.0f}")

Output:

base model + long prompt: $0.01275 per call
fine-tuned + short prompt: $0.00517 per call  (saves $0.00757)
break-even volume: 105,611 calls
     50,000 calls: net       -421
    200,000 calls: net       +715
  1,000,000 calls: net     +6,775

Price the cheap prompt first

Before concluding tuning is cheaper, price the long prompt with prompt caching and per-request example selection; either can remove the argument.

Quick check: Below the break-even volume, which option is usually cheaper?

  • Fine-tuning
  • Staying with prompting
  • Both cost exactly the same
  • Neither can be compared
Answer

Staying with prompting — Fixed training costs are not recovered at low volume.