SkillByAIOpen interactive version →

Lesson 6 / 25

Cost and Latency Budgets

Price every stage before you ramp.

Per request times users times days

Measure average input and output tokens per request in testing, multiply by model prices, expected requests per user and the number of users at each stage. Costs that look trivial in dogfooding can become large at full rollout. Do the same for latency: measure p50 and p95 end to end, including retrieval and tool calls, against a budget users will accept. If budgets break at full scale, fix it before ramping: shorter prompts, smaller models for some steps, caching, or a narrower launch.

Monthly cost at each rollout stage, run

I ran this with Python 3 (scipy 1.18.1 where imported) on example numbers, not data from a real product. With placeholder prices and 2,800 input plus 350 output tokens per request, each call costs $0.0136. At 4 requests per user per day that is about $328 a month for 200 internal users, but about $409,500 a month at 250,000 users.

price_in, price_out = 3.00, 15.00     # $ per million tokens (placeholder prices)
avg_in, avg_out = 2800, 350            # tokens per request measured in testing
requests_per_user_day = 4
users = {"internal (200)": 200, "beta (2,000)": 2000, "10% rollout (25,000)": 25000, "100% (250,000)": 250000}
per_req = avg_in * price_in / 1e6 + avg_out * price_out / 1e6
print(f"cost per request ${per_req:.4f}")
for stage, n in users.items():
    month = per_req * requests_per_user_day * n * 30
    print(f"{stage:<22} ~${month:>10,.0f} per month")

Output:

cost per request $0.0136
internal (200)         ~$       328 per month
beta (2,000)           ~$     3,276 per month
10% rollout (25,000)   ~$    40,950 per month
100% (250,000)         ~$   409,500 per month

Put the full-rollout cost in the plan

Executives approving a beta should see the projected cost at 100%, not only the pilot bill.

Quick check: Why compute cost per rollout stage before launching?

  • It replaces evaluation
  • Costs never change with usage
  • Providers do not charge for tokens
  • Costs that are small in testing can become very large at full scale
Answer

Costs that are small in testing can become very large at full scale — Cost scales with exposure.