पाठ 24 / 26

Estimating Cost and Latency per Run

Add up the nodes.

Tokens per node, runs per month

A workflow's cost is the sum over its nodes: each model call costs input and output tokens at that model's price, plus any tool fees, multiplied by loop iterations and retries. Latency similarly adds up along the path (parallel branches aside). Estimate from traces: average tokens per node, how often each branch runs, and how many iterations loops take. The usual savings: smaller models for classification and guardrails, shorter instructions, fewer loop iterations, cached stable prompts, and skipping expensive nodes when a cheap check suffices.

Run it economically and safely

Estimate cost, walk through a realistic design, and review with a checklist.

Three ideas: cost, case study, checklist.
Figure 8.1 — Cost, case study and checklist.

Cost per run from per-node tokens, run

I ran this with plain Python 3. It is a small model of the workflow idea, not Agent Builder itself, and no model is called. With invented prices of $1 per million input and $4 per million output tokens, four nodes cost $0.00954 per run, about $286 for 30,000 runs a month; the specialist node is about two thirds of it. Use your provider's current prices.

# Cost per workflow run from token counts per node (prices are placeholders, not real rates).
price_in, price_out = 1.00, 4.00          # $ per million tokens, invented
nodes = {"guardrail": (300, 20), "classifier": (900, 60), "specialist": (4200, 500), "judge": (1500, 80)}
per_run = sum(i * price_in / 1e6 + o * price_out / 1e6 for i, o in nodes.values())
for name, (i, o) in nodes.items():
    print(f"{name:<11} in {i:>5}  out {o:>4}  ${i * price_in / 1e6 + o * price_out / 1e6:.5f}")
print(f"per run ${per_run:.5f} | 30,000 runs/month ${per_run * 30000:,.2f}")

Output:

guardrail   in   300  out   20  $0.00038
classifier  in   900  out   60  $0.00114
specialist  in  4200  out  500  $0.00620
judge       in  1500  out   80  $0.00182
per run $0.00954 | 30,000 runs/month $286.20

Optimise the biggest node first

Sort nodes by cost share; shaving the large specialist call beats tuning a tiny guardrail.

त्वरित जाँच: How is workflow cost per run estimated?

  • Use the canvas size
  • Count the number of nodes only
  • Sum each node's token costs, accounting for branch frequency and loop iterations
  • It is always free
Answer

Sum each node's token costs, accounting for branch frequency and loop iterations — Traces give the numbers.