पाठ 24 / 25
Cost per Run
Find the expensive node.
Tokens by node
A workflow's cost is the sum of its model calls (input and output tokens times each model's price) plus any paid tools. Logs show token usage per node, so you can find where money goes. Typical savings: use small models for classification, extraction and checks; shorten prompts and retrieved context; lower top K; cap iteration and agent loops; cache answers to frequent questions. Set budgets and alerts at the model provider too.
Cost per run from node token counts, run
I ran this with plain Python 3 (scikit-learn 1.9.1 where imported). It models or tests one piece of a Dify app locally; Dify itself was not running. With invented prices, four nodes cost $0.00915 per run, about $457 per 50,000 runs, and the large answer model accounts for 98% of it. Shrinking its context would matter far more than optimising the classifier.
# Token usage per node from run logs; prices are placeholders, not real rates
price = {"small": (0.15, 0.60), "large": (2.50, 10.00)} # $ per million input / output tokens (invented)
runs = [ # (node, model, input tokens, output tokens)
("question classifier", "small", 350, 10), ("knowledge retrieval", None, 0, 0),
("answer LLM", "large", 2400, 300), ("tone check LLM", "small", 500, 20)]
total = 0.0
for node, model, tin, tout in runs:
cost = 0 if model is None else tin * price[model][0] / 1e6 + tout * price[model][1] / 1e6
total += cost
print(f"{node:<20} ${cost:.5f}")
print(f"per run ${total:.5f} | 50,000 runs ${total * 50000:,.2f}")
print(f"share of the answer LLM: {(2400 * 2.5 / 1e6 + 300 * 10 / 1e6) / total:.0%}")
Output:
question classifier $0.00006 knowledge retrieval $0.00000 answer LLM $0.00900 tone check LLM $0.00009 per run $0.00915 | 50,000 runs $457.28 share of the answer LLM: 98%
Optimise the biggest node first
Sort nodes by cost share and work on the top one; small nodes rarely matter.
त्वरित जाँच: Which change usually saves the most in the example?
- Renaming variables
- Removing the free retrieval node
- Reducing the large answer model's input tokens
- Changing the app icon
Answer
Reducing the large answer model's input tokens — Target the dominant cost.