पाठ 21 / 26
Evaluating Workflows With Datasets
Measure routing and answers on a fixed set.
Datasets, graders, per-step metrics
Before publishing, run the workflow on an evaluation dataset: realistic inputs with expected outcomes (correct route, required facts, allowed actions). Grade with code where possible (did it pick the billing branch? did it call the refund tool with the right amount?) and with calibrated model graders for open-ended answers. Measure per step, not only end to end: routing accuracy, tool-call correctness, guardrail false positives, and final answer quality. OpenAI's evals tooling can grade datasets and traces; any harness that stores inputs, outputs and scores works.
Routing accuracy and per-route precision, run
I ran this with plain Python 3. It is a small model of the workflow idea, not Agent Builder itself, and no model is called. On eight labelled requests, the router is right 6 times (0.75). Billing precision is only 0.60 because a password reset and a gift card question were sent to billing.
cases = [ # (input, expected route, route the workflow chose)
("I was charged twice", "billing", "billing"), ("App crashes on login", "technical", "technical"),
("Refund for order 77", "billing", "billing"), ("Can't reset password", "technical", "billing"),
("Where is your office?", "other", "other"), ("Invoice shows wrong VAT", "billing", "billing"),
("Error 500 on export", "technical", "technical"), ("Do you sell gift cards?", "other", "billing"),
]
correct = sum(e == g for _, e, g in cases)
print(f"routing accuracy: {correct}/{len(cases)} = {correct / len(cases):.2f}")
for label in ["billing", "technical", "other"]:
tp = sum(e == g == label for _, e, g in cases)
pred = sum(g == label for *_, g in cases)
print(f" {label:<9} precision {tp / pred:.2f}" if pred else f" {label:<9} never predicted")
print("misroutes:", [i for i, e, g in cases if e != g])
Output:
routing accuracy: 6/8 = 0.75 billing precision 0.60 technical precision 1.00 other precision 1.00 misroutes: ["Can't reset password", 'Do you sell gift cards?']
Grade the router separately
Most multi-agent failures start with a wrong route; a routing metric finds them faster than end-to-end scores.
त्वरित जाँच: Why measure routing accuracy separately?
- Routing is never wrong
- A wrong route causes downstream failures that end-to-end scores hide
- It replaces answer evaluation
- It reduces token cost
Answer
A wrong route causes downstream failures that end-to-end scores hide — Per-step metrics locate problems.