पाठ 7 / 25
Format and Constraint Checks
Parse it, validate it, measure it.
Code-checkable criteria first
Before any subjective grading, check what code can decide exactly: does the output parse as JSON, does it have the required fields, are values from the allowed set, is it within the length limit, does it avoid forbidden phrases, does it include a citation id that exists? Report the pass rate for each check separately, so you know which instruction the prompt fails to enforce. These checks are free, deterministic and fast enough to run on every output in production too.
Free, exact, repeatable
Many quality aspects can be checked with code, no judgement needed.
Per-check pass rates on five outputs, run
I ran this with Python 3 (scipy 1.18.1 where imported) on example or seeded simulated data, not results from a real product. Across five example outputs, JSON validity, required keys and length each pass 80%, but allowed urgency values only 60% (one output says "urgent"); only 2 of 5 outputs pass every check. The breakdown shows which instruction needs strengthening.
import json, re
outputs = [
'{"issue": "login crash", "customer_ask": "fix", "urgency": "high"}',
'{"issue": "refund", "customer_ask": "money back", "urgency": "urgent"}',
'Sure! Here is the summary: {"issue": "slow app"}',
'{"issue": "billing", "customer_ask": "unknown", "urgency": "low"}',
'{"issue": "x", "customer_ask": "y", "urgency": "medium", "note": "' + "a" * 600 + '"}',
]
def checks(o):
r = {"valid_json": False, "required_keys": False, "allowed_urgency": False, "length_ok": len(o) <= 400}
try:
d = json.loads(o); r["valid_json"] = True
r["required_keys"] = {"issue", "customer_ask", "urgency"} <= d.keys()
r["allowed_urgency"] = d.get("urgency") in {"low", "medium", "high"}
except json.JSONDecodeError:
pass
return r
rows = [checks(o) for o in outputs]
for k in rows[0]:
print(f"{k:<16} pass rate {sum(r[k] for r in rows) / len(rows):.0%}")
print("all checks pass:", f"{sum(all(r.values()) for r in rows)}/{len(rows)} outputs")
Output:
valid_json pass rate 80% required_keys pass rate 80% allowed_urgency pass rate 60% length_ok pass rate 80% all checks pass: 2/5 outputs
Use structured output features too
Where the API supports JSON schemas or structured outputs, use them, and keep the code checks as a safety net.
त्वरित जाँच: Why report pass rates per check rather than one combined number?
- It shows which specific instruction the prompt fails to enforce
- Combined numbers are illegal
- Per-check rates are always higher
- It hides failures
Answer
It shows which specific instruction the prompt fails to enforce — Specific failures suggest specific fixes.