Lesson 7 / 25

Format and Constraint Checks

Parse it, validate it, measure it.

Code-checkable criteria first

Before any subjective grading, check what code can decide exactly: does the output parse as JSON, does it have the required fields, are values from the allowed set, is it within the length limit, does it avoid forbidden phrases, does it include a citation id that exists? Report the pass rate for each check separately, so you know which instruction the prompt fails to enforce. These checks are free, deterministic and fast enough to run on every output in production too.

Free, exact, repeatable

Many quality aspects can be checked with code, no judgement needed.

Three ideas: format checks, reference metrics, assertion tests.
Figure 3.1 — Format checks, reference metrics and assertions.

Per-check pass rates on five outputs, run

I ran this with Python 3 (scipy 1.18.1 where imported) on example or seeded simulated data, not results from a real product. Across five example outputs, JSON validity, required keys and length each pass 80%, but allowed urgency values only 60% (one output says "urgent"); only 2 of 5 outputs pass every check. The breakdown shows which instruction needs strengthening.

import json, re
outputs = [
    '{"issue": "login crash", "customer_ask": "fix", "urgency": "high"}',
    '{"issue": "refund", "customer_ask": "money back", "urgency": "urgent"}',
    'Sure! Here is the summary: {"issue": "slow app"}',
    '{"issue": "billing", "customer_ask": "unknown", "urgency": "low"}',
    '{"issue": "x", "customer_ask": "y", "urgency": "medium", "note": "' + "a" * 600 + '"}',
]
def checks(o):
    r = {"valid_json": False, "required_keys": False, "allowed_urgency": False, "length_ok": len(o) <= 400}
    try:
        d = json.loads(o); r["valid_json"] = True
        r["required_keys"] = {"issue", "customer_ask", "urgency"} <= d.keys()
        r["allowed_urgency"] = d.get("urgency") in {"low", "medium", "high"}
    except json.JSONDecodeError:
        pass
    return r
rows = [checks(o) for o in outputs]
for k in rows[0]:
    print(f"{k:<16} pass rate {sum(r[k] for r in rows) / len(rows):.0%}")
print("all checks pass:", f"{sum(all(r.values()) for r in rows)}/{len(rows)} outputs")

Output:

valid_json       pass rate 80%
required_keys    pass rate 80%
allowed_urgency  pass rate 60%
length_ok        pass rate 80%
all checks pass: 2/5 outputs

Use structured output features too

Where the API supports JSON schemas or structured outputs, use them, and keep the code checks as a safety net.

Quick check: Why report pass rates per check rather than one combined number?

  • It shows which specific instruction the prompt fails to enforce
  • Combined numbers are illegal
  • Per-check rates are always higher
  • It hides failures
Answer

It shows which specific instruction the prompt fails to enforce — Specific failures suggest specific fixes.