Lesson 11 / 25

Weighted Scores and Must-Pass Criteria

Some failures cannot be averaged away.

Weights for trade-offs, gates for deal-breakers

A weighted score combines criteria by importance, useful for ranking versions. But averaging hides critical failures: a fluent, well-formatted answer that invents a fact can still get a decent average. Mark criteria such as faithfulness, safety and format validity as must-pass: any failure fails the output regardless of the weighted score. Report both: the share of outputs passing all must-pass criteria, and the average weighted score among them.

Weighted scores with must-pass gates, run

I ran this with Python 3 (scipy 1.18.1 where imported) on example or seeded simulated data, not results from a real product. Answer B scores 0.65 on the weighted rubric, not far below the others, but fails the must-pass faithfulness criterion and is rejected. Answers A (0.89) and C (0.82) pass.

rubric = {  # criterion: (weight, must_pass)
    "faithful to source": (0.35, True), "addresses the ask": (0.25, False),
    "correct format": (0.15, True), "concise": (0.15, False), "tone": (0.10, False)}
scores = {  # 0-1 per criterion from human or judge grading
    "answer A": {"faithful to source": 1.0, "addresses the ask": 0.8, "correct format": 1.0, "concise": 0.6, "tone": 1.0},
    "answer B": {"faithful to source": 0.0, "addresses the ask": 1.0, "correct format": 1.0, "concise": 1.0, "tone": 1.0},
    "answer C": {"faithful to source": 1.0, "addresses the ask": 0.5, "correct format": 1.0, "concise": 1.0, "tone": 0.5},
}
for name, s in scores.items():
    total = sum(w * s[c] for c, (w, _) in rubric.items())
    failed = [c for c, (_, must) in rubric.items() if must and s[c] < 1.0]
    verdict = "FAIL (must-pass: " + ", ".join(failed) + ")" if failed else "pass"
    print(f"{name}: weighted score {total:.2f} -> {verdict}")

Output:

answer A: weighted score 0.89 -> pass
answer B: weighted score 0.65 -> FAIL (must-pass: faithful to source)
answer C: weighted score 0.82 -> pass

Report the must-pass rate first

Lead reports with "share of outputs passing all must-pass criteria"; weighted averages come second.

Quick check: Why mark faithfulness as must-pass instead of just weighting it?

  • Must-pass criteria are optional
  • Faithfulness is unimportant
  • Weights cannot be numbers
  • An invented fact should fail the output regardless of other strengths
Answer

An invented fact should fail the output regardless of other strengths — Deal-breakers are gates, not weights.