पाठ 24 / 25
Case Study: Scoring a Ticket-Summary Prompt
The full method on one feature.
From v1 to a gated v3
A hypothetical team scores a support-ticket summary prompt. Linting shows v1 lacks format and constraints. v2 adds a JSON schema and delimiters; deterministic checks reveal invalid urgency values, fixed by listing allowed values. Rubric grading with faithfulness as must-pass finds invented order numbers; v3 adds "use only facts in the ticket". A judge calibrated on 120 human-graded outputs (weighted kappa 0.84) scores all cases; a length check shows no bias after adding a concision criterion. McNemar on 120 cases shows v3 beats v2, and the regression gate confirms no critical case broke. The numbers are illustrative.
The scoring history
Each version moves through the same pipeline.
version lint format pass must-pass rubric avg judge vs v-prev gate
v1 1/6 20% 55% 0.58 - -
v2 4/6 60% 71% 0.71 v2 > v1 (p < 0.01) pass
v3 6/6 96% 88% 0.83 v3 > v2 (p = 0.015) pass (0 critical regressions)Change one thing per version
Small, single-purpose prompt changes make it clear which edit caused which score change.
त्वरित जाँच: In the case study, what found the invented order numbers?
- Counting words
- The prompt linter
- Rubric grading with faithfulness as a must-pass criterion
- Latency monitoring
Answer
Rubric grading with faithfulness as a must-pass criterion — Must-pass criteria catch serious errors.