# Case Study: Scoring a Ticket-Summary Prompt — Prompt Quality Scoring

Source: https://www.skillbyai.com/en/prompt-quality-scoring/x-case

> The full method on one feature.

## From v1 to a gated v3

A hypothetical team scores a support-ticket summary prompt. Linting shows v1 lacks format and constraints. v2 adds a JSON schema and delimiters; deterministic checks reveal invalid urgency values, fixed by listing allowed values. Rubric grading with faithfulness as must-pass finds invented order numbers; v3 adds "use only facts in the ticket". A judge calibrated on 120 human-graded outputs (weighted kappa 0.84) scores all cases; a length check shows no bias after adding a concision criterion. McNemar on 120 cases shows v3 beats v2, and the regression gate confirms no critical case broke. The numbers are illustrative.

## The scoring history

Each version moves through the same pipeline.

```text
version  lint  format pass  must-pass  rubric avg  judge vs v-prev        gate
v1       1/6   20%          55%        0.58        -                      -
v2       4/6   60%          71%        0.71        v2 > v1 (p < 0.01)     pass
v3       6/6   96%          88%        0.83        v3 > v2 (p = 0.015)    pass (0 critical regressions)
```

## Change one thing per version

Small, single-purpose prompt changes make it clear which edit caused which score change.

**Quiz:** In the case study, what found the invented order numbers?

- [ ] Counting words
- [ ] The prompt linter
- [x] Rubric grading with faithfulness as a must-pass criterion
- [ ] Latency monitoring

*Answer:* Rubric grading with faithfulness as a must-pass criterion. Must-pass criteria catch serious errors.
