# The Problem With Eyeballing Prompts — Prompt Quality Scoring

Source: https://www.skillbyai.com/en/prompt-quality-scoring/i-why

> Five examples are not evidence.

## Anecdotes mislead

The usual prompt workflow is to try a change on a few inputs, like the outputs, and ship it. But LLM outputs vary between runs, a change that fixes one case often breaks another, and people remember striking examples. **Prompt quality scoring** replaces impressions with measurements: a fixed set of test inputs, defined criteria, repeatable scoring and comparisons that account for noise. It does not remove judgement; it makes judgement visible and consistent.

## From hunches to numbers

Prompt changes feel better or worse; scoring tells you whether they are.

![Three ideas: why measure, what to score, the scoring pipeline.](assets/figures/prompt-quality-scoring/section-1-map.svg) — Figure 1.1 — Why, what and how to score.

## Tasting one spoonful

A chef who tastes one spoonful may miss that the other half of the pot is over-salted; good kitchens taste systematically, at several points, against a standard.

## Keep a "before" snapshot

Before editing a prompt, score the current version on your test set so every change has a baseline.

**Quiz:** Why is checking a prompt change on a handful of examples risky?

- [ ] Small samples are always representative
- [ ] LLMs always give identical answers
- [x] Outputs vary and a fix for one case can silently break others
- [ ] Prompts cannot be changed

*Answer:* Outputs vary and a fix for one case can silently break others. Measure on a representative set.
