SkillByAIOpen interactive version →

Lesson 1 / 25

The Problem With Eyeballing Prompts

Five examples are not evidence.

Anecdotes mislead

The usual prompt workflow is to try a change on a few inputs, like the outputs, and ship it. But LLM outputs vary between runs, a change that fixes one case often breaks another, and people remember striking examples. Prompt quality scoring replaces impressions with measurements: a fixed set of test inputs, defined criteria, repeatable scoring and comparisons that account for noise. It does not remove judgement; it makes judgement visible and consistent.

From hunches to numbers

Prompt changes feel better or worse; scoring tells you whether they are.

Figure 1.1 — Why, what and how to score.

Tasting one spoonful

A chef who tastes one spoonful may miss that the other half of the pot is over-salted; good kitchens taste systematically, at several points, against a standard.

Keep a "before" snapshot

Before editing a prompt, score the current version on your test set so every change has a baseline.

Quick check: Why is checking a prompt change on a handful of examples risky?

  • Small samples are always representative
  • LLMs always give identical answers
  • Outputs vary and a fix for one case can silently break others
  • Prompts cannot be changed
Answer

Outputs vary and a fix for one case can silently break others — Measure on a representative set.