पाठ 1 / 25
The Problem With Eyeballing Prompts
Five examples are not evidence.
Anecdotes mislead
The usual prompt workflow is to try a change on a few inputs, like the outputs, and ship it. But LLM outputs vary between runs, a change that fixes one case often breaks another, and people remember striking examples. Prompt quality scoring replaces impressions with measurements: a fixed set of test inputs, defined criteria, repeatable scoring and comparisons that account for noise. It does not remove judgement; it makes judgement visible and consistent.
From hunches to numbers
Prompt changes feel better or worse; scoring tells you whether they are.
Tasting one spoonful
A chef who tastes one spoonful may miss that the other half of the pot is over-salted; good kitchens taste systematically, at several points, against a standard.
Keep a "before" snapshot
Before editing a prompt, score the current version on your test set so every change has a baseline.
त्वरित जाँच: Why is checking a prompt change on a handful of examples risky?
- Small samples are always representative
- LLMs always give identical answers
- Outputs vary and a fix for one case can silently break others
- Prompts cannot be changed
Answer
Outputs vary and a fix for one case can silently break others — Measure on a representative set.