Lesson 19 / 25

Run-to-Run Variance and Repeated Sampling

The same prompt scores differently each time.

Sample more than once

With non-zero temperature, and sometimes even at zero, the same prompt gives different outputs on different runs, so a single run per input is a noisy estimate. Repeating the whole evaluation can shift the average by more than the difference you are trying to detect. Run several samples per input for important comparisons, report the spread, and use the same seeds or settings where the API allows. Measure how much scores vary before trusting small differences.

How much scores move between runs, run

I ran this with Python 3 (scipy 1.18.1 where imported) on example or seeded simulated data, not results from a real product. In a simulation of 30 inputs with run-to-run noise, the same item's scores vary with a standard deviation of 0.68 across 5 runs. Five repeated single-run evaluations of the same prompt give averages from 3.35 to 3.60, a range larger than many claimed "improvements".

import random, statistics
rng = random.Random(5)
items = 30; samples = 5
true_item_quality = [rng.uniform(2.5, 4.5) for _ in range(items)]
scores = [[min(5, max(1, q + rng.gauss(0, 0.7))) for _ in range(samples)] for q in true_item_quality]
one_sample = [s[0] for s in scores]
per_item_sd = statistics.mean(statistics.pstdev(s) for s in scores)
print(f"average spread of scores for the SAME item across {samples} runs: sd {per_item_sd:.2f}")
print(f"mean score using 1 run per item : {statistics.mean(one_sample):.2f}")
print(f"mean score using {samples} runs per item: {statistics.mean(statistics.mean(s) for s in scores):.2f}")
rng2 = random.Random(9)
reruns = [statistics.mean(min(5, max(1, q + rng2.gauss(0, 0.7))) for q in true_item_quality) for _ in range(5)]
print("five repeated single-run evaluations:", [round(r, 2) for r in reruns])

Output:

average spread of scores for the SAME item across 5 runs: sd 0.68
mean score using 1 run per item : 3.66
mean score using 5 runs per item: 3.49
five repeated single-run evaluations: [3.38, 3.49, 3.35, 3.6, 3.48]

Estimate noise with an A/A test

Score the same prompt twice as if it were two versions; any "difference" you see is your noise floor.

Quick check: What is an A/A test in prompt evaluation?

  • Deleting the baseline
  • Testing two unrelated prompts
  • Testing only one input
  • Comparing a prompt with itself to measure the noise in scores
Answer

Comparing a prompt with itself to measure the noise in scores — Know your noise floor.