Lesson 13 / 25
A/B Tests and Sample Size
How many users you need to see a real effect.
Size before you start
To show that an AI feature actually improves a success metric, run an A/B test: randomly assign users to the feature or not, and compare. The number of users needed depends on the baseline rate, the size of effect you want to detect, the significance level and the power. Small effects need very large samples. Compute the sample size in advance, avoid stopping early when results look good, and pre-register the primary metric.
Users per arm for different effect sizes, run
I ran this with Python 3 (scipy 1.18.1 where imported) on example numbers, not data from a real product. With a 20% baseline completion rate, 5% significance and 80% power, detecting a 20% relative lift needs about 1,683 users per arm, a 10% lift about 6,510, a 5% lift about 25,583, and a 2% lift about 158,150.
from math import ceil, sqrt
from scipy.stats import norm
def n_per_arm(p0, lift, alpha=0.05, power=0.8):
p1 = p0 * (1 + lift); pbar = (p0 + p1) / 2
za, zb = norm.ppf(1 - alpha / 2), norm.ppf(power)
return ceil((za * sqrt(2 * pbar * (1 - pbar)) + zb * sqrt(p0 * (1 - p0) + p1 * (1 - p1))) ** 2 / (p1 - p0) ** 2)
base = 0.20 # baseline task-completion rate
for lift in [0.20, 0.10, 0.05, 0.02]:
print(f"detect a {lift:.0%} relative lift on {base:.0%}: ~{n_per_arm(base, lift):,} users per arm")
Output:
detect a 20% relative lift on 20%: ~1,683 users per arm detect a 10% relative lift on 20%: ~6,510 users per arm detect a 5% relative lift on 20%: ~25,583 users per arm detect a 2% relative lift on 20%: ~158,150 users per arm
Do not peek and stop
Checking results repeatedly and stopping at the first good-looking moment inflates false positives; use a fixed horizon or a sequential testing method.
Quick check: What happens to the required sample size as the effect you want to detect gets smaller?
- It shrinks
- It grows rapidly
- It stays the same
- It becomes zero
Answer
It grows rapidly — Small effects need large samples.