# Canary Comparisons — Safe Rollout Plans for AI Features

Source: https://www.skillbyai.com/en/ai-feature-rollouts/s-canary

> Is the exposed group doing worse?

## Compare guardrails with statistics

During each ramp stage, compare the exposed group (canary) with a control group on guardrail and success metrics. Use statistical tests suited to small samples, decide one-sided "is it worse?" checks for guardrails in advance, and look at practical size as well as significance. A guardrail that is significantly worse stops the ramp; a success metric that is not yet significantly better usually just means waiting for more data.

## Guardrail and success checks on a canary, run

I ran this with Python 3 (scipy 1.18.1 where imported) on example numbers, not data from a real product. With example counts, the canary's thumbs-down rate (4.1% versus 3.1%) is significantly worse (one-sided p = 0.043), so the ramp stops. Escalations show no significant harm, and task completion is slightly higher but not yet significant.

```python
from math import sqrt
from scipy.stats import norm
metrics = {  # metric: (control events, control n, canary events, canary n, higher_is_bad)
    "thumbs down": (310, 10000, 41, 1000, True),
    "escalation to human": (520, 10000, 49, 1000, True),
    "task completed": (7400, 10000, 760, 1000, False)}
for m, (c, n1, k, n2, bad_up) in metrics.items():
    p1, p2 = c / n1, k / n2; p = (c + k) / (n1 + n2)
    z = (p2 - p1) / sqrt(p * (1 - p) * (1 / n1 + 1 / n2))
    worse = z > 0 if bad_up else z < 0
    pval = 1 - norm.cdf(abs(z))
    verdict = "WORSE (stop)" if worse and pval < 0.05 else "no significant harm"
    print(f"{m:<20} control {p1:.3f} canary {p2:.3f} one-sided p {pval:.3f} -> {verdict}")
```

Output:

```
thumbs down          control 0.031 canary 0.041 one-sided p 0.043 -> WORSE (stop)
escalation to human  control 0.052 canary 0.049 one-sided p 0.341 -> no significant harm
task completed       control 0.740 canary 0.760 one-sided p 0.084 -> no significant harm
```

## Watch practical significance too

With huge samples tiny differences become significant; decide in advance what size of change matters.

**Quiz:** A guardrail metric is significantly worse in the canary. What should happen?

- [ ] Advance to 100% quickly
- [x] Stop the ramp and investigate
- [ ] Ignore it because success metrics matter more
- [ ] Delete the metric

*Answer:* Stop the ramp and investigate. Guardrails are stop signs.
