Lesson 11 / 25

Shadow Mode

Run on real traffic without showing results.

Compare silently

In shadow mode the AI feature runs on real requests but its output is only logged, not shown, while users keep the existing behaviour. This tests real-world inputs, latency, cost and error rates at no user risk, and lets you compare with the current system: where they disagree, humans can judge which was right. It cannot measure how users react, and it adds cost, so keep shadow periods targeted.

Agreement between the current system and the AI in shadow, run

I ran this with Python 3 (scipy 1.18.1 where imported) on example numbers, not data from a real product. On 500 simulated tickets, an LLM classifier in shadow mode agrees with the existing rules engine 90.4% of the time; the 48 disagreements are queued for human review to see which system is right.

import random
rng = random.Random(4)
tickets = 500
old = [rng.choice(["billing", "technical", "other"]) for _ in range(tickets)]   # current rules engine
new = [o if rng.random() < 0.9 else rng.choice(["billing", "technical", "other"]) for o in old]  # LLM in shadow
agree = sum(a == b for a, b in zip(old, new)) / tickets
diffs = [(i, a, b) for i, (a, b) in enumerate(zip(old, new)) if a != b]
print(f"shadow mode: {tickets} real tickets, agreement {agree:.1%}, {len(diffs)} disagreements")
print("sample for human review:", diffs[:3])

Output:

shadow mode: 500 real tickets, agreement 90.4%, 48 disagreements
sample for human review: [(9, 'billing', 'technical'), (35, 'other', 'technical'), (38, 'technical', 'billing')]

Review disagreements, not agreements

Disagreements are where you learn whether the AI is better or worse than the current system.

Quick check: What can shadow mode NOT tell you?

  • Disagreement with the current system
  • Latency on real traffic
  • Cost per request
  • How users react to seeing the AI output
Answer

How users react to seeing the AI output — Users never see shadow results.