Lesson 15 / 25
Shadow Deployment
Run the new model on real traffic without affecting users.
Predict silently, compare later
In a shadow deployment, the new model receives a copy of live requests and its predictions are logged but never shown. You can then check real-traffic behaviour: error rates, latency, prediction distributions, and disagreements with the live model, which are excellent cases to review or label. Shadowing finds integration bugs (missing features, skew) safely. It doubles inference cost for the shadowed traffic and cannot measure effects on user behaviour, which needs a canary or A/B test.
Agreement between live and shadow models, run
I ran this with Python 3, MLflow 3.16.1, scikit-learn 1.9.1, scipy 1.18.1 and numpy 2.5.3, using bundled or seeded synthetic data and a local SQLite tracking store. On 228 held-out cases treated as live requests, a logistic regression shadow agrees with the random forest that users see on 93.9% of requests; the 14 disagreements are listed for review.
import numpy as np
from sklearn.datasets import load_breast_cancer
from sklearn.ensemble import RandomForestClassifier
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
X, y = load_breast_cancer(return_X_y=True)
X_tr, X_live, y_tr, _ = train_test_split(X, y, test_size=0.4, random_state=1)
live_model = RandomForestClassifier(random_state=0).fit(X_tr, y_tr)
shadow_model = make_pipeline(StandardScaler(), LogisticRegression()).fit(X_tr, y_tr)
served = live_model.predict(X_live) # users see these
shadow = shadow_model.predict(X_live) # logged only, never shown
disagree = np.where(served != shadow)[0]
print(f"requests {len(X_live)} | agreement {np.mean(served == shadow):.3f} | disagreements {len(disagree)}")
print("disagreement rows to review:", disagree[:5].tolist())
Output:
requests 228 | agreement 0.939 | disagreements 14 disagreement rows to review: [0, 4, 38, 76, 110]
Label the disagreements
Send shadow-versus-live disagreements to reviewers; they are the most informative examples for deciding which model is better.
Quick check: What can a shadow deployment NOT measure?
- How users react to the new model's predictions
- Latency on real traffic
- Disagreement with the live model
- Errors from missing features
Answer
How users react to the new model's predictions — Users never see shadow predictions.