पाठ 8 / 25
Training-Serving Skew
The model sees different features in production.
Same feature, different computation
Training-serving skew happens when features in production differ from those used in training: a different code path, different units, a changed column order, statistics recomputed on live data, or missing values filled differently. The model gets inputs it was never trained on and quality drops silently. Prevent it by packaging preprocessing with the model, sharing one feature implementation between training and serving, validating input schemas by name rather than position, and comparing feature distributions between training data and live requests.
Three serving mistakes, run
I ran this with Python 3, MLflow 3.16.1, scikit-learn 1.9.1, scipy 1.18.1 and numpy 2.5.3, using bundled or seeded synthetic data and a local SQLite tracking store. A logistic regression scores 0.958 when served with the training statistics. Re-computing scaling statistics on the live batch gives 0.951, a small loss here because the batch resembles the training data (it can be much larger for small or shifted batches). Swapping two columns in a new export drops accuracy to 0.371.
import numpy as np
from sklearn.datasets import load_breast_cancer
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split
X, y = load_breast_cancer(return_X_y=True)
X_tr, X_te, y_tr, y_te = train_test_split(X, y, random_state=0, stratify=y)
mu, sd = X_tr.mean(0), X_tr.std(0)
model = LogisticRegression(max_iter=1000).fit((X_tr - mu) / sd, y_tr)
print("serving with training statistics :", round(model.score((X_te - mu) / sd, y_te), 3))
mu_live, sd_live = X_te.mean(0), X_te.std(0) # bug: re-computed on live batch
print("serving with live-batch statistics:", round(model.score((X_te - mu_live) / sd_live, y_te), 3))
X_bad = X_te.copy(); X_bad[:, [0, 3]] = X_bad[:, [3, 0]] # bug: two columns swapped by a new export
print("two columns swapped (radius, area) :", round(model.score((X_bad - mu) / sd, y_te), 3))
Output:
serving with training statistics : 0.958 serving with live-batch statistics: 0.951 two columns swapped (radius, area) : 0.371
Send features by name
Use named fields (JSON objects, dataframes with checked columns) at the serving boundary, never bare positional arrays.
त्वरित जाँच: What is the most robust way to prevent skew in preprocessing?
- Recompute statistics on each request batch
- Re-implement it separately in the serving language
- Package the preprocessing with the model so serving uses the same code
- Skip preprocessing in production
Answer
Package the preprocessing with the model so serving uses the same code — One implementation, used in both places.