# Training-Serving Skew — MLOps

Source: https://www.skillbyai.com/en/mlops/d-skew

> The model sees different features in production.

## Same feature, different computation

**Training-serving skew** happens when features in production differ from those used in training: a different code path, different units, a changed column order, statistics recomputed on live data, or missing values filled differently. The model gets inputs it was never trained on and quality drops silently. Prevent it by **packaging preprocessing with the model**, sharing one feature implementation between training and serving, validating input schemas by name rather than position, and comparing feature distributions between training data and live requests.

## Three serving mistakes, run

I ran this with Python 3, MLflow 3.16.1, scikit-learn 1.9.1, scipy 1.18.1 and numpy 2.5.3, using bundled or seeded synthetic data and a local SQLite tracking store. A logistic regression scores 0.958 when served with the training statistics. Re-computing scaling statistics on the live batch gives 0.951, a small loss here because the batch resembles the training data (it can be much larger for small or shifted batches). Swapping two columns in a new export drops accuracy to 0.371.

```python
import numpy as np
from sklearn.datasets import load_breast_cancer
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split
X, y = load_breast_cancer(return_X_y=True)
X_tr, X_te, y_tr, y_te = train_test_split(X, y, random_state=0, stratify=y)
mu, sd = X_tr.mean(0), X_tr.std(0)
model = LogisticRegression(max_iter=1000).fit((X_tr - mu) / sd, y_tr)
print("serving with training statistics  :", round(model.score((X_te - mu) / sd, y_te), 3))
mu_live, sd_live = X_te.mean(0), X_te.std(0)                 # bug: re-computed on live batch
print("serving with live-batch statistics:", round(model.score((X_te - mu_live) / sd_live, y_te), 3))
X_bad = X_te.copy(); X_bad[:, [0, 3]] = X_bad[:, [3, 0]]     # bug: two columns swapped by a new export
print("two columns swapped (radius, area) :", round(model.score((X_bad - mu) / sd, y_te), 3))
```

Output:

```
serving with training statistics  : 0.958
serving with live-batch statistics: 0.951
two columns swapped (radius, area) : 0.371
```

## Send features by name

Use named fields (JSON objects, dataframes with checked columns) at the serving boundary, never bare positional arrays.

**Quiz:** What is the most robust way to prevent skew in preprocessing?

- [ ] Recompute statistics on each request batch
- [ ] Re-implement it separately in the serving language
- [x] Package the preprocessing with the model so serving uses the same code
- [ ] Skip preprocessing in production

*Answer:* Package the preprocessing with the model so serving uses the same code. One implementation, used in both places.
