पाठ 18 / 25
Detecting Data Drift
Compare live inputs with training data.
PSI and statistical tests
Data drift means the distribution of inputs has changed from what the model was trained on. Common measures: the Population Stability Index (PSI) over binned values (a common rule of thumb reads below 0.1 as stable, 0.1 to 0.25 as moderate and above 0.25 as major), and statistical tests such as Kolmogorov-Smirnov. With large samples, tests flag tiny, harmless shifts as significant, so pair them with effect-size measures like PSI and with model quality. Drift is a reason to investigate, not automatically to retrain.
PSI and KS test on three live samples, run
I ran this with Python 3, MLflow 3.16.1, scikit-learn 1.9.1, scipy 1.18.1 and numpy 2.5.3, using bundled or seeded synthetic data and a local SQLite tracking store. Against a reference sample, live data from the same distribution has PSI 0.007 and KS p-value 0.2. A mean shift of 3 gives PSI 0.120 but an extremely small p-value; a shift of 10 gives PSI 0.939. The tiny p-value for a modest shift shows why significance alone is a poor alert.
import numpy as np
from scipy.stats import ks_2samp
rng = np.random.default_rng(0)
reference = rng.normal(50, 10, 5000) # e.g. order value at training time
def psi(ref, live, bins=10):
edges = np.quantile(ref, np.linspace(0, 1, bins + 1)); edges[0], edges[-1] = -np.inf, np.inf
r = np.histogram(ref, edges)[0] / len(ref); l = np.histogram(live, edges)[0] / len(live)
r, l = np.clip(r, 1e-6, None), np.clip(l, 1e-6, None)
return float(np.sum((l - r) * np.log(l / r)))
for name, live in [("same distribution", rng.normal(50, 10, 2000)),
("mean +3", rng.normal(53, 10, 2000)),
("mean +10", rng.normal(60, 10, 2000))]:
print(f"{name:<18} PSI {psi(reference, live):.3f} KS p-value {ks_2samp(reference, live).pvalue:.3g}")
Output:
same distribution PSI 0.007 KS p-value 0.2 mean +3 PSI 0.120 KS p-value 5.41e-26 mean +10 PSI 0.939 KS p-value 2.37e-188
Weight drift by importance
Drift in a feature the model barely uses matters less; prioritise alerts on the most important features.
त्वरित जाँच: Why not alert on KS p-values alone with large samples?
- Tiny, harmless shifts become statistically significant
- p-values cannot be computed on large data
- KS only works on images
- Large samples hide all drift
Answer
Tiny, harmless shifts become statistically significant — Use effect sizes and model impact as well.