पाठ 5 / 25

Data Leakage

When the answer sneaks into the features.

Too good to be true

Data leakage happens when information that would not be available at prediction time, or that is derived from the label, gets into the features or the training process. Examples: a "refund issued" column when predicting complaints, statistics computed on the whole dataset before splitting, duplicate rows across train and test, or using future data to predict the past. Leakage gives excellent test scores that collapse in production. Be suspicious of near-perfect results and ask, for each feature, whether you would really know it at the moment of prediction.

A leaked column makes the test score perfect, run

I ran this with Python 3, numpy 2.5.3 and scikit-learn 1.9.1, using fixed random seeds. With five honest features a logistic regression scores 0.716 on test data. Adding one column computed from the label pushes it to 1.000, a result that would vanish in real use where that column is unknown.

import numpy as np
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split
rng = np.random.default_rng(0)
n = 1000
X = rng.normal(size=(n, 5))
y = (X[:, 0] + rng.normal(size=n) > 0).astype(int)
leaky = np.column_stack([X, y + rng.normal(scale=0.1, size=n)])   # a column computed from the label
for name, data in [("honest features", X), ("with leaked column", leaky)]:
    X_tr, X_te, y_tr, y_te = train_test_split(data, y, random_state=0)
    acc = LogisticRegression().fit(X_tr, y_tr).score(X_te, y_te)
    print(f"{name:<20} test accuracy {acc:.3f}")

Output:

honest features      test accuracy 0.716
with leaked column   test accuracy 1.000

Fit preprocessing on training data only

Put scaling and encoding inside a pipeline so they learn their statistics from the training split, not the whole dataset.

त्वरित जाँच: What is a warning sign of data leakage?

  • A model that underfits
  • Suspiciously near-perfect test results
  • Slow training
  • Missing values in one column
Answer

Suspiciously near-perfect test results — Too good to be true usually is.