# Data Leakage — NumPy / Pandas / scikit-learn

Source: https://www.skillbyai.com/en/numpy-pandas-sklearn/x-leakage

> When the model sees the answers.

## Fit everything inside the folds

**Leakage** happens when information from validation or test data reaches training: scaling, imputing or selecting features on the full dataset before splitting, using columns that are only known after the outcome, or splitting related rows (same customer) across train and test. Leaky results look excellent and collapse in production. Put every data-dependent step inside a Pipeline and split by time or group when appropriate.

## Honest results, reusable models

Leakage inflates results, saved models need versions, and a checklist keeps workflows sound.

![Three ideas: data leakage, saving models, checklist.](assets/figures/numpy-pandas-sklearn/section-8-map.svg) — Figure 8.1 — Leakage, persistence and checklist.

## Feature selection before cross-validation, run

I ran this with Python 3.12.3 and scikit-learn 1.9.1 on a dataset bundled with scikit-learn (no download). This uses random features and random labels, so no real signal exists and honest accuracy should be about 50%. Selecting the 20 best features on all data first gives 0.85; doing the selection inside each fold gives 0.48.

```python
import numpy as np
from sklearn.feature_selection import SelectKBest, f_classif
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import cross_val_score
from sklearn.pipeline import make_pipeline

rng = np.random.default_rng(0)
X = rng.normal(size=(100, 5000))          # pure noise features
y = rng.integers(0, 2, size=100)          # random labels: true accuracy is 50%

X_sel = SelectKBest(f_classif, k=20).fit_transform(X, y)    # WRONG: uses all labels
leaky = cross_val_score(LogisticRegression(), X_sel, y, cv=5).mean()

pipe = make_pipeline(SelectKBest(f_classif, k=20), LogisticRegression())
honest = cross_val_score(pipe, X, y, cv=5).mean()           # selection inside each fold
print(f"leaky CV accuracy : {leaky:.2f}")
print(f"honest CV accuracy: {honest:.2f}")
```

Output:

```
leaky CV accuracy : 0.85
honest CV accuracy: 0.48
```

## Be suspicious of great results

If a model looks too good, look for leakage before celebrating.

**Quiz:** How do you prevent leakage from feature selection or scaling?

- [x] Put them in a Pipeline so they are fitted only on training folds
- [ ] Apply them to all data before splitting
- [ ] Skip cross-validation
- [ ] Use a larger test set only

*Answer:* Put them in a Pipeline so they are fitted only on training folds. Fit preprocessing on training data only.
