Lesson 21 / 25
Cross-Validation
More than one split.
k folds, mean and spread
A single train/test split can be lucky or unlucky. k-fold cross-validation trains on k-1 folds and validates on the remaining one, k times, giving a mean and a spread. Use StratifiedKFold for classification, GroupKFold when rows from the same user or patient must stay together, and TimeSeriesSplit for time-ordered data. Cross-validate the whole pipeline so preprocessing is refitted in every fold.
Five-fold cross-validation of a pipeline, run
I ran this with Python 3.12.3 and scikit-learn 1.9.1 on a dataset bundled with scikit-learn (no download). Fold accuracies range from 0.956 to 1.0; the mean of 0.979 with a standard deviation of 0.014 is a more reliable estimate than any single split.
from sklearn.datasets import load_breast_cancer
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import StratifiedKFold, cross_val_score
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
X, y = load_breast_cancer(return_X_y=True)
pipe = make_pipeline(StandardScaler(), LogisticRegression())
cv = StratifiedKFold(n_splits=5, shuffle=True, random_state=0)
scores = cross_val_score(pipe, X, y, cv=cv, scoring="accuracy")
print(scores.round(3))
print(f"mean={scores.mean():.3f} std={scores.std():.3f}")
f1 = cross_val_score(pipe, X, y, cv=cv, scoring="f1")
print(f"f1 mean={f1.mean():.3f}")
Output:
[0.956 0.974 0.982 1. 0.982] mean=0.979 std=0.014 f1 mean=0.983
Report the spread
A mean without its standard deviation hides how much results depend on the split.
Quick check: Which splitter suits data ordered in time?
- Random KFold with shuffling
- TimeSeriesSplit
- StratifiedKFold
- LeaveOneOut on shuffled data
Answer
TimeSeriesSplit — Never train on the future.