# Cross-Validation — NumPy / Pandas / scikit-learn

Source: https://www.skillbyai.com/en/numpy-pandas-sklearn/r-cv

> More than one split.

## k folds, mean and spread

A single train/test split can be lucky or unlucky. **k-fold cross-validation** trains on k-1 folds and validates on the remaining one, k times, giving a mean and a spread. Use `StratifiedKFold` for classification, `GroupKFold` when rows from the same user or patient must stay together, and `TimeSeriesSplit` for time-ordered data. Cross-validate the whole pipeline so preprocessing is refitted in every fold.

## Five-fold cross-validation of a pipeline, run

I ran this with Python 3.12.3 and scikit-learn 1.9.1 on a dataset bundled with scikit-learn (no download). Fold accuracies range from 0.956 to 1.0; the mean of 0.979 with a standard deviation of 0.014 is a more reliable estimate than any single split.

```python
from sklearn.datasets import load_breast_cancer
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import StratifiedKFold, cross_val_score
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler

X, y = load_breast_cancer(return_X_y=True)
pipe = make_pipeline(StandardScaler(), LogisticRegression())
cv = StratifiedKFold(n_splits=5, shuffle=True, random_state=0)
scores = cross_val_score(pipe, X, y, cv=cv, scoring="accuracy")
print(scores.round(3))
print(f"mean={scores.mean():.3f} std={scores.std():.3f}")
f1 = cross_val_score(pipe, X, y, cv=cv, scoring="f1")
print(f"f1 mean={f1.mean():.3f}")
```

Output:

```
[0.956 0.974 0.982 1.    0.982]
mean=0.979 std=0.014
f1 mean=0.983
```

## Report the spread

A mean without its standard deviation hides how much results depend on the split.

**Quiz:** Which splitter suits data ordered in time?

- [ ] Random KFold with shuffling
- [x] TimeSeriesSplit
- [ ] StratifiedKFold
- [ ] LeaveOneOut on shuffled data

*Answer:* TimeSeriesSplit. Never train on the future.
