# Cross-Validation — Machine Learning Basics

Source: https://www.skillbyai.com/en/machine-learning/g-cv

> Average over several splits for a steadier estimate.

## k folds, k scores

A single train/test split gives a noisy estimate: a lucky or unlucky split can move the score by several points. **k-fold cross-validation** splits the training data into k parts, trains on k-1 and validates on the remaining one, k times, and averages the scores. Use it to compare models and tune settings, report the mean and spread, and keep a final **test set** aside for one last check. For time series use time-ordered splits; for grouped data (several rows per customer) use group-aware folds.

## Single splits versus 5-fold CV, run

I ran this with Python 3, numpy 2.5.3 and scikit-learn 1.9.1, using fixed random seeds. The same decision tree scores anywhere from 0.881 to 0.951 on four different random splits. Five-fold cross-validation gives a mean of 0.917 with a standard deviation of 0.016.

```python
from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import cross_val_score, train_test_split
from sklearn.tree import DecisionTreeClassifier
X, y = load_breast_cancer(return_X_y=True)
single = [round(DecisionTreeClassifier(random_state=0).fit(a, c).score(b, d), 3)
          for a, b, c, d in [train_test_split(X, y, random_state=s) for s in range(4)]]
print("four different single splits:", single)
scores = cross_val_score(DecisionTreeClassifier(random_state=0), X, y, cv=5)
print("5-fold scores:", scores.round(3), "mean", round(scores.mean(), 3), "std", round(scores.std(), 3))
```

Output:

```
four different single splits: [0.881, 0.951, 0.916, 0.944]
5-fold scores: [0.904 0.921 0.912 0.947 0.903] mean 0.917 std 0.016
```

## Report mean and spread

Write 0.917 plus or minus 0.016 rather than a single number, so readers see how stable it is.

**Quiz:** Why use cross-validation instead of one split?

- [x] It averages over several splits, giving a more reliable estimate
- [ ] It makes training faster
- [ ] It removes the need for data
- [ ] It guarantees zero error

*Answer:* It averages over several splits, giving a more reliable estimate. More splits, less luck.
