# Features, Labels and Train/Test Splits — Machine Learning Basics

Source: https://www.skillbyai.com/en/machine-learning/d-split

> Measure on data the model has not seen.

## The test set is your honest exam

A dataset is a table: each row an example, columns are **features** (inputs) and one column is the **label** (target). To measure how a model will do on **new** data, hold some rows out as a **test set** that is never used for training or tuning. A model scored on its own training data can look perfect simply by memorising. Typical splits are 70-80% train and 20-30% test; for classification use a **stratified** split so class proportions stay the same, and for time-based data split by time (train on the past, test on the future).

## Features, splits, transformations

How you prepare and split data decides whether your results can be trusted.

![Three ideas: split honestly, avoid leakage, transform features.](assets/figures/machine-learning/section-2-map.svg) — Figure 2.1 — Split, leakage and transformation.

## Training accuracy versus test accuracy, run

I ran this with Python 3, numpy 2.5.3 and scikit-learn 1.9.1, using fixed random seeds. On the bundled breast cancer dataset (569 rows, 30 features), an unrestricted decision tree scores 1.0 on its training data but 0.902 on the held-out test set. The test number is the one that matters.

```python
from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import train_test_split
from sklearn.tree import DecisionTreeClassifier
X, y = load_breast_cancer(return_X_y=True)
print("dataset:", X.shape[0], "rows,", X.shape[1], "features")
X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.25, random_state=0, stratify=y)
model = DecisionTreeClassifier(random_state=0).fit(X_tr, y_tr)
print("accuracy on training data:", round(model.score(X_tr, y_tr), 3))
print("accuracy on unseen test  :", round(model.score(X_te, y_te), 3))
```

Output:

```
dataset: 569 rows, 30 features
accuracy on training data: 1.0
accuracy on unseen test  : 0.902
```

## Fix the random seed

Set random_state so splits are reproducible and comparisons between models are fair.

**Quiz:** Why is training accuracy a poor measure of model quality?

- [ ] Training data is always wrong
- [x] The model may have memorised the training rows
- [ ] It is always lower than test accuracy
- [ ] It cannot be computed

*Answer:* The model may have memorised the training rows. Judge models on unseen data.
