पाठ 4 / 25

Features, Labels and Train/Test Splits

Measure on data the model has not seen.

The test set is your honest exam

A dataset is a table: each row an example, columns are features (inputs) and one column is the label (target). To measure how a model will do on new data, hold some rows out as a test set that is never used for training or tuning. A model scored on its own training data can look perfect simply by memorising. Typical splits are 70-80% train and 20-30% test; for classification use a stratified split so class proportions stay the same, and for time-based data split by time (train on the past, test on the future).

Features, splits, transformations

How you prepare and split data decides whether your results can be trusted.

Three ideas: split honestly, avoid leakage, transform features.
Figure 2.1 — Split, leakage and transformation.

Training accuracy versus test accuracy, run

I ran this with Python 3, numpy 2.5.3 and scikit-learn 1.9.1, using fixed random seeds. On the bundled breast cancer dataset (569 rows, 30 features), an unrestricted decision tree scores 1.0 on its training data but 0.902 on the held-out test set. The test number is the one that matters.

from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import train_test_split
from sklearn.tree import DecisionTreeClassifier
X, y = load_breast_cancer(return_X_y=True)
print("dataset:", X.shape[0], "rows,", X.shape[1], "features")
X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.25, random_state=0, stratify=y)
model = DecisionTreeClassifier(random_state=0).fit(X_tr, y_tr)
print("accuracy on training data:", round(model.score(X_tr, y_tr), 3))
print("accuracy on unseen test  :", round(model.score(X_te, y_te), 3))

Output:

dataset: 569 rows, 30 features
accuracy on training data: 1.0
accuracy on unseen test  : 0.902

Fix the random seed

Set random_state so splits are reproducible and comparisons between models are fair.

त्वरित जाँच: Why is training accuracy a poor measure of model quality?

  • Training data is always wrong
  • The model may have memorised the training rows
  • It is always lower than test accuracy
  • It cannot be computed
Answer

The model may have memorised the training rows — Judge models on unseen data.