पाठ 4 / 25
Features, Labels and Train/Test Splits
Measure on data the model has not seen.
The test set is your honest exam
A dataset is a table: each row an example, columns are features (inputs) and one column is the label (target). To measure how a model will do on new data, hold some rows out as a test set that is never used for training or tuning. A model scored on its own training data can look perfect simply by memorising. Typical splits are 70-80% train and 20-30% test; for classification use a stratified split so class proportions stay the same, and for time-based data split by time (train on the past, test on the future).
Features, splits, transformations
How you prepare and split data decides whether your results can be trusted.
Training accuracy versus test accuracy, run
I ran this with Python 3, numpy 2.5.3 and scikit-learn 1.9.1, using fixed random seeds. On the bundled breast cancer dataset (569 rows, 30 features), an unrestricted decision tree scores 1.0 on its training data but 0.902 on the held-out test set. The test number is the one that matters.
from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import train_test_split
from sklearn.tree import DecisionTreeClassifier
X, y = load_breast_cancer(return_X_y=True)
print("dataset:", X.shape[0], "rows,", X.shape[1], "features")
X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.25, random_state=0, stratify=y)
model = DecisionTreeClassifier(random_state=0).fit(X_tr, y_tr)
print("accuracy on training data:", round(model.score(X_tr, y_tr), 3))
print("accuracy on unseen test :", round(model.score(X_te, y_te), 3))
Output:
dataset: 569 rows, 30 features accuracy on training data: 1.0 accuracy on unseen test : 0.902
Fix the random seed
Set random_state so splits are reproducible and comparisons between models are fair.
त्वरित जाँच: Why is training accuracy a poor measure of model quality?
- Training data is always wrong
- The model may have memorised the training rows
- It is always lower than test accuracy
- It cannot be computed
Answer
The model may have memorised the training rows — Judge models on unseen data.