Lesson 18 / 25
Hyperparameter Search
Choose settings with cross-validation, then test once.
Grid and random search
Hyperparameters are settings you choose rather than learn: tree depth, number of trees, regularisation strength, k in k-NN. Grid search tries every combination from lists you give; random search samples combinations and often finds good settings faster when there are many. Both should evaluate with cross-validation on the training data only. After choosing, evaluate once on the untouched test set; repeatedly checking the test set while tuning turns it into training data and inflates results.
Grid search with cross-validation, run
I ran this with Python 3, numpy 2.5.3 and scikit-learn 1.9.1, using fixed random seeds. Eight random forest settings are compared with 5-fold CV on the training split. The best (depth 3, min_samples_leaf 5, 50 trees) scores 0.958 in CV and 0.93 on the test set, a little lower, as is common.
from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import GridSearchCV, train_test_split
from sklearn.ensemble import RandomForestClassifier
X, y = load_breast_cancer(return_X_y=True)
X_tr, X_te, y_tr, y_te = train_test_split(X, y, random_state=0, stratify=y)
grid = GridSearchCV(RandomForestClassifier(random_state=0),
{"n_estimators": [50, 200], "max_depth": [3, None], "min_samples_leaf": [1, 5]}, cv=5)
grid.fit(X_tr, y_tr)
print("best params :", grid.best_params_)
print("best CV acc :", round(grid.best_score_, 3))
print("test acc :", round(grid.score(X_te, y_te), 3), "(test set used once, at the end)")
Output:
best params : {'max_depth': 3, 'min_samples_leaf': 5, 'n_estimators': 50}
best CV acc : 0.958
test acc : 0.93 (test set used once, at the end)Expect a small drop on test
The best CV score is slightly optimistic because it was chosen as the maximum; a modest drop on the test set is normal.
Quick check: Why evaluate on the test set only once, at the end?
- Tuning against it would make the final score optimistic
- The test set expires
- It is too large to use twice
- Cross-validation requires it
Answer
Tuning against it would make the final score optimistic — Keep the final exam unseen.