पाठ 18 / 25

Hyperparameter Search

Choose settings with cross-validation, then test once.

Grid and random search

Hyperparameters are settings you choose rather than learn: tree depth, number of trees, regularisation strength, k in k-NN. Grid search tries every combination from lists you give; random search samples combinations and often finds good settings faster when there are many. Both should evaluate with cross-validation on the training data only. After choosing, evaluate once on the untouched test set; repeatedly checking the test set while tuning turns it into training data and inflates results.

Grid search with cross-validation, run

I ran this with Python 3, numpy 2.5.3 and scikit-learn 1.9.1, using fixed random seeds. Eight random forest settings are compared with 5-fold CV on the training split. The best (depth 3, min_samples_leaf 5, 50 trees) scores 0.958 in CV and 0.93 on the test set, a little lower, as is common.

from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import GridSearchCV, train_test_split
from sklearn.ensemble import RandomForestClassifier
X, y = load_breast_cancer(return_X_y=True)
X_tr, X_te, y_tr, y_te = train_test_split(X, y, random_state=0, stratify=y)
grid = GridSearchCV(RandomForestClassifier(random_state=0),
                    {"n_estimators": [50, 200], "max_depth": [3, None], "min_samples_leaf": [1, 5]}, cv=5)
grid.fit(X_tr, y_tr)
print("best params :", grid.best_params_)
print("best CV acc :", round(grid.best_score_, 3))
print("test acc    :", round(grid.score(X_te, y_te), 3), "(test set used once, at the end)")

Output:

best params : {'max_depth': 3, 'min_samples_leaf': 5, 'n_estimators': 50}
best CV acc : 0.958
test acc    : 0.93 (test set used once, at the end)

Expect a small drop on test

The best CV score is slightly optimistic because it was chosen as the maximum; a modest drop on the test set is normal.

त्वरित जाँच: Why evaluate on the test set only once, at the end?

  • Tuning against it would make the final score optimistic
  • The test set expires
  • It is too large to use twice
  • Cross-validation requires it
Answer

Tuning against it would make the final score optimistic — Keep the final exam unseen.