पाठ 22 / 25
Hyperparameter Search
GridSearchCV.
Search with cross-validation, test once
GridSearchCV tries every combination of hyperparameters with cross-validation on the training set, refits the best one on all training data, and exposes best_params_, best_score_ and cv_results_. Name parameters inside pipelines with step__param. Evaluate the chosen model once on the held-out test set; RandomizedSearchCV or HalvingGridSearchCV scale better to large search spaces.
Tuning an SVM in a pipeline, run
I ran this with Python 3.12.3 and scikit-learn 1.9.1 on a dataset bundled with scikit-learn (no download). Six combinations were tried with 5-fold CV; a linear kernel with C=0.1 won with 0.988 CV accuracy, and the untouched test set gave 0.958.
from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import GridSearchCV, train_test_split
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.svm import SVC
X, y = load_breast_cancer(return_X_y=True)
X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.25, random_state=0, stratify=y)
pipe = Pipeline([("scale", StandardScaler()), ("svc", SVC())])
grid = GridSearchCV(pipe, {"svc__C": [0.1, 1, 10], "svc__kernel": ["linear", "rbf"]}, cv=5)
grid.fit(X_tr, y_tr)
print("best params:", grid.best_params_)
print("best CV accuracy:", round(grid.best_score_, 3))
print("test accuracy :", round(grid.score(X_te, y_te), 3))
print("candidates tried:", len(grid.cv_results_["params"]))
Output:
best params: {'svc__C': 0.1, 'svc__kernel': 'linear'}
best CV accuracy: 0.988
test accuracy : 0.958
candidates tried: 6Expect test scores to be a bit lower
The best CV score is optimistically biased because it was selected; the test score is the honest estimate.
त्वरित जाँच: How do you name the C parameter of an SVC step called "svc" in a pipeline grid?
- pipeline_C
- C
- svc.C
- svc__C
Answer
svc__C — step__parameter.