Lesson 17 / 25

The Estimator API

Split, fit, score, predict.

One interface for every model

In scikit-learn, features X are a 2-D array or DataFrame and targets y a 1-D array. Every estimator follows the same pattern: create it with hyperparameters, fit(X_train, y_train), then predict, predict_proba or score. Hold out a test set with train_test_split (use stratify=y for classification so class proportions match and random_state for reproducibility) and keep it untouched until the final evaluation.

fit, predict, evaluate

scikit-learn gives every model the same API, so the workflow stays the same while models change.

Three ideas: the estimator API, classification metrics, regression and baselines.
Figure 6.1 — Fit, evaluate and compare to a baseline.

Training a classifier on Iris, run

I ran this with Python 3.12.3 and scikit-learn 1.9.1 on a dataset bundled with scikit-learn (no download). A 25% stratified test split; logistic regression scores 0.955 on training data and 1.0 on this small test set of 38 flowers, which is a lucky split rather than a promise of perfection.

from sklearn.datasets import load_iris
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split

X, y = load_iris(return_X_y=True)
print(X.shape, y.shape)
X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.25, random_state=0, stratify=y)
model = LogisticRegression(max_iter=1000)
model.fit(X_train, y_train)
print("train accuracy:", round(model.score(X_train, y_train), 3))
print("test accuracy :", round(model.score(X_test, y_test), 3))
print(model.predict(X_test[:5]), y_test[:5])
print(model.predict_proba(X_test[:1]).round(3))

Output:

(150, 4) (150,)
train accuracy: 0.955
test accuracy : 1.0
[0 0 0 0 1] [0 0 0 0 1]
[[0.968 0.032 0.   ]]

Mock exam and final exam

Training data is the textbook, the test set is the final exam. Looking at exam questions while studying makes the grade meaningless.

Quick check: Why use stratify=y in train_test_split?

  • To remove outliers
  • To make training faster
  • To keep class proportions the same in train and test sets
  • To scale features
Answer

To keep class proportions the same in train and test sets — Stratification preserves class balance.