Lesson 17 / 25
The Estimator API
Split, fit, score, predict.
One interface for every model
In scikit-learn, features X are a 2-D array or DataFrame and targets y a 1-D array. Every estimator follows the same pattern: create it with hyperparameters, fit(X_train, y_train), then predict, predict_proba or score. Hold out a test set with train_test_split (use stratify=y for classification so class proportions match and random_state for reproducibility) and keep it untouched until the final evaluation.
fit, predict, evaluate
scikit-learn gives every model the same API, so the workflow stays the same while models change.
Training a classifier on Iris, run
I ran this with Python 3.12.3 and scikit-learn 1.9.1 on a dataset bundled with scikit-learn (no download). A 25% stratified test split; logistic regression scores 0.955 on training data and 1.0 on this small test set of 38 flowers, which is a lucky split rather than a promise of perfection.
from sklearn.datasets import load_iris
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split
X, y = load_iris(return_X_y=True)
print(X.shape, y.shape)
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.25, random_state=0, stratify=y)
model = LogisticRegression(max_iter=1000)
model.fit(X_train, y_train)
print("train accuracy:", round(model.score(X_train, y_train), 3))
print("test accuracy :", round(model.score(X_test, y_test), 3))
print(model.predict(X_test[:5]), y_test[:5])
print(model.predict_proba(X_test[:1]).round(3))
Output:
(150, 4) (150,) train accuracy: 0.955 test accuracy : 1.0 [0 0 0 0 1] [0 0 0 0 1] [[0.968 0.032 0. ]]
Mock exam and final exam
Training data is the textbook, the test set is the final exam. Looking at exam questions while studying makes the grade meaningless.
Quick check: Why use stratify=y in train_test_split?
- To remove outliers
- To make training faster
- To keep class proportions the same in train and test sets
- To scale features
Answer
To keep class proportions the same in train and test sets — Stratification preserves class balance.