# The Estimator API — NumPy / Pandas / scikit-learn

Source: https://www.skillbyai.com/en/numpy-pandas-sklearn/s-fit

> Split, fit, score, predict.

## One interface for every model

In scikit-learn, features `X` are a 2-D array or DataFrame and targets `y` a 1-D array. Every estimator follows the same pattern: create it with hyperparameters, `fit(X_train, y_train)`, then `predict`, `predict_proba` or `score`. Hold out a **test set** with `train_test_split` (use `stratify=y` for classification so class proportions match and `random_state` for reproducibility) and keep it untouched until the final evaluation.

## fit, predict, evaluate

scikit-learn gives every model the same API, so the workflow stays the same while models change.

![Three ideas: the estimator API, classification metrics, regression and baselines.](assets/figures/numpy-pandas-sklearn/section-6-map.svg) — Figure 6.1 — Fit, evaluate and compare to a baseline.

## Training a classifier on Iris, run

I ran this with Python 3.12.3 and scikit-learn 1.9.1 on a dataset bundled with scikit-learn (no download). A 25% stratified test split; logistic regression scores 0.955 on training data and 1.0 on this small test set of 38 flowers, which is a lucky split rather than a promise of perfection.

```python
from sklearn.datasets import load_iris
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split

X, y = load_iris(return_X_y=True)
print(X.shape, y.shape)
X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.25, random_state=0, stratify=y)
model = LogisticRegression(max_iter=1000)
model.fit(X_train, y_train)
print("train accuracy:", round(model.score(X_train, y_train), 3))
print("test accuracy :", round(model.score(X_test, y_test), 3))
print(model.predict(X_test[:5]), y_test[:5])
print(model.predict_proba(X_test[:1]).round(3))
```

Output:

```
(150, 4) (150,)
train accuracy: 0.955
test accuracy : 1.0
[0 0 0 0 1] [0 0 0 0 1]
[[0.968 0.032 0.   ]]
```

## Mock exam and final exam

Training data is the textbook, the test set is the final exam. Looking at exam questions while studying makes the grade meaningless.

**Quiz:** Why use stratify=y in train_test_split?

- [ ] To remove outliers
- [ ] To make training faster
- [x] To keep class proportions the same in train and test sets
- [ ] To scale features

*Answer:* To keep class proportions the same in train and test sets. Stratification preserves class balance.
