# Pipelines and ColumnTransformer — NumPy / Pandas / scikit-learn

Source: https://www.skillbyai.com/en/numpy-pandas-sklearn/r-pipeline

> Preprocessing and model as one object.

## Different steps for different columns

Real data mixes numeric and categorical columns. A **ColumnTransformer** applies imputation and scaling to numeric columns and one-hot encoding to categories; a **Pipeline** chains it with a model. The whole pipeline is fitted with one `fit` call, applies identical transformations at prediction time, and can be cross-validated and saved as one object. `OneHotEncoder(handle_unknown="ignore")` tolerates categories unseen during training.

## Repeatable and fair

Pipelines bundle preprocessing with models; cross-validation and grid search choose models fairly.

![Three ideas: pipelines, cross-validation, hyperparameter search.](assets/figures/numpy-pandas-sklearn/section-7-map.svg) — Figure 7.1 — Pipelines, cross-validation and search.

## A churn model on mixed columns, run

I ran this with Python 3.12.3 and scikit-learn 1.9.1 on a dataset bundled with scikit-learn (no download). This uses a small invented churn table rather than a bundled dataset: the pipeline imputes the missing age, scales numbers, one-hot encodes plans and predicts for new customers, including an unseen "enterprise" plan.

```python
import pandas as pd
from sklearn.compose import ColumnTransformer
from sklearn.impute import SimpleImputer
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import Pipeline, make_pipeline
from sklearn.preprocessing import OneHotEncoder, StandardScaler

df = pd.DataFrame({
    "age": [25, 40, None, 35, 52, 23, 44, 31],
    "plan": ["basic", "pro", "pro", "basic", "team", "basic", "team", "pro"],
    "monthly_spend": [10, 45, 50, 12, 120, 8, 110, 40],
    "churned": [1, 0, 0, 1, 0, 1, 0, 0],
})
X, y = df.drop(columns="churned"), df["churned"]
pre = ColumnTransformer([
    ("num", make_pipeline(SimpleImputer(strategy="median"), StandardScaler()), ["age", "monthly_spend"]),
    ("cat", OneHotEncoder(handle_unknown="ignore"), ["plan"]),
])
model = Pipeline([("pre", pre), ("clf", LogisticRegression())]).fit(X, y)
print(model.named_steps["pre"].get_feature_names_out().tolist())
new = pd.DataFrame({"age": [29, None], "plan": ["basic", "enterprise"], "monthly_spend": [9, 200]})
print(model.predict(new))
```

Output:

```
['num__age', 'num__monthly_spend', 'cat__plan_basic', 'cat__plan_pro', 'cat__plan_team']
[1 0]
```

## Feed DataFrames to pipelines

Selecting columns by name in a ColumnTransformer is clearer and safer than relying on column positions.

**Quiz:** What does handle_unknown="ignore" do in OneHotEncoder?

- [ ] Drops the whole row
- [x] Encodes unseen categories as all zeros instead of raising an error
- [ ] Adds a new column at predict time
- [ ] Raises a warning and stops

*Answer:* Encodes unseen categories as all zeros instead of raising an error. Robust to new categories.
