पाठ 20 / 25

Pipelines and ColumnTransformer

Preprocessing and model as one object.

Different steps for different columns

Real data mixes numeric and categorical columns. A ColumnTransformer applies imputation and scaling to numeric columns and one-hot encoding to categories; a Pipeline chains it with a model. The whole pipeline is fitted with one fit call, applies identical transformations at prediction time, and can be cross-validated and saved as one object. OneHotEncoder(handle_unknown="ignore") tolerates categories unseen during training.

Repeatable and fair

Pipelines bundle preprocessing with models; cross-validation and grid search choose models fairly.

Three ideas: pipelines, cross-validation, hyperparameter search.
Figure 7.1 — Pipelines, cross-validation and search.

A churn model on mixed columns, run

I ran this with Python 3.12.3 and scikit-learn 1.9.1 on a dataset bundled with scikit-learn (no download). This uses a small invented churn table rather than a bundled dataset: the pipeline imputes the missing age, scales numbers, one-hot encodes plans and predicts for new customers, including an unseen "enterprise" plan.

import pandas as pd
from sklearn.compose import ColumnTransformer
from sklearn.impute import SimpleImputer
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import Pipeline, make_pipeline
from sklearn.preprocessing import OneHotEncoder, StandardScaler

df = pd.DataFrame({
    "age": [25, 40, None, 35, 52, 23, 44, 31],
    "plan": ["basic", "pro", "pro", "basic", "team", "basic", "team", "pro"],
    "monthly_spend": [10, 45, 50, 12, 120, 8, 110, 40],
    "churned": [1, 0, 0, 1, 0, 1, 0, 0],
})
X, y = df.drop(columns="churned"), df["churned"]
pre = ColumnTransformer([
    ("num", make_pipeline(SimpleImputer(strategy="median"), StandardScaler()), ["age", "monthly_spend"]),
    ("cat", OneHotEncoder(handle_unknown="ignore"), ["plan"]),
])
model = Pipeline([("pre", pre), ("clf", LogisticRegression())]).fit(X, y)
print(model.named_steps["pre"].get_feature_names_out().tolist())
new = pd.DataFrame({"age": [29, None], "plan": ["basic", "enterprise"], "monthly_spend": [9, 200]})
print(model.predict(new))

Output:

['num__age', 'num__monthly_spend', 'cat__plan_basic', 'cat__plan_pro', 'cat__plan_team']
[1 0]

Feed DataFrames to pipelines

Selecting columns by name in a ColumnTransformer is clearer and safer than relying on column positions.

त्वरित जाँच: What does handle_unknown="ignore" do in OneHotEncoder?

  • Drops the whole row
  • Encodes unseen categories as all zeros instead of raising an error
  • Adds a new column at predict time
  • Raises a warning and stops
Answer

Encodes unseen categories as all zeros instead of raising an error — Robust to new categories.