Lesson 23 / 25

An End-to-End Mini Project

Every step in one pipeline.

Pipeline, validation, final test

A clean scikit-learn project: load data; split off a test set (stratified); build a pipeline of preprocessing (imputation, scaling, encoding) and a model; compare a baseline and a few models with cross-validation; tune the best with a search; evaluate once on the test set; analyse errors (which cases fail, and why); save the fitted pipeline with its version and metrics. Pipelines prevent leakage and make the exact same steps run in production.

Project, responsibility, checklist

Combine the steps into a reliable project and think about its impact on people.

Three ideas: end-to-end project, responsible ML, checklist.
Figure 8.1 — Project, responsibility and checklist.

A project skeleton (sketch)

Structure to copy; fill in your data and models. Not run here.

X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.2, stratify=y, random_state=0)
pre = ColumnTransformer([("num", make_pipeline(SimpleImputer(), StandardScaler()), num_cols),
                         ("cat", OneHotEncoder(handle_unknown="ignore"), cat_cols)])
candidates = {"baseline": DummyClassifier(), "logreg": LogisticRegression(max_iter=1000),
              "forest": RandomForestClassifier(random_state=0)}
for name, model in candidates.items():
    print(name, cross_val_score(make_pipeline(pre, model), X_tr, y_tr, cv=5, scoring="f1").mean())
# tune the best with GridSearchCV, then: best.score(X_te, y_te) once; joblib.dump(best, "model-v1.joblib")

Save metrics with the model

Store the model file together with its training date, data version, CV and test scores, so you can compare future versions.

Quick check: Why put preprocessing inside a pipeline?

  • The same steps are fit on training folds only and reused identically in production
  • Pipelines make models larger
  • It is required for decision trees only
  • It removes the need for a test set
Answer

The same steps are fit on training folds only and reused identically in production — Pipelines prevent leakage and mismatch.