Lesson 23 / 25
An End-to-End Mini Project
Every step in one pipeline.
Pipeline, validation, final test
A clean scikit-learn project: load data; split off a test set (stratified); build a pipeline of preprocessing (imputation, scaling, encoding) and a model; compare a baseline and a few models with cross-validation; tune the best with a search; evaluate once on the test set; analyse errors (which cases fail, and why); save the fitted pipeline with its version and metrics. Pipelines prevent leakage and make the exact same steps run in production.
Project, responsibility, checklist
Combine the steps into a reliable project and think about its impact on people.
A project skeleton (sketch)
Structure to copy; fill in your data and models. Not run here.
X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.2, stratify=y, random_state=0)
pre = ColumnTransformer([("num", make_pipeline(SimpleImputer(), StandardScaler()), num_cols),
("cat", OneHotEncoder(handle_unknown="ignore"), cat_cols)])
candidates = {"baseline": DummyClassifier(), "logreg": LogisticRegression(max_iter=1000),
"forest": RandomForestClassifier(random_state=0)}
for name, model in candidates.items():
print(name, cross_val_score(make_pipeline(pre, model), X_tr, y_tr, cv=5, scoring="f1").mean())
# tune the best with GridSearchCV, then: best.score(X_te, y_te) once; joblib.dump(best, "model-v1.joblib")Save metrics with the model
Store the model file together with its training date, data version, CV and test scores, so you can compare future versions.
Quick check: Why put preprocessing inside a pipeline?
- The same steps are fit on training folds only and reused identically in production
- Pipelines make models larger
- It is required for decision trees only
- It removes the need for a test set
Answer
The same steps are fit on training folds only and reused identically in production — Pipelines prevent leakage and mismatch.