Lesson 3 / 25
The Machine Learning Workflow
From question to monitored model.
Eight steps, mostly not modelling
A typical project: (1) define the question and a success metric; (2) collect and understand data; (3) clean it and build features; (4) split into training and test sets; (5) train a simple baseline, then better models; (6) evaluate on held-out data and analyse errors; (7) deploy; (8) monitor and retrain as data changes. Most time goes into data and evaluation, not into choosing algorithms. A fair evaluation on data the model has never seen is the heart of the whole process.
The workflow as a checklist
Keep it beside every project.
1 question + metric "predict churn within 30 days; measure recall at 20% precision"
2 data sources, size, label quality, time range
3 features cleaning, encoding, scaling
4 split train / validation / test, no leakage
5 baseline -> models dummy model first, then real models
6 evaluate held-out metrics + error analysis
7 deploy batch job or API
8 monitor drift, performance, retraining planStart with a baseline
A model that always predicts the average or the most common class tells you whether your real model is actually learning anything.
Quick check: Where does most of the effort in an ML project usually go?
- Buying GPUs
- Choosing a fancy algorithm
- Writing the user interface
- Data preparation and evaluation
Answer
Data preparation and evaluation — Good data and honest evaluation matter most.