Lesson 25 / 25
A Data Science Workflow Checklist
Review before trusting results.
Questions to ask
Are dtypes correct and missing values handled deliberately? Is code vectorised rather than looping? Are joins validated and row counts checked? Is the test set held out from all decisions? Is every preprocessing step inside a pipeline? Are results cross-validated with spreads reported and compared to a baseline? Do metrics match the business cost of errors? Are seeds, versions and data snapshots recorded with the saved model?
The checklist
Use it in reviews.
[ ] dtypes checked; numbers not stored as text; dates parsed
[ ] missing values counted and handled per column
[ ] vectorised NumPy/pandas instead of Python loops
[ ] merges validated (validate=, row counts, indicator)
[ ] loc for assignment (copy-on-write in pandas 3)
[ ] test set held out until the end; stratified/grouped/time splits
[ ] all preprocessing inside a Pipeline (no leakage)
[ ] cross-validated mean +- std vs a Dummy baseline
[ ] metrics match error costs (recall, precision, MAE)
[ ] seeds, library versions and data snapshot saved with the modelAutomate the boring checks
Assertions on shapes, dtypes and value ranges in your pipeline code catch data problems before they reach the model.
Quick check: Which item belongs on a data science checklist?
- Scale data before splitting
- Tune hyperparameters on the test set
- Compare the model to a dummy baseline
- Ignore missing values
Answer
Compare the model to a dummy baseline — A baseline shows whether the model adds value.