# A Data Science Workflow Checklist — NumPy / Pandas / scikit-learn

Source: https://www.skillbyai.com/en/numpy-pandas-sklearn/x-check

> Review before trusting results.

## Questions to ask

Are dtypes correct and missing values handled deliberately? Is code vectorised rather than looping? Are joins validated and row counts checked? Is the test set held out from all decisions? Is every preprocessing step inside a pipeline? Are results cross-validated with spreads reported and compared to a baseline? Do metrics match the business cost of errors? Are seeds, versions and data snapshots recorded with the saved model?

## The checklist

Use it in reviews.

```text
[ ] dtypes checked; numbers not stored as text; dates parsed
[ ] missing values counted and handled per column
[ ] vectorised NumPy/pandas instead of Python loops
[ ] merges validated (validate=, row counts, indicator)
[ ] loc for assignment (copy-on-write in pandas 3)
[ ] test set held out until the end; stratified/grouped/time splits
[ ] all preprocessing inside a Pipeline (no leakage)
[ ] cross-validated mean +- std vs a Dummy baseline
[ ] metrics match error costs (recall, precision, MAE)
[ ] seeds, library versions and data snapshot saved with the model
```

## Automate the boring checks

Assertions on shapes, dtypes and value ranges in your pipeline code catch data problems before they reach the model.

**Quiz:** Which item belongs on a data science checklist?

- [ ] Scale data before splitting
- [ ] Tune hyperparameters on the test set
- [x] Compare the model to a dummy baseline
- [ ] Ignore missing values

*Answer:* Compare the model to a dummy baseline. A baseline shows whether the model adds value.
