NumPy / Pandas / scikit-learn
The core Python data stack, hands on: fast arrays with NumPy, data wrangling with pandas, and honest machine learning with scikit-learn, with every example run on NumPy 2.5, pandas 3.0 and scikit-learn 1.9.
What you'll learn
- Create, reshape, index and broadcast NumPy arrays, and write vectorised code instead of Python loops.
- Load, select, clean and reshape tabular data with pandas, including missing values and types.
- Summarise data with groupby, merges, pivot tables and time-series resampling.
- Train and evaluate scikit-learn models with train/test splits, metrics and baselines.
- Build leak-free pipelines with preprocessing, cross-validation and hyperparameter search.
- Save models reproducibly and review a data-science workflow with a checklist.