Machine Learning Basics
Learn machine learning by doing it: data splits, regression, classification, metrics, overfitting, cross-validation, ensembles and unsupervised learning, with every scikit-learn example run on real bundled datasets.
What you'll learn
- Explain what machine learning is, the main problem types, and when ML beats hand-written rules.
- Prepare data correctly: features and labels, train/test splits, scaling and encoding, and avoiding leakage.
- Train and evaluate regression and classification models with the right metrics and a baseline.
- Recognise overfitting and underfitting and use cross-validation and regularisation to control them.
- Use ensembles, hyperparameter search and permutation importance, and apply clustering, PCA and anomaly detection.
- Run a small end-to-end project responsibly, including fairness and monitoring considerations.
Syllabus
What Machine Learning Is
- Rules From Data Instead of Rules by Hand
- Supervised, Unsupervised and Reinforcement Learning
- The Machine Learning Workflow
Preparing Data
Regression: Predicting Numbers
Classification: Predicting Categories
- Logistic Regression and Probabilities
- Decision Trees and k-Nearest Neighbours
- Confusion Matrix, Precision and Recall
- Choosing the Decision Threshold
Generalisation: Overfitting and Validation
Ensembles, Tuning and Interpretation
- Random Forests and Gradient Boosting
- Hyperparameter Search
- Which Features Matter? Permutation Importance