SkillByAIOpen interactive version →

NumPy / Pandas / scikit-learn

The core Python data stack, hands on: fast arrays with NumPy, data wrangling with pandas, and honest machine learning with scikit-learn, with every example run on NumPy 2.5, pandas 3.0 and scikit-learn 1.9.

Start course →

What you'll learn

Syllabus

NumPy Arrays

  1. Creating and Shaping Arrays
  2. Vectorised Maths and Broadcasting
  3. Indexing, Masks, Views and Copies

Computing With NumPy

  1. Aggregation and Random Numbers
  2. Floating Point, Overflow and NaN
  3. Vectorisation Instead of Loops

pandas Fundamentals

  1. Series and DataFrames
  2. Selecting Rows and Columns
  3. Reading and Writing Data

Cleaning Data With pandas

  1. Handling Missing Values
  2. Types, Strings and Categories
  3. Copy-on-Write in pandas 3

Analysing Data With pandas

  1. Grouping and Aggregating
  2. Merging and Concatenating
  3. Pivot Tables and Reshaping
  4. Working With Time Series

Machine Learning With scikit-learn

  1. The Estimator API
  2. Classification Metrics
  3. Regression and Baselines

Pipelines and Model Selection

  1. Pipelines and ColumnTransformer
  2. Cross-Validation
  3. Hyperparameter Search

Avoiding Pitfalls and Shipping Models

  1. Data Leakage
  2. Saving and Loading Models
  3. A Data Science Workflow Checklist