SkillByAIOpen interactive version →

MLOps

Take models from notebook to reliable production: experiment tracking, reproducibility, data validation, registries, CI gates, serving, shadow and canary releases, drift monitoring and retraining, with MLflow and scikit-learn examples run.

Start course →

What you'll learn

Syllabus

Why MLOps Exists

  1. How ML Systems Fail in Production
  2. The MLOps Lifecycle
  3. Maturity Levels and Roles

Experiment Tracking and Reproducibility

  1. Experiment Tracking With MLflow
  2. Reproducible Training
  3. Packaging Models With Metadata

Data in Production

  1. Data Validation
  2. Training-Serving Skew
  3. Feature Stores and Point-in-Time Correctness

Model Registry and Continuous Integration

  1. Model Registry, Versions and Aliases
  2. Testing ML Code, Data and Models
  3. Automated Evaluation Gates

Deploying Models Safely

  1. Batch, Online and Streaming Serving
  2. Serving Performance: Latency and Batching
  3. Shadow Deployment
  4. Canary Releases and A/B Tests

Monitoring Models in Production

  1. What to Monitor
  2. Detecting Data Drift
  3. Delayed Labels and Proxy Metrics

Retraining, Rollback and Incidents

  1. Concept Drift and Model Decay
  2. Retraining Pipelines and Triggers
  3. Rollback and Incident Response

Governance, LLMOps and a Checklist

  1. Governance: Lineage, Model Cards and Access
  2. How LLMOps Differs
  3. An MLOps Readiness Checklist