# Retraining Pipelines and Triggers — MLOps

Source: https://www.skillbyai.com/en/mlops/t-pipeline

> Automate the path from new data to a gated candidate.

## Scheduled or triggered, always gated

A **retraining pipeline** runs the same steps every time: pull a validated data snapshot, build features, train, evaluate against the champion, register the candidate and (if gates pass) promote or request approval. Trigger it on a **schedule**, on **drift or performance alerts**, or when enough new labels arrive. Orchestrators such as Airflow, Kubeflow Pipelines, Prefect or cloud pipeline services run and record these steps. Never let retraining bypass the evaluation gate: automatic retraining on bad data is a fast way to ship a bad model.

## A retraining pipeline outline

Each step logs to the tracker.

```text
trigger: monthly schedule OR PSI > 0.2 on key features OR 5,000 new labels
 1 snapshot data (versioned) -> validate (schema, ranges, volume)
 2 build features (same code as serving)
 3 train with tracked params + data fingerprint
 4 evaluate candidate vs champion on recent holdout + slices
 5 register candidate version
 6 if gates pass -> shadow -> canary -> move champion alias
 7 else -> report + alert the owning team
```

## Keep a human in the loop for high-risk models

For credit, health or safety models, require a named approver before promotion even when gates pass.

**Quiz:** Why must automated retraining still pass an evaluation gate?

- [ ] Gates make training faster
- [x] Training on bad or shifted data can produce a worse model
- [ ] Retrained models are always better
- [ ] It is required by the GPU

*Answer:* Training on bad or shifted data can produce a worse model. Automation needs guardrails.
