# Automated Evaluation Gates — MLOps

Source: https://www.skillbyai.com/en/mlops/g-gate

> The candidate must beat the champion on agreed rules.

## Champion versus challenger

Before promotion, evaluate the **candidate** and the current **champion** on the same recent holdout data and apply written rules: overall metric not worse beyond a tolerance, critical metrics (for example recall on the dangerous class) not worse at all, slice floors met, latency and size within limits. If the rules pass, promotion can be automatic or require one approval; if not, the candidate is blocked with a report. Using the same data for both models makes the comparison fair.

## Champion versus candidate gate, run

I ran this with Python 3, MLflow 3.16.1, scikit-learn 1.9.1, scipy 1.18.1 and numpy 2.5.3, using bundled or seeded synthetic data and a local SQLite tracking store. On the same holdout, the random forest champion scores 0.953 accuracy and the gradient boosting candidate 0.936, with equal recall on malignant cases (0.922). The accuracy drop exceeds the 0.005 tolerance, so the candidate is not promoted.

```python
from sklearn.datasets import load_breast_cancer
from sklearn.ensemble import GradientBoostingClassifier, RandomForestClassifier
from sklearn.metrics import recall_score, accuracy_score
from sklearn.model_selection import train_test_split
X, y = load_breast_cancer(return_X_y=True)
X_tr, X_ho, y_tr, y_ho = train_test_split(X, y, test_size=0.3, random_state=0, stratify=y)
champion = RandomForestClassifier(n_estimators=200, random_state=0).fit(X_tr, y_tr)
candidate = GradientBoostingClassifier(random_state=0).fit(X_tr, y_tr)
def report(m):
    p = m.predict(X_ho)
    return {"accuracy": accuracy_score(y_ho, p), "recall_malignant": recall_score(y_ho, p, pos_label=0)}
c, n = report(champion), report(candidate)
for k in c:
    print(f"{k:<17} champion {c[k]:.3f}  candidate {n[k]:.3f}")
ok = n["accuracy"] >= c["accuracy"] - 0.005 and n["recall_malignant"] >= c["recall_malignant"]
print("promote candidate:", ok)
```

Output:

```
accuracy          champion 0.953  candidate 0.936
recall_malignant  champion 0.922  candidate 0.922
promote candidate: False
```

## Use a recent holdout

Evaluate on the most recent labelled data you have; it best reflects the traffic the model will face.

**Quiz:** Why compare candidate and champion on the same data?

- [ ] It doubles the dataset
- [x] So differences reflect the models, not the data
- [ ] It makes training faster
- [ ] It is required by Python

*Answer:* So differences reflect the models, not the data. Fair comparisons use identical evaluation sets.
