पाठ 12 / 25
Automated Evaluation Gates
The candidate must beat the champion on agreed rules.
Champion versus challenger
Before promotion, evaluate the candidate and the current champion on the same recent holdout data and apply written rules: overall metric not worse beyond a tolerance, critical metrics (for example recall on the dangerous class) not worse at all, slice floors met, latency and size within limits. If the rules pass, promotion can be automatic or require one approval; if not, the candidate is blocked with a report. Using the same data for both models makes the comparison fair.
Champion versus candidate gate, run
I ran this with Python 3, MLflow 3.16.1, scikit-learn 1.9.1, scipy 1.18.1 and numpy 2.5.3, using bundled or seeded synthetic data and a local SQLite tracking store. On the same holdout, the random forest champion scores 0.953 accuracy and the gradient boosting candidate 0.936, with equal recall on malignant cases (0.922). The accuracy drop exceeds the 0.005 tolerance, so the candidate is not promoted.
from sklearn.datasets import load_breast_cancer
from sklearn.ensemble import GradientBoostingClassifier, RandomForestClassifier
from sklearn.metrics import recall_score, accuracy_score
from sklearn.model_selection import train_test_split
X, y = load_breast_cancer(return_X_y=True)
X_tr, X_ho, y_tr, y_ho = train_test_split(X, y, test_size=0.3, random_state=0, stratify=y)
champion = RandomForestClassifier(n_estimators=200, random_state=0).fit(X_tr, y_tr)
candidate = GradientBoostingClassifier(random_state=0).fit(X_tr, y_tr)
def report(m):
p = m.predict(X_ho)
return {"accuracy": accuracy_score(y_ho, p), "recall_malignant": recall_score(y_ho, p, pos_label=0)}
c, n = report(champion), report(candidate)
for k in c:
print(f"{k:<17} champion {c[k]:.3f} candidate {n[k]:.3f}")
ok = n["accuracy"] >= c["accuracy"] - 0.005 and n["recall_malignant"] >= c["recall_malignant"]
print("promote candidate:", ok)
Output:
accuracy champion 0.953 candidate 0.936 recall_malignant champion 0.922 candidate 0.922 promote candidate: False
Use a recent holdout
Evaluate on the most recent labelled data you have; it best reflects the traffic the model will face.
त्वरित जाँच: Why compare candidate and champion on the same data?
- It doubles the dataset
- So differences reflect the models, not the data
- It makes training faster
- It is required by Python
Answer
So differences reflect the models, not the data — Fair comparisons use identical evaluation sets.