पाठ 4 / 25

Experiment Tracking With MLflow

Every run: parameters, metrics, artefacts.

Stop losing results in notebooks

An experiment tracker records each training run: parameters, metrics, code version, data reference, and artefacts such as the model file and plots. You can then compare runs, find the best one and see exactly how it was produced. MLflow is a widely used open-source option (others include Weights & Biases and cloud ML platforms). Log automatically from training code, not by hand, and include the data version and git commit as tags.

Record everything that made the model

Tracking runs and fingerprinting inputs lets you explain, reproduce and compare models.

Three ideas: tracking, reproducibility, packaging.
Figure 2.1 — Tracking, reproducibility and packaging.

Logging three runs and ranking them, run

I ran this with Python 3, MLflow 3.16.1, scikit-learn 1.9.1, scipy 1.18.1 and numpy 2.5.3, using bundled or seeded synthetic data and a local SQLite tracking store. Three random forest runs with different max_depth are logged to a local SQLite-backed MLflow store, then queried in order: unlimited depth scores 0.9631 in cross-validation, depth 4 0.9578 and depth 2 0.9491.

import os, tempfile, logging
os.environ["MLFLOW_DISABLE_AGENT_HINT"] = "1"; logging.getLogger("mlflow").setLevel(logging.ERROR)
import mlflow
from sklearn.datasets import load_breast_cancer
from sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import cross_val_score
mlflow.set_tracking_uri("sqlite:///" + os.path.join(tempfile.mkdtemp(), "mlflow.db"))
mlflow.set_experiment("cancer-classifier")
X, y = load_breast_cancer(return_X_y=True)
for depth in [2, 4, None]:
    with mlflow.start_run(run_name=f"rf-depth-{depth}"):
        params = {"n_estimators": 100, "max_depth": depth, "random_state": 0}
        score = cross_val_score(RandomForestClassifier(**params), X, y, cv=5).mean()
        mlflow.log_params(params)
        mlflow.log_metric("cv_accuracy", score)
runs = mlflow.search_runs(order_by=["metrics.cv_accuracy DESC"])
print(runs[["tags.mlflow.runName", "params.max_depth", "metrics.cv_accuracy"]].round(4).to_string(index=False))

Output:

tags.mlflow.runName params.max_depth  metrics.cv_accuracy
      rf-depth-None             None               0.9631
         rf-depth-4                4               0.9578
         rf-depth-2                2               0.9491

Tag runs with git commit and data version

A metric without the code and data that produced it cannot be reproduced or trusted later.

त्वरित जाँच: What should an experiment tracker record for each run?

  • Only the final accuracy
  • Parameters, metrics, artefacts, and code and data versions
  • Only the model file name
  • Nothing until deployment
Answer

Parameters, metrics, artefacts, and code and data versions — Complete records make results reproducible.