# Experiment Tracking With MLflow — MLOps

Source: https://www.skillbyai.com/en/mlops/r-track

> Every run: parameters, metrics, artefacts.

## Stop losing results in notebooks

An **experiment tracker** records each training run: parameters, metrics, code version, data reference, and artefacts such as the model file and plots. You can then compare runs, find the best one and see exactly how it was produced. **MLflow** is a widely used open-source option (others include Weights & Biases and cloud ML platforms). Log automatically from training code, not by hand, and include the data version and git commit as tags.

## Record everything that made the model

Tracking runs and fingerprinting inputs lets you explain, reproduce and compare models.

![Three ideas: tracking, reproducibility, packaging.](assets/figures/mlops/section-2-map.svg) — Figure 2.1 — Tracking, reproducibility and packaging.

## Logging three runs and ranking them, run

I ran this with Python 3, MLflow 3.16.1, scikit-learn 1.9.1, scipy 1.18.1 and numpy 2.5.3, using bundled or seeded synthetic data and a local SQLite tracking store. Three random forest runs with different max_depth are logged to a local SQLite-backed MLflow store, then queried in order: unlimited depth scores 0.9631 in cross-validation, depth 4 0.9578 and depth 2 0.9491.

```python
import os, tempfile, logging
os.environ["MLFLOW_DISABLE_AGENT_HINT"] = "1"; logging.getLogger("mlflow").setLevel(logging.ERROR)
import mlflow
from sklearn.datasets import load_breast_cancer
from sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import cross_val_score
mlflow.set_tracking_uri("sqlite:///" + os.path.join(tempfile.mkdtemp(), "mlflow.db"))
mlflow.set_experiment("cancer-classifier")
X, y = load_breast_cancer(return_X_y=True)
for depth in [2, 4, None]:
    with mlflow.start_run(run_name=f"rf-depth-{depth}"):
        params = {"n_estimators": 100, "max_depth": depth, "random_state": 0}
        score = cross_val_score(RandomForestClassifier(**params), X, y, cv=5).mean()
        mlflow.log_params(params)
        mlflow.log_metric("cv_accuracy", score)
runs = mlflow.search_runs(order_by=["metrics.cv_accuracy DESC"])
print(runs[["tags.mlflow.runName", "params.max_depth", "metrics.cv_accuracy"]].round(4).to_string(index=False))
```

Output:

```
tags.mlflow.runName params.max_depth  metrics.cv_accuracy
      rf-depth-None             None               0.9631
         rf-depth-4                4               0.9578
         rf-depth-2                2               0.9491
```

## Tag runs with git commit and data version

A metric without the code and data that produced it cannot be reproduced or trusted later.

**Quiz:** What should an experiment tracker record for each run?

- [ ] Only the final accuracy
- [x] Parameters, metrics, artefacts, and code and data versions
- [ ] Only the model file name
- [ ] Nothing until deployment

*Answer:* Parameters, metrics, artefacts, and code and data versions. Complete records make results reproducible.
