पाठ 5 / 25
Reproducible Training
Same inputs, same model.
Pin data, code, config, environment and seeds
A run is reproducible when you can recreate the same model from recorded inputs: the exact data (a snapshot or versioned dataset, identified by a hash), the code (git commit), the configuration (hyperparameters), the environment (library versions, container image) and random seeds. Tools like DVC, lakeFS or table formats with time travel version data; containers and lock files pin environments. Some GPU operations are not bit-for-bit deterministic, so record tolerances for those cases.
Fingerprints and seeds, run
I ran this with Python 3, MLflow 3.16.1, scikit-learn 1.9.1, scipy 1.18.1 and numpy 2.5.3, using bundled or seeded synthetic data and a local SQLite tracking store. A SHA-256 fingerprint over data and config identifies the training inputs. Two forests with the same seed give identical predictions; a different seed changes them. Changing one data value by 0.001 changes the fingerprint completely.
import hashlib, json
import numpy as np
from sklearn.datasets import load_breast_cancer
from sklearn.ensemble import RandomForestClassifier
X, y = load_breast_cancer(return_X_y=True)
def fingerprint(X, y, config):
h = hashlib.sha256()
h.update(np.ascontiguousarray(X).tobytes()); h.update(np.ascontiguousarray(y).tobytes())
h.update(json.dumps(config, sort_keys=True).encode())
return h.hexdigest()[:12]
cfg = {"model": "rf", "n_estimators": 50, "random_state": 0}
a = RandomForestClassifier(n_estimators=50, random_state=0).fit(X, y).predict_proba(X[:3])[:, 1]
b = RandomForestClassifier(n_estimators=50, random_state=0).fit(X, y).predict_proba(X[:3])[:, 1]
c = RandomForestClassifier(n_estimators=50, random_state=1).fit(X, y).predict_proba(X[:3])[:, 1]
print("data+config fingerprint:", fingerprint(X, y, cfg))
print("same seed, same predictions :", np.array_equal(a, b))
print("different seed, same predictions:", np.array_equal(a, c))
X2 = X.copy(); X2[0, 0] += 0.001
print("one value changed -> fingerprint:", fingerprint(X2, y, cfg))
Output:
data+config fingerprint: 04cf6d426dd1 same seed, same predictions : True different seed, same predictions: False one value changed -> fingerprint: f0e22dcdc486
Store the fingerprint with the model
Log the data hash with every run and model version; it answers "was this trained on the corrected data?" in seconds.
त्वरित जाँच: Which is NOT needed to reproduce a training run?
- The data version
- The name of the person who watched the dashboard
- The code commit
- The random seed
Answer
The name of the person who watched the dashboard — Data, code, config, environment and seeds define a run.