पाठ 6 / 25

Packaging Models With Metadata

A model file alone is not enough.

Artefact plus a manifest

Package each model as an artefact plus metadata: the serialised pipeline (including preprocessing), library versions, expected input schema, model version, training data fingerprint, metrics and a checksum. Loading code should verify the checksum and the library versions before serving. Serialisation formats like pickle and joblib can execute code when loaded, so only load artefacts from trusted storage; consider formats such as ONNX or the MLflow model format for portability.

Saving a pipeline with a checksummed manifest, run

I ran this with Python 3, MLflow 3.16.1, scikit-learn 1.9.1, scipy 1.18.1 and numpy 2.5.3, using bundled or seeded synthetic data and a local SQLite tracking store. A scaler-plus-logistic-regression pipeline is saved with joblib next to a JSON manifest (name, version, scikit-learn version, feature count, checksum). On load, the checksum is verified and predictions match the original.

import hashlib, json, os, tempfile, joblib, sklearn
from sklearn.datasets import load_breast_cancer
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
X, y = load_breast_cancer(return_X_y=True)
pipe = make_pipeline(StandardScaler(), LogisticRegression()).fit(X, y)
d = tempfile.mkdtemp(); path = os.path.join(d, "model.joblib")
joblib.dump(pipe, path)
meta = {"name": "cancer-clf", "version": "1.4.0", "sklearn": sklearn.__version__,
        "features": int(X.shape[1]), "sha256": hashlib.sha256(open(path, "rb").read()).hexdigest()[:16]}
json.dump(meta, open(os.path.join(d, "model.json"), "w"))
loaded = joblib.load(path)
ok = hashlib.sha256(open(path, "rb").read()).hexdigest()[:16] == meta["sha256"]
print("metadata keys:", sorted(meta))
print("checksum verified:", ok, "| same predictions:", (loaded.predict(X) == pipe.predict(X)).all())

Output:

metadata keys: ['features', 'name', 'sha256', 'sklearn', 'version']
checksum verified: True | same predictions: True

Never unpickle untrusted files

Loading a pickle or joblib file can run arbitrary code; restrict who can write to the model store.

त्वरित जाँच: Why include preprocessing in the packaged model?

  • To hide the model
  • To make the file smaller
  • Because metrics require it
  • So serving applies exactly the same transformations as training
Answer

So serving applies exactly the same transformations as training — Packaging the whole pipeline prevents skew.