पाठ 6 / 25
Packaging Models With Metadata
A model file alone is not enough.
Artefact plus a manifest
Package each model as an artefact plus metadata: the serialised pipeline (including preprocessing), library versions, expected input schema, model version, training data fingerprint, metrics and a checksum. Loading code should verify the checksum and the library versions before serving. Serialisation formats like pickle and joblib can execute code when loaded, so only load artefacts from trusted storage; consider formats such as ONNX or the MLflow model format for portability.
Saving a pipeline with a checksummed manifest, run
I ran this with Python 3, MLflow 3.16.1, scikit-learn 1.9.1, scipy 1.18.1 and numpy 2.5.3, using bundled or seeded synthetic data and a local SQLite tracking store. A scaler-plus-logistic-regression pipeline is saved with joblib next to a JSON manifest (name, version, scikit-learn version, feature count, checksum). On load, the checksum is verified and predictions match the original.
import hashlib, json, os, tempfile, joblib, sklearn
from sklearn.datasets import load_breast_cancer
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
X, y = load_breast_cancer(return_X_y=True)
pipe = make_pipeline(StandardScaler(), LogisticRegression()).fit(X, y)
d = tempfile.mkdtemp(); path = os.path.join(d, "model.joblib")
joblib.dump(pipe, path)
meta = {"name": "cancer-clf", "version": "1.4.0", "sklearn": sklearn.__version__,
"features": int(X.shape[1]), "sha256": hashlib.sha256(open(path, "rb").read()).hexdigest()[:16]}
json.dump(meta, open(os.path.join(d, "model.json"), "w"))
loaded = joblib.load(path)
ok = hashlib.sha256(open(path, "rb").read()).hexdigest()[:16] == meta["sha256"]
print("metadata keys:", sorted(meta))
print("checksum verified:", ok, "| same predictions:", (loaded.predict(X) == pipe.predict(X)).all())
Output:
metadata keys: ['features', 'name', 'sha256', 'sklearn', 'version'] checksum verified: True | same predictions: True
Never unpickle untrusted files
Loading a pickle or joblib file can run arbitrary code; restrict who can write to the model store.
त्वरित जाँच: Why include preprocessing in the packaged model?
- To hide the model
- To make the file smaller
- Because metrics require it
- So serving applies exactly the same transformations as training
Answer
So serving applies exactly the same transformations as training — Packaging the whole pipeline prevents skew.