Lesson 24 / 25
Saving and Loading Models
joblib and versions.
Persist the whole pipeline
Save the fitted pipeline (not just the model) with joblib.dump, together with metadata: scikit-learn version, training date, data snapshot and metrics. Load it with the same library versions: pickled models are not guaranteed to work across scikit-learn versions. Only load model files from trusted sources, because unpickling can execute code; for cross-language serving, consider ONNX or skops.
Saving, reloading and inspecting a forest, run
I ran this with Python 3.12.3 and scikit-learn 1.9.1 on a dataset bundled with scikit-learn (no download). Predictions are identical after reload, the version is stored with the model, and feature importances show petal measurements dominate.
import os, tempfile
import joblib, sklearn
from sklearn.datasets import load_iris
from sklearn.ensemble import RandomForestClassifier
X, y = load_iris(return_X_y=True)
model = RandomForestClassifier(n_estimators=50, random_state=0).fit(X, y)
path = os.path.join(tempfile.mkdtemp(), "iris_rf.joblib")
joblib.dump({"model": model, "sklearn_version": sklearn.__version__}, path)
bundle = joblib.load(path)
same = (bundle["model"].predict(X) == model.predict(X)).all()
print("predictions identical after reload:", same)
print("saved with scikit-learn", bundle["sklearn_version"])
for name, imp in zip(load_iris().feature_names, model.feature_importances_):
print(f"{name:18} {imp:.3f}")
Output:
predictions identical after reload: True saved with scikit-learn 1.9.1 sepal length (cm) 0.083 sepal width (cm) 0.027 petal length (cm) 0.442 petal width (cm) 0.448
Pin versions in requirements
Record exact numpy, pandas and scikit-learn versions next to every saved model.
Quick check: Why should you never load a pickled model from an untrusted source?
- It deletes training data
- It is always too large
- Unpickling can execute arbitrary code
- It changes the model's accuracy
Answer
Unpickling can execute arbitrary code — Pickle is not a safe format for untrusted input.