# Which Features Matter? Permutation Importance — Machine Learning Basics

Source: https://www.skillbyai.com/en/machine-learning/e-importance

> Inspect what the model relies on.

## Shuffle a column, watch the score drop

**Permutation importance** measures how much a model's score drops when one feature's values are randomly shuffled, breaking its link to the target. Large drops mean the model relies on that feature. Compute it on held-out data. Interpret carefully: correlated features share importance (shuffling one may hurt little because another carries the same information), and importance shows what the model uses, not what causes the outcome. Use it to sanity-check models, spot leakage and explain behaviour.

## Top permutation importances, run

I ran this with Python 3, numpy 2.5.3 and scikit-learn 1.9.1, using fixed random seeds. For a random forest on the breast cancer test set, shuffling worst area lowers accuracy by about 0.010, and the concave-points features by 0.007 to 0.009. The drops are small because many features are correlated and can substitute for each other.

```python
from sklearn.datasets import load_breast_cancer
from sklearn.ensemble import RandomForestClassifier
from sklearn.inspection import permutation_importance
from sklearn.model_selection import train_test_split
data = load_breast_cancer()
X_tr, X_te, y_tr, y_te = train_test_split(data.data, data.target, random_state=0, stratify=data.target)
m = RandomForestClassifier(n_estimators=200, random_state=0).fit(X_tr, y_tr)
r = permutation_importance(m, X_te, y_te, n_repeats=10, random_state=0)
for i in r.importances_mean.argsort()[::-1][:4]:
    print(f"{data.feature_names[i]:<22} accuracy drop when shuffled {r.importances_mean[i]:.3f}")
```

Output:

```
worst area             accuracy drop when shuffled 0.010
worst concave points   accuracy drop when shuffled 0.009
mean concave points    accuracy drop when shuffled 0.007
mean concavity         accuracy drop when shuffled 0.007
```

## Investigate surprising top features

If an ID column or a timestamp ranks highly, suspect leakage before celebrating.

**Quiz:** What does high permutation importance tell you?

- [ ] The feature should be deleted
- [ ] The feature causes the outcome
- [ ] The feature has missing values
- [x] The model relies heavily on that feature for its predictions

*Answer:* The model relies heavily on that feature for its predictions. Importance is about the model, not causation.
