Lesson 21 / 25
Dimensionality Reduction With PCA
Keep most of the information in fewer columns.
New axes ordered by variance
Principal component analysis (PCA) finds new axes (components) that capture the most variance in the data, ordered from most to least. Keeping the first few components compresses the data while preserving most of its variation, which helps visualisation (2 components for a scatter plot), speeds up models and reduces noise. Components are combinations of the original features, so they are harder to interpret. Scale features first, and choose the number of components from the cumulative explained variance.
Variance kept by components on 64-pixel digits, run
I ran this with Python 3, numpy 2.5.3 and scikit-learn 1.9.1, using fixed random seeds. The bundled digits dataset has 64 pixel features. Two components keep 29% of the variance, 10 keep 74%, 20 keep 89% and 40 keep 99%.
from sklearn.datasets import load_digits
from sklearn.decomposition import PCA
X, _ = load_digits(return_X_y=True)
pca = PCA().fit(X)
cum = pca.explained_variance_ratio_.cumsum()
print("original dimensions:", X.shape[1])
for n in [2, 10, 20, 40]:
print(f"{n:>2} components keep {cum[n - 1]:.0%} of the variance")
Output:
original dimensions: 64 2 components keep 29% of the variance 10 components keep 74% of the variance 20 components keep 89% of the variance 40 components keep 99% of the variance
Fit PCA inside the pipeline
Like scaling, fit PCA on training data only and apply it to test data through a pipeline.
Quick check: What does PCA order its components by?
- Alphabetical feature name
- The amount of variance each captures
- Training time
- Class labels
Answer
The amount of variance each captures — The first component captures the most variance.