पाठ 21 / 25

Dimensionality Reduction With PCA

Keep most of the information in fewer columns.

New axes ordered by variance

Principal component analysis (PCA) finds new axes (components) that capture the most variance in the data, ordered from most to least. Keeping the first few components compresses the data while preserving most of its variation, which helps visualisation (2 components for a scatter plot), speeds up models and reduces noise. Components are combinations of the original features, so they are harder to interpret. Scale features first, and choose the number of components from the cumulative explained variance.

Variance kept by components on 64-pixel digits, run

I ran this with Python 3, numpy 2.5.3 and scikit-learn 1.9.1, using fixed random seeds. The bundled digits dataset has 64 pixel features. Two components keep 29% of the variance, 10 keep 74%, 20 keep 89% and 40 keep 99%.

from sklearn.datasets import load_digits
from sklearn.decomposition import PCA
X, _ = load_digits(return_X_y=True)
pca = PCA().fit(X)
cum = pca.explained_variance_ratio_.cumsum()
print("original dimensions:", X.shape[1])
for n in [2, 10, 20, 40]:
    print(f"{n:>2} components keep {cum[n - 1]:.0%} of the variance")

Output:

original dimensions: 64
 2 components keep 29% of the variance
10 components keep 74% of the variance
20 components keep 89% of the variance
40 components keep 99% of the variance

Fit PCA inside the pipeline

Like scaling, fit PCA on training data only and apply it to test data through a pipeline.

त्वरित जाँच: What does PCA order its components by?

  • Alphabetical feature name
  • The amount of variance each captures
  • Training time
  • Class labels
Answer

The amount of variance each captures — The first component captures the most variance.