पाठ 6 / 25
Scaling and Encoding Features
Turn raw columns into numbers models can use well.
Comparable numbers and one-hot categories
Many algorithms need numeric input on comparable scales. Scaling (for example standardising each feature to mean 0 and standard deviation 1) matters for distance-based models like k-nearest neighbours and for regularised linear models; tree models do not need it. Categorical columns such as city are usually one-hot encoded: one 0/1 column per category. Wrap these steps with the model in a pipeline so the same transformations are applied consistently in training, cross-validation and prediction.
Scaling changes k-NN accuracy, run
I ran this with Python 3, numpy 2.5.3 and scikit-learn 1.9.1, using fixed random seeds. On the bundled wine dataset, feature maxima range from 0.66 to 1,680, so the largest features dominate distances. k-NN scores 0.691 in 5-fold cross-validation without scaling and 0.949 with a StandardScaler in a pipeline.
from sklearn.datasets import load_wine
from sklearn.model_selection import cross_val_score
from sklearn.neighbors import KNeighborsClassifier
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
X, y = load_wine(return_X_y=True)
print("feature ranges: smallest max", round(X.max(axis=0).min(), 2), "| largest max", round(X.max(axis=0).max(), 1))
raw = cross_val_score(KNeighborsClassifier(), X, y, cv=5).mean()
scaled = cross_val_score(make_pipeline(StandardScaler(), KNeighborsClassifier()), X, y, cv=5).mean()
print(f"k-NN without scaling: {raw:.3f}")
print(f"k-NN with scaling : {scaled:.3f}")
Output:
feature ranges: smallest max 0.66 | largest max 1680.0 k-NN without scaling: 0.691 k-NN with scaling : 0.949
One-hot encoding and standardising, run
I ran this with Python 3, numpy 2.5.3 and scikit-learn 1.9.1, using fixed random seeds. Four rows with a city and an area: the city becomes three 0/1 columns, and the area becomes standardised values around zero.
import numpy as np
from sklearn.preprocessing import OneHotEncoder, StandardScaler
cities = np.array([["Pune"], ["Delhi"], ["Pune"], ["Mumbai"]])
area = np.array([[650], [1200], [900], [400]])
enc = OneHotEncoder().fit(cities)
print("one-hot columns:", list(enc.get_feature_names_out(["city"])))
print(enc.transform(cities).toarray().astype(int))
print("scaled area:", StandardScaler().fit_transform(area).ravel().round(2))
Output:
one-hot columns: ['city_Delhi', 'city_Mumbai', 'city_Pune'] [[0 0 1] [1 0 0] [0 0 1] [0 1 0]] scaled area: [-0.46 1.39 0.38 -1.31]
त्वरित जाँच: Which model is most sensitive to feature scaling?
- A model that always predicts the mean
- Decision tree
- Random forest
- k-nearest neighbours
Answer
k-nearest neighbours — Distances are dominated by large-scale features.