# Scaling and Encoding Features — Machine Learning Basics

Source: https://www.skillbyai.com/en/machine-learning/d-prep

> Turn raw columns into numbers models can use well.

## Comparable numbers and one-hot categories

Many algorithms need numeric input on comparable scales. **Scaling** (for example standardising each feature to mean 0 and standard deviation 1) matters for distance-based models like k-nearest neighbours and for regularised linear models; tree models do not need it. **Categorical** columns such as city are usually **one-hot encoded**: one 0/1 column per category. Wrap these steps with the model in a **pipeline** so the same transformations are applied consistently in training, cross-validation and prediction.

## Scaling changes k-NN accuracy, run

I ran this with Python 3, numpy 2.5.3 and scikit-learn 1.9.1, using fixed random seeds. On the bundled wine dataset, feature maxima range from 0.66 to 1,680, so the largest features dominate distances. k-NN scores 0.691 in 5-fold cross-validation without scaling and 0.949 with a StandardScaler in a pipeline.

```python
from sklearn.datasets import load_wine
from sklearn.model_selection import cross_val_score
from sklearn.neighbors import KNeighborsClassifier
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
X, y = load_wine(return_X_y=True)
print("feature ranges: smallest max", round(X.max(axis=0).min(), 2), "| largest max", round(X.max(axis=0).max(), 1))
raw = cross_val_score(KNeighborsClassifier(), X, y, cv=5).mean()
scaled = cross_val_score(make_pipeline(StandardScaler(), KNeighborsClassifier()), X, y, cv=5).mean()
print(f"k-NN without scaling: {raw:.3f}")
print(f"k-NN with scaling   : {scaled:.3f}")
```

Output:

```
feature ranges: smallest max 0.66 | largest max 1680.0
k-NN without scaling: 0.691
k-NN with scaling   : 0.949
```

## One-hot encoding and standardising, run

I ran this with Python 3, numpy 2.5.3 and scikit-learn 1.9.1, using fixed random seeds. Four rows with a city and an area: the city becomes three 0/1 columns, and the area becomes standardised values around zero.

```python
import numpy as np
from sklearn.preprocessing import OneHotEncoder, StandardScaler
cities = np.array([["Pune"], ["Delhi"], ["Pune"], ["Mumbai"]])
area = np.array([[650], [1200], [900], [400]])
enc = OneHotEncoder().fit(cities)
print("one-hot columns:", list(enc.get_feature_names_out(["city"])))
print(enc.transform(cities).toarray().astype(int))
print("scaled area:", StandardScaler().fit_transform(area).ravel().round(2))
```

Output:

```
one-hot columns: ['city_Delhi', 'city_Mumbai', 'city_Pune']
[[0 0 1]
 [1 0 0]
 [0 0 1]
 [0 1 0]]
scaled area: [-0.46  1.39  0.38 -1.31]
```

**Quiz:** Which model is most sensitive to feature scaling?

- [ ] A model that always predicts the mean
- [ ] Decision tree
- [ ] Random forest
- [x] k-nearest neighbours

*Answer:* k-nearest neighbours. Distances are dominated by large-scale features.
