पाठ 20 / 25
Clustering With k-Means
Group similar points around centres.
Centres, assignments, repeat
k-means picks k centres, assigns each point to its nearest centre, moves each centre to the mean of its points, and repeats until stable. You must choose k: look at inertia (within-cluster spread, which always falls as k grows, so look for an elbow) and the silhouette score (how well points fit their own cluster versus the next one, higher is better). k-means prefers round, similar-sized clusters and needs scaled features. Clusters are a tool for exploration and segmentation; give them names only after inspecting them.
Structure without labels
Group, compress and spot the unusual when there are no labels.
Choosing k with inertia and silhouette, run
I ran this with Python 3, numpy 2.5.3 and scikit-learn 1.9.1, using fixed random seeds. On 600 synthetic points from 4 groups, inertia keeps falling as k grows (9,253 at k = 2 to 980 at k = 6), but the silhouette score peaks clearly at k = 4 (0.728).
from sklearn.cluster import KMeans
from sklearn.datasets import make_blobs
from sklearn.metrics import silhouette_score
X, _ = make_blobs(n_samples=600, centers=4, cluster_std=1.0, random_state=3)
print(" k | inertia | silhouette")
for k in [2, 3, 4, 5, 6]:
km = KMeans(n_clusters=k, n_init=10, random_state=0).fit(X)
print(f"{k:>2} | {km.inertia_:>7.0f} | {silhouette_score(X, km.labels_):.3f}")
Output:
k | inertia | silhouette 2 | 9253 | 0.636 3 | 4374 | 0.608 4 | 1205 | 0.728 5 | 1086 | 0.624 6 | 980 | 0.521
Describe each cluster
Print the average feature values per cluster; a cluster you cannot describe is not useful to the business.
त्वरित जाँच: Why not choose k simply by the lowest inertia?
- Inertia is always zero
- Inertia always decreases as k increases
- Inertia increases with k
- k-means has no inertia
Answer
Inertia always decreases as k increases — Use an elbow or silhouette instead.