MPE StudioMath of Planet Earth
Machine Learning Toolkit · 05

K-Means in High Dimensions

When irrelevant coordinates overwhelm distance

K-means can fail even when one coordinate separates groups clearly. In many dimensions, accumulated noise dominates Euclidean distance and makes candidate neighbors look alike.

watchone idea
→
manipulateone example
→
leave withone intuition

The essential idea: distance concentration weakens the comparisons that distance-based clustering depends on.

Watch the concept

One Concept · One Example

K-Means in High Dimensions video thumbnail▶

K-Means in High Dimensions

Presented by Charlotte Moser

Watch on YouTube ↗

What to notice

The idea in 30 seconds

The curse appears in the distance

Squared distance

Every coordinate contributes a nonnegative squared difference.

‖x−y‖2=∑i=1d(xi−yi)2

Concentration

The standard deviation becomes small relative to the mean as dimension grows.

σμ=Cd→0
Explore

The same first two dimensions, different clustering

Keep the first two coordinates visible while increasing how many noisy dimensions k-means uses. Compare the data-generating groups with the resulting cluster labels.

Data-generating groupsDim 1 (signal)Dim 2 (noise)K-means using 2D distancesDim 1 (signal)Dim 2 (noise)
KEY TAKEAWAY

In high dimensions, irrelevant features can make distances nearly indistinguishable and undermine k-means.