Skip to content

Clinic 19

Cluster Stability

Compare separation and stability before committing to a customer segmentation. State the selection rule before reading the result.

Situation

A marketing team wants a segmentation of 40,000 customers with 22 scaled numeric features. The team agrees to choose the largest k with mean pairwise ARI at least 0.80, mean silhouette at least 0.25, and positive gap, then validate usefulness with campaign owners. These are this exercise's operating criteria, not universal statistical cutoffs.

Artifact Packet

The table is a fixed illustrative summary of 20 seed runs per k, each with explicitly specified KMeans(n_init=10). Ten is an explicit setting, not the current default. The runner displays these summaries; it does not execute the seed sweep.

k mean silhouette mean pairwise ARI gap statistic
2 0.28 0.94 0.18
3 0.30 0.82 0.21
4 0.34 0.71 0.15
5 0.31 0.56 0.09
6 0.27 0.48 0.04

Silhouette summarizes within-cluster cohesion relative to nearest-cluster separation. ARI compares pairwise assignments with a chance adjustment: 1 indicates identical partitions, chance-like agreement has expected value near 0, and negative values are possible. ARI 0.71 does not mean 29% of customers moved. To quantify reassignment, align cluster labels and inspect a cross-tab or an explicit reassignment rate. ARI definition

The gap statistic is E_reference[log W_k] − log W_k, where W_k measures within-cluster dispersion. It is not a difference between silhouettes. Its canonical one-standard-error rule chooses the smallest k with Gap(k) ≥ Gap(k+1) − s_(k+1). This packet lacks the reference-simulation standard errors required for that rule; use the separately declared exercise criteria here. Gap statistic explanation

A consensus matrix records pairwise co-clustering frequency. Its diagonal is always 1; inspect stable off-diagonal blocks of pairs assigned together, not a demand that every off-diagonal cell be low.

Decision Prompt

  1. Which k passes this exercise’s decision rule?
  2. What does ARI measure, and what does it not measure?
  3. What would you need to apply the canonical gap selection rule?
  4. What should you report if no candidate passes the agreed criteria?

Run The Clinic In Browser

The runner prints this fixed illustrative packet and recalculates any derived columns. It does not run a new training experiment. Edit its PACKET values to explore the decision.

Reference Reveal

Open after writing your note Choose **k=3 under the stated rule**. Both k=2 and k=3 pass; k=3 is the largest passing k. The silhouette winner k=4 fails the agreed stability floor. ARI does not directly tell you a customer-switching percentage. Before a real campaign, repeat preprocessing and clustering under resampling and relevant time windows, inspect segment sizes and meaning, and compare suitable alternative algorithms. Seed stability assesses optimization variability, not all sampling uncertainty. Agreement across algorithms supports a result but does not prove unique or naturally discrete customer groups. If no candidate passes, report that this representation and procedure have not produced a sufficiently stable, useful segmentation. Investigate features, distance choices, and the business objective. Failure of a particular sweep does not prove there are no clusters; a continuous score is an alternative only if it serves the task.

What To Do Next

  1. open Clustering and Low-Dimensional Views for the elbow/silhouette/gap workflow
  2. open Advanced Clustering and Dimensionality Reduction for hierarchical and density-based alternatives
  3. open Manifold Choice — the adjacent clinic on projection-method selection
  4. run the stability sweep on your own data; compare seed stability, resampling stability, and usefulness for the intended task