Chapter 4 · Section 3 practice
K-Means: Assign, Average, Repeat
Work from question 1 to 10. Edit your Python box, click Run Code, review the result, then click Submit attempt. Reference answers unlock after all ten attempts are submitted. This records completion, not correctness. Each box runs independently. Progress is saved in this browser.
1. Count labels
Warm-up. Print the number of fitted labels.
Write your answer, then run it.
2. Read centre shape
Warm-up. Print the centre array shape.
Write your answer, then run it.
3. Count group sizes
Warm-up. Print sorted cluster membership counts.
Write your answer, then run it.
4. Calculate a manual centroid
Build your skills. Print the coordinate means for A,B,C.
Write your answer, then run it.
5. Read the objective
Build your skills. Print inertia rounded to three decimals.
Write your answer, then run it.
6. Assign a new nearby point
Build your skills. Predict the cluster of (1.5,1.5).
Write your answer, then run it.
7. Check a new point manually
Build your skills. Calculate squared distances from (1.5,1.5) to both centres; print the nearest-centre index.
Write your answer, then run it.
8. Verify group means
Challenge. Print whether group means equal the fitted centres in label order.
Write your answer, then run it.
9. Reconstruct WCSS
Challenge. Use each row’s assigned centre to calculate the total squared distance.
Write your answer, then run it.
10. Distinguish assignment and refitting
Challenge. Save centres, predict a new point, and print whether centres stayed unchanged.
Write your answer, then run it.
Reference answers are locked until all 10 attempts are submitted.
1. Count labels
import pandas as pd
import numpy as np
points = pd.DataFrame({"X1": [1, 1, 2, 7, 8, 8],
"X2": [1, 2, 1, 7, 7, 8]},
index=["A", "B", "C", "D", "E", "F"])
from sklearn.cluster import KMeans
km = KMeans(n_clusters=2, random_state=42, n_init=10)
points["Cluster"] = km.fit_predict(points[["X1", "X2"]])
print(len(km.labels_))Expected output
6There is one label per training point.
2. Read centre shape
import pandas as pd
import numpy as np
points = pd.DataFrame({"X1": [1, 1, 2, 7, 8, 8],
"X2": [1, 2, 1, 7, 7, 8]},
index=["A", "B", "C", "D", "E", "F"])
from sklearn.cluster import KMeans
km = KMeans(n_clusters=2, random_state=42, n_init=10)
points["Cluster"] = km.fit_predict(points[["X1", "X2"]])
print(km.cluster_centers_.shape)Expected output
(2, 2)K=2 and two features gives a 2 by 2 array.
3. Count group sizes
import pandas as pd
import numpy as np
points = pd.DataFrame({"X1": [1, 1, 2, 7, 8, 8],
"X2": [1, 2, 1, 7, 7, 8]},
index=["A", "B", "C", "D", "E", "F"])
from sklearn.cluster import KMeans
km = KMeans(n_clusters=2, random_state=42, n_init=10)
points["Cluster"] = km.fit_predict(points[["X1", "X2"]])
print(points["Cluster"].value_counts().sort_index())Expected output
Cluster
0 3
1 3
Name: count, dtype: int64Label values are arbitrary; counts describe the fitted partition.
4. Calculate a manual centroid
import pandas as pd
import numpy as np
points = pd.DataFrame({"X1": [1, 1, 2, 7, 8, 8],
"X2": [1, 2, 1, 7, 7, 8]},
index=["A", "B", "C", "D", "E", "F"])
from sklearn.cluster import KMeans
km = KMeans(n_clusters=2, random_state=42, n_init=10)
points["Cluster"] = km.fit_predict(points[["X1", "X2"]])
print(points.loc[["A", "B", "C"], ["X1", "X2"]].mean().round(3))Expected output
X1 1.333
X2 1.333
dtype: float64A centre is the mean of assigned coordinates.
5. Read the objective
import pandas as pd
import numpy as np
points = pd.DataFrame({"X1": [1, 1, 2, 7, 8, 8],
"X2": [1, 2, 1, 7, 7, 8]},
index=["A", "B", "C", "D", "E", "F"])
from sklearn.cluster import KMeans
km = KMeans(n_clusters=2, random_state=42, n_init=10)
points["Cluster"] = km.fit_predict(points[["X1", "X2"]])
print(round(km.inertia_, 3))Expected output
2.667Inertia is WCSS, a sum of squared Euclidean gaps.
6. Assign a new nearby point
import pandas as pd
import numpy as np
points = pd.DataFrame({"X1": [1, 1, 2, 7, 8, 8],
"X2": [1, 2, 1, 7, 7, 8]},
index=["A", "B", "C", "D", "E", "F"])
from sklearn.cluster import KMeans
km = KMeans(n_clusters=2, random_state=42, n_init=10)
points["Cluster"] = km.fit_predict(points[["X1", "X2"]])
new = pd.DataFrame({"X1": [1.5], "X2": [1.5]})
print(km.predict(new).tolist())Expected output
[0]Prediction uses existing fitted centres.
7. Check a new point manually
import pandas as pd
import numpy as np
points = pd.DataFrame({"X1": [1, 1, 2, 7, 8, 8],
"X2": [1, 2, 1, 7, 7, 8]},
index=["A", "B", "C", "D", "E", "F"])
from sklearn.cluster import KMeans
km = KMeans(n_clusters=2, random_state=42, n_init=10)
points["Cluster"] = km.fit_predict(points[["X1", "X2"]])
p = np.array([1.5, 1.5])
d2 = ((km.cluster_centers_ - p) ** 2).sum(axis=1)
print(d2.round(3).tolist())
print(int(d2.argmin()))Expected output
[0.056, 72.056]
0This reproduces the nearest-centre decision.
8. Verify group means
import pandas as pd
import numpy as np
points = pd.DataFrame({"X1": [1, 1, 2, 7, 8, 8],
"X2": [1, 2, 1, 7, 7, 8]},
index=["A", "B", "C", "D", "E", "F"])
from sklearn.cluster import KMeans
km = KMeans(n_clusters=2, random_state=42, n_init=10)
points["Cluster"] = km.fit_predict(points[["X1", "X2"]])
means = pd.DataFrame([
points.loc[points["Cluster"] == 0, ["X1", "X2"]].mean(),
points.loc[points["Cluster"] == 1, ["X1", "X2"]].mean()
], index=[0, 1])
print(np.allclose(means.to_numpy(), km.cluster_centers_))Expected output
TrueThe means correspond to each final fitted group.
9. Reconstruct WCSS
import pandas as pd
import numpy as np
points = pd.DataFrame({"X1": [1, 1, 2, 7, 8, 8],
"X2": [1, 2, 1, 7, 7, 8]},
index=["A", "B", "C", "D", "E", "F"])
from sklearn.cluster import KMeans
km = KMeans(n_clusters=2, random_state=42, n_init=10)
points["Cluster"] = km.fit_predict(points[["X1", "X2"]])
assigned = km.cluster_centers_[km.labels_]
wcss = ((points[["X1", "X2"]].to_numpy() - assigned) ** 2).sum()
print(round(wcss, 3))Expected output
2.667The reconstructed value matches inertia_.
10. Distinguish assignment and refitting
import pandas as pd
import numpy as np
points = pd.DataFrame({"X1": [1, 1, 2, 7, 8, 8],
"X2": [1, 2, 1, 7, 7, 8]},
index=["A", "B", "C", "D", "E", "F"])
from sklearn.cluster import KMeans
km = KMeans(n_clusters=2, random_state=42, n_init=10)
points["Cluster"] = km.fit_predict(points[["X1", "X2"]])
before = km.cluster_centers_.copy()
km.predict(pd.DataFrame({"X1": [20], "X2": [20]}))
print(np.allclose(before, km.cluster_centers_))Expected output
Truepredict does not retrain even when the new point is far away.