Chapter 4 quiz
Ten questions to bring the chapter together.
Work from question 1 to 10. Edit your Python box, click Run Code, review the result, then click Submit attempt. Reference answers unlock after all ten attempts are submitted. This records completion, not correctness. Each box runs independently. Progress is saved in this browser.
Question 1
Count wholesale customers whose Grocery spending is at least 10,000 m.u.
Write your answer, then run it.
Question 2
For A=(1,1), B=(4,5), print Euclidean and Manhattan distances.
Write your answer, then run it.
Question 3
Manually standardise classroom Income using ddof=0 and compare with StandardScaler.
Write your answer, then run it.
Question 4
Assign (1.5,1.5) to the fitted six-point model and print its distances to the centres.
Write your answer, then run it.
Question 5
Reconstruct WCSS for the six-point fitted model.
Write your answer, then run it.
Question 6
Print mall WCSS for K=4 and K=5 and a sentence explaining why the smaller value does not settle the business choice.
Write your answer, then run it.
Question 7
Print wholesale original-unit mean Milk and Grocery spending for each working group.
Write your answer, then run it.
Question 8
Inverse-transform wholesale centroids and verify that they match group means.
Write your answer, then run it.
Question 9
Print the largest wholesale group size and a cautious business proposal.
Write your answer, then run it.
Question 10
For the six-point Ward tree, print group counts at cut heights 0.5,3,20.
Write your answer, then run it.
Reference answers are locked until all 10 attempts are submitted.
Question 1
import pandas as pd
import numpy as np
from sklearn.preprocessing import StandardScaler
from sklearn.cluster import KMeans
wholesale = pd.read_csv(
"https://busanalytics-book.pages.dev/data/chapter4-wholesale.csv")
features = ["Fresh", "Milk", "Grocery", "Frozen", "Detergents_Paper", "Delicassen"]
scaler = StandardScaler()
X = scaler.fit_transform(wholesale[features])
print(int((wholesale["Grocery"] >= 10000).sum()))Expected output
117A simple threshold rule uses one chosen feature.
Question 2
import numpy as np
a = np.array([1, 1])
b = np.array([4, 5])
print(np.sqrt(((a - b) ** 2).sum()))
print(np.abs(a - b).sum())Expected output
5.0
7Euclidean uses squared gaps and a square root; Manhattan sums absolute gaps.
Question 3
import pandas as pd
import numpy as np
customers = pd.DataFrame({"Age": [35, 28, 60, 45, 22],
"Income": [60000, 80000, 30000, 55000, 40000],
"Score": [75, 85, 20, 60, 90]}, index=["A", "B", "C", "D", "E"])
from sklearn.preprocessing import StandardScaler
scaler = StandardScaler()
scaled = pd.DataFrame(scaler.fit_transform(customers),
index=customers.index, columns=customers.columns)
manual = (customers["Income"] - customers["Income"].mean()) / customers["Income"].std(
ddof=0
)
print(np.allclose(manual, scaled["Income"]))Expected output
TrueThe population SD convention must match.
Question 4
import pandas as pd
import numpy as np
points = pd.DataFrame({"X1": [1, 1, 2, 7, 8, 8],
"X2": [1, 2, 1, 7, 7, 8]},
index=["A", "B", "C", "D", "E", "F"])
from sklearn.cluster import KMeans
km = KMeans(n_clusters=2, random_state=42, n_init=10)
points["Cluster"] = km.fit_predict(points[["X1", "X2"]])
p = np.array([1.5, 1.5])
d = np.sqrt(((km.cluster_centers_ - p) ** 2).sum(axis=1))
print(d.round(3).tolist())
print(km.predict(pd.DataFrame({"X1": [1.5], "X2": [1.5]})).tolist())Expected output
[0.236, 8.489]
[0]The assigned centre is the nearest in the selected feature space.
Question 5
import pandas as pd
import numpy as np
points = pd.DataFrame({"X1": [1, 1, 2, 7, 8, 8],
"X2": [1, 2, 1, 7, 7, 8]},
index=["A", "B", "C", "D", "E", "F"])
from sklearn.cluster import KMeans
km = KMeans(n_clusters=2, random_state=42, n_init=10)
points["Cluster"] = km.fit_predict(points[["X1", "X2"]])
assigned = km.cluster_centers_[km.labels_]
print(round(((points[["X1", "X2"]].to_numpy() - assigned) ** 2).sum(), 3))Expected output
2.667Use squared distances to each assigned centre.
Question 6
import pandas as pd
import numpy as np
from sklearn.preprocessing import StandardScaler
from sklearn.cluster import KMeans
mall = pd.read_csv(
"https://busanalytics-book.pages.dev/data/chapter4-mall-classroom.csv",
index_col="CustomerID")
features = ["Annual_Income", "Spending_Score"]
scaler = StandardScaler()
X = scaler.fit_transform(mall[features])
a = KMeans(n_clusters=4, random_state=42, n_init=10).fit(X)
b = KMeans(n_clusters=5, random_state=42, n_init=10).fit(X)
print(round(a.inertia_, 3), round(b.inertia_, 3))
print("Also inspect profiles, sizes, stability, and actionable differences.")Expected output
74.922 37.976
Also inspect profiles, sizes, stability, and actionable differences.More groups generally reduce the fitted objective without proving usefulness.
Question 7
import pandas as pd
import numpy as np
from sklearn.preprocessing import StandardScaler
from sklearn.cluster import KMeans
wholesale = pd.read_csv(
"https://busanalytics-book.pages.dev/data/chapter4-wholesale.csv")
features = ["Fresh", "Milk", "Grocery", "Frozen", "Detergents_Paper", "Delicassen"]
scaler = StandardScaler()
X = scaler.fit_transform(wholesale[features])
km = KMeans(n_clusters=3, random_state=42, n_init=10)
wholesale["Cluster"] = km.fit_predict(X)
print(pd.DataFrame([
wholesale.loc[wholesale["Cluster"] == 0, ["Milk", "Grocery"]].mean(),
wholesale.loc[wholesale["Cluster"] == 1, ["Milk", "Grocery"]].mean(),
wholesale.loc[wholesale["Cluster"] == 2, ["Milk", "Grocery"]].mean()
], index=[0, 1, 2]).round(1))Expected output
Milk Grocery
0 19386.4 28656.1
1 4115.1 5535.0
2 30367.0 16898.0Profiles return to business units.
Question 8
import pandas as pd
import numpy as np
from sklearn.preprocessing import StandardScaler
from sklearn.cluster import KMeans
wholesale = pd.read_csv(
"https://busanalytics-book.pages.dev/data/chapter4-wholesale.csv")
features = ["Fresh", "Milk", "Grocery", "Frozen", "Detergents_Paper", "Delicassen"]
scaler = StandardScaler()
X = scaler.fit_transform(wholesale[features])
km = KMeans(n_clusters=3, random_state=42, n_init=10)
wholesale["Cluster"] = km.fit_predict(X)
centres = scaler.inverse_transform(km.cluster_centers_)
means = pd.DataFrame([
wholesale.loc[wholesale["Cluster"] == 0, features].mean(),
wholesale.loc[wholesale["Cluster"] == 1, features].mean(),
wholesale.loc[wholesale["Cluster"] == 2, features].mean()
], index=[0, 1, 2]).to_numpy()
print(np.allclose(centres, means))Expected output
TrueRetain fitted scaler and cluster-label order.
Question 9
import pandas as pd
import numpy as np
from sklearn.preprocessing import StandardScaler
from sklearn.cluster import KMeans
wholesale = pd.read_csv(
"https://busanalytics-book.pages.dev/data/chapter4-wholesale.csv")
features = ["Fresh", "Milk", "Grocery", "Frozen", "Detergents_Paper", "Delicassen"]
scaler = StandardScaler()
X = scaler.fit_transform(wholesale[features])
km = KMeans(n_clusters=3, random_state=42, n_init=10)
wholesale["Cluster"] = km.fit_predict(X)
print(int(wholesale["Cluster"].value_counts().max()))
print(
"Use the measured category profile to propose a pilot and measure incremental margin."
)Expected output
393
Use the measured category profile to propose a pilot and measure incremental margin.A descriptive grouping supports a hypothesis, not a proven campaign effect.
Question 10
import pandas as pd
import numpy as np
points = pd.DataFrame({"X1": [1, 1, 2, 7, 8, 8],
"X2": [1, 2, 1, 7, 7, 8]},
index=["A", "B", "C", "D", "E", "F"])
from scipy.cluster.hierarchy import linkage, fcluster
Z = linkage(points[["X1", "X2"]], method="ward")
counts = [len(np.unique(fcluster(Z, t=h, criterion="distance"))) for h in [0.5, 3, 20]]
print(counts)Expected output
[6, 2, 1]A higher tree cut merges groups; this is extension material.