Chapter 4 · Section 5 practice
Customer Profiles and Business Actions
Work from question 1 to 10. Edit your Python box, click Run Code, review the result, then click Submit attempt. Reference answers unlock after all ten attempts are submitted. This records completion, not correctness. Each box runs independently. Progress is saved in this browser.
1. Count fitted groups
Warm-up. Print membership counts.
Write your answer, then run it.
2. Profile Grocery
Warm-up. Print average Grocery spending per group.
Write your answer, then run it.
3. Profile two categories
Warm-up. Print mean Milk and Grocery spending by group.
Write your answer, then run it.
4. Read original-unit centres
Build your skills. Inverse-transform the centres and print the first two categories.
Write your answer, then run it.
5. Identify a descriptive group
Build your skills. Print the cluster with the highest mean Grocery spending.
Write your answer, then run it.
6. Check group context
Build your skills. For the highest-mean-Grocery group, print its size and mean Milk spending.
Write your answer, then run it.
7. Inspect channel mix
Build your skills. Print a Channel-by-Cluster cross-tabulation.
Write your answer, then run it.
8. Inspect proportions
Challenge. Print the proportion of each Channel within each Cluster.
Write your answer, then run it.
9. Check centres against profiles
Challenge. Print whether inverse-transformed centres match group means in label order.
Write your answer, then run it.
10. Write an evidence-based proposal
Challenge. Print the high-Grocery group’s profile, followed by a bundle-pilot proposal that avoids claiming proven profit.
Write your answer, then run it.
Reference answers are locked until all 10 attempts are submitted.
1. Count fitted groups
import pandas as pd
import numpy as np
from sklearn.preprocessing import StandardScaler
from sklearn.cluster import KMeans
wholesale = pd.read_csv(
"https://busanalytics-book.pages.dev/data/chapter4-wholesale.csv")
features = ["Fresh", "Milk", "Grocery", "Frozen", "Detergents_Paper", "Delicassen"]
scaler = StandardScaler()
X = scaler.fit_transform(wholesale[features])
km = KMeans(n_clusters=3, random_state=42, n_init=10)
wholesale["Cluster"] = km.fit_predict(X)
print(wholesale["Cluster"].value_counts().sort_index())Expected output
Cluster
0 45
1 393
2 2
Name: count, dtype: int64Group sizes help interpret an average.
2. Profile Grocery
import pandas as pd
import numpy as np
from sklearn.preprocessing import StandardScaler
from sklearn.cluster import KMeans
wholesale = pd.read_csv(
"https://busanalytics-book.pages.dev/data/chapter4-wholesale.csv")
features = ["Fresh", "Milk", "Grocery", "Frozen", "Detergents_Paper", "Delicassen"]
scaler = StandardScaler()
X = scaler.fit_transform(wholesale[features])
km = KMeans(n_clusters=3, random_state=42, n_init=10)
wholesale["Cluster"] = km.fit_predict(X)
print(pd.Series([
wholesale.loc[wholesale["Cluster"] == 0, "Grocery"].mean(),
wholesale.loc[wholesale["Cluster"] == 1, "Grocery"].mean(),
wholesale.loc[wholesale["Cluster"] == 2, "Grocery"].mean()
], index=[0, 1, 2]).round(1))Expected output
0 28656.1
1 5535.0
2 16898.0
dtype: float64The amounts are annual m.u., not standardised scores.
3. Profile two categories
import pandas as pd
import numpy as np
from sklearn.preprocessing import StandardScaler
from sklearn.cluster import KMeans
wholesale = pd.read_csv(
"https://busanalytics-book.pages.dev/data/chapter4-wholesale.csv")
features = ["Fresh", "Milk", "Grocery", "Frozen", "Detergents_Paper", "Delicassen"]
scaler = StandardScaler()
X = scaler.fit_transform(wholesale[features])
km = KMeans(n_clusters=3, random_state=42, n_init=10)
wholesale["Cluster"] = km.fit_predict(X)
print(pd.DataFrame([
wholesale.loc[wholesale["Cluster"] == 0, ["Milk", "Grocery"]].mean(),
wholesale.loc[wholesale["Cluster"] == 1, ["Milk", "Grocery"]].mean(),
wholesale.loc[wholesale["Cluster"] == 2, ["Milk", "Grocery"]].mean()
], index=[0, 1, 2]).round(1))Expected output
Milk Grocery
0 19386.4 28656.1
1 4115.1 5535.0
2 30367.0 16898.0Several features give a clearer profile than one label.
4. Read original-unit centres
import pandas as pd
import numpy as np
from sklearn.preprocessing import StandardScaler
from sklearn.cluster import KMeans
wholesale = pd.read_csv(
"https://busanalytics-book.pages.dev/data/chapter4-wholesale.csv")
features = ["Fresh", "Milk", "Grocery", "Frozen", "Detergents_Paper", "Delicassen"]
scaler = StandardScaler()
X = scaler.fit_transform(wholesale[features])
km = KMeans(n_clusters=3, random_state=42, n_init=10)
wholesale["Cluster"] = km.fit_predict(X)
centres = pd.DataFrame(scaler.inverse_transform(km.cluster_centers_), columns=features)
print(centres[["Fresh", "Milk"]].round(1))Expected output
Fresh Milk
0 10440.9 19386.4
1 12062.9 4115.1
2 34782.0 30367.0Use the fitted scaler rather than fitting a new one.
5. Identify a descriptive group
import pandas as pd
import numpy as np
from sklearn.preprocessing import StandardScaler
from sklearn.cluster import KMeans
wholesale = pd.read_csv(
"https://busanalytics-book.pages.dev/data/chapter4-wholesale.csv")
features = ["Fresh", "Milk", "Grocery", "Frozen", "Detergents_Paper", "Delicassen"]
scaler = StandardScaler()
X = scaler.fit_transform(wholesale[features])
km = KMeans(n_clusters=3, random_state=42, n_init=10)
wholesale["Cluster"] = km.fit_predict(X)
profile = pd.DataFrame([
wholesale.loc[wholesale["Cluster"] == 0, features].mean(),
wholesale.loc[wholesale["Cluster"] == 1, features].mean(),
wholesale.loc[wholesale["Cluster"] == 2, features].mean()
], index=[0, 1, 2])
print(int(profile["Grocery"].idxmax()))Expected output
0The ID is a label selected by a measured profile, not a ranking built into KMeans.
6. Check group context
import pandas as pd
import numpy as np
from sklearn.preprocessing import StandardScaler
from sklearn.cluster import KMeans
wholesale = pd.read_csv(
"https://busanalytics-book.pages.dev/data/chapter4-wholesale.csv")
features = ["Fresh", "Milk", "Grocery", "Frozen", "Detergents_Paper", "Delicassen"]
scaler = StandardScaler()
X = scaler.fit_transform(wholesale[features])
km = KMeans(n_clusters=3, random_state=42, n_init=10)
wholesale["Cluster"] = km.fit_predict(X)
profile = pd.DataFrame([
wholesale.loc[wholesale["Cluster"] == 0, features].mean(),
wholesale.loc[wholesale["Cluster"] == 1, features].mean(),
wholesale.loc[wholesale["Cluster"] == 2, features].mean()
], index=[0, 1, 2])
chosen = profile["Grocery"].idxmax()
print(int((wholesale["Cluster"] == chosen).sum()))
print(round(profile.loc[chosen, "Milk"], 1))Expected output
45
19386.4A supported memo reports both amount and group size.
7. Inspect channel mix
import pandas as pd
import numpy as np
from sklearn.preprocessing import StandardScaler
from sklearn.cluster import KMeans
wholesale = pd.read_csv(
"https://busanalytics-book.pages.dev/data/chapter4-wholesale.csv")
features = ["Fresh", "Milk", "Grocery", "Frozen", "Detergents_Paper", "Delicassen"]
scaler = StandardScaler()
X = scaler.fit_transform(wholesale[features])
km = KMeans(n_clusters=3, random_state=42, n_init=10)
wholesale["Cluster"] = km.fit_predict(X)
print(pd.crosstab(wholesale["Channel"], wholesale["Cluster"]))Expected output
Cluster 0 1 2
Channel
1 1 295 2
2 44 98 0Channel is used after fitting to interpret the groups.
8. Inspect proportions
import pandas as pd
import numpy as np
from sklearn.preprocessing import StandardScaler
from sklearn.cluster import KMeans
wholesale = pd.read_csv(
"https://busanalytics-book.pages.dev/data/chapter4-wholesale.csv")
features = ["Fresh", "Milk", "Grocery", "Frozen", "Detergents_Paper", "Delicassen"]
scaler = StandardScaler()
X = scaler.fit_transform(wholesale[features])
km = KMeans(n_clusters=3, random_state=42, n_init=10)
wholesale["Cluster"] = km.fit_predict(X)
print(pd.crosstab(wholesale["Channel"], wholesale["Cluster"], normalize="columns").round(3))Expected output
Cluster 0 1 2
Channel
1 0.022 0.751 1.0
2 0.978 0.249 0.0Column normalisation answers the within-cluster composition question.
9. Check centres against profiles
import pandas as pd
import numpy as np
from sklearn.preprocessing import StandardScaler
from sklearn.cluster import KMeans
wholesale = pd.read_csv(
"https://busanalytics-book.pages.dev/data/chapter4-wholesale.csv")
features = ["Fresh", "Milk", "Grocery", "Frozen", "Detergents_Paper", "Delicassen"]
scaler = StandardScaler()
X = scaler.fit_transform(wholesale[features])
km = KMeans(n_clusters=3, random_state=42, n_init=10)
wholesale["Cluster"] = km.fit_predict(X)
centres = scaler.inverse_transform(km.cluster_centers_)
means = pd.DataFrame([
wholesale.loc[wholesale["Cluster"] == 0, features].mean(),
wholesale.loc[wholesale["Cluster"] == 1, features].mean(),
wholesale.loc[wholesale["Cluster"] == 2, features].mean()
], index=[0, 1, 2]).to_numpy()
print(np.allclose(centres, means))Expected output
TrueCoordinate means and centre values agree for this converged fit.
10. Write an evidence-based proposal
import pandas as pd
import numpy as np
from sklearn.preprocessing import StandardScaler
from sklearn.cluster import KMeans
wholesale = pd.read_csv(
"https://busanalytics-book.pages.dev/data/chapter4-wholesale.csv")
features = ["Fresh", "Milk", "Grocery", "Frozen", "Detergents_Paper", "Delicassen"]
scaler = StandardScaler()
X = scaler.fit_transform(wholesale[features])
km = KMeans(n_clusters=3, random_state=42, n_init=10)
wholesale["Cluster"] = km.fit_predict(X)
profile = pd.DataFrame([
wholesale.loc[wholesale["Cluster"] == 0, features].mean(),
wholesale.loc[wholesale["Cluster"] == 1, features].mean(),
wholesale.loc[wholesale["Cluster"] == 2, features].mean()
], index=[0, 1, 2])
chosen = profile["Grocery"].idxmax()
print(profile.loc[chosen, ["Milk", "Grocery"]].round(1))
print("Test a Milk–Grocery bundle and measure incremental margin before rollout.")Expected output
Milk 19386.4
Grocery 28656.1
Name: 0, dtype: float64
Test a Milk–Grocery bundle and measure incremental margin before rollout.The action is a testable hypothesis, not a causal conclusion from clustering.