Chapter 4 · Section 1 practice
Business Groups, Simple Rules, and Features
Work from question 1 to 10. Edit your Python box, click Run Code, review the result, then click Submit attempt. Reference answers unlock after all ten attempts are submitted. This records completion, not correctness. Each box runs independently. Progress is saved in this browser.
1. Count customers
Warm-up. Print the number of customer rows.
Write your answer, then run it.
2. Select spending features
Warm-up. Print the first two rows of the six spending columns.
Write your answer, then run it.
3. Identify coded categories
Warm-up. Print the unique Channel codes.
Write your answer, then run it.
4. Apply a simple rule
Build your skills. Count customers with Grocery at least 10,000 m.u.
Write your answer, then run it.
5. Count both rule groups
Build your skills. Print counts for both sides of the Grocery rule.
Write your answer, then run it.
6. Profile a rule
Build your skills. Print average Milk spending for both Grocery-rule groups.
Write your answer, then run it.
7. Change the rule
Build your skills. Count customers classified high at cutoffs 8,000 and 12,000.
Write your answer, then run it.
8. Check feature completeness
Challenge. Print the total missing entries in the selected feature table.
Write your answer, then run it.
9. Compare distribution spread
Challenge. Print the mean and sample SD of Fresh and Grocery.
Write your answer, then run it.
10. Separate inputs from interpretation
Challenge. Print whether Channel or Region appears in features; add a sentence about their later use.
Write your answer, then run it.
Reference answers are locked until all 10 attempts are submitted.
1. Count customers
import pandas as pd
import numpy as np
from sklearn.preprocessing import StandardScaler
from sklearn.cluster import KMeans
wholesale = pd.read_csv(
"https://busanalytics-book.pages.dev/data/chapter4-wholesale.csv")
features = ["Fresh", "Milk", "Grocery", "Frozen", "Detergents_Paper", "Delicassen"]
scaler = StandardScaler()
X = scaler.fit_transform(wholesale[features])
print(len(wholesale))Expected output
440The UCI snapshot contains 440 customer records.
2. Select spending features
import pandas as pd
import numpy as np
from sklearn.preprocessing import StandardScaler
from sklearn.cluster import KMeans
wholesale = pd.read_csv(
"https://busanalytics-book.pages.dev/data/chapter4-wholesale.csv")
features = ["Fresh", "Milk", "Grocery", "Frozen", "Detergents_Paper", "Delicassen"]
scaler = StandardScaler()
X = scaler.fit_transform(wholesale[features])
print(wholesale[features].head(2).to_string(index=False))Expected output
Fresh Milk Grocery Frozen Detergents_Paper Delicassen
12669 9656 7561 214 2674 1338
7057 9810 9568 1762 3293 1776Only the selected spending measures define similarity.
3. Identify coded categories
import pandas as pd
import numpy as np
from sklearn.preprocessing import StandardScaler
from sklearn.cluster import KMeans
wholesale = pd.read_csv(
"https://busanalytics-book.pages.dev/data/chapter4-wholesale.csv")
features = ["Fresh", "Milk", "Grocery", "Frozen", "Detergents_Paper", "Delicassen"]
scaler = StandardScaler()
X = scaler.fit_transform(wholesale[features])
print(sorted(wholesale["Channel"].unique().tolist()))Expected output
[1, 2]Channel is a coded category, not an annual spending amount.
4. Apply a simple rule
import pandas as pd
import numpy as np
from sklearn.preprocessing import StandardScaler
from sklearn.cluster import KMeans
wholesale = pd.read_csv(
"https://busanalytics-book.pages.dev/data/chapter4-wholesale.csv")
features = ["Fresh", "Milk", "Grocery", "Frozen", "Detergents_Paper", "Delicassen"]
scaler = StandardScaler()
X = scaler.fit_transform(wholesale[features])
print(int((wholesale["Grocery"] >= 10000).sum()))Expected output
117The threshold is a business rule, not an estimated cluster boundary.
5. Count both rule groups
import pandas as pd
import numpy as np
from sklearn.preprocessing import StandardScaler
from sklearn.cluster import KMeans
wholesale = pd.read_csv(
"https://busanalytics-book.pages.dev/data/chapter4-wholesale.csv")
features = ["Fresh", "Milk", "Grocery", "Frozen", "Detergents_Paper", "Delicassen"]
scaler = StandardScaler()
X = scaler.fit_transform(wholesale[features])
rule = wholesale["Grocery"] >= 10000
print(rule.value_counts().sort_index())Expected output
Grocery
False 323
True 117
Name: count, dtype: int64Every customer is assigned using this single cutoff.
6. Profile a rule
import pandas as pd
import numpy as np
from sklearn.preprocessing import StandardScaler
from sklearn.cluster import KMeans
wholesale = pd.read_csv(
"https://busanalytics-book.pages.dev/data/chapter4-wholesale.csv")
features = ["Fresh", "Milk", "Grocery", "Frozen", "Detergents_Paper", "Delicassen"]
scaler = StandardScaler()
X = scaler.fit_transform(wholesale[features])
wholesale["HighGrocery"] = wholesale["Grocery"] >= 10000
print(pd.Series([
wholesale.loc[wholesale["HighGrocery"] == False, "Milk"].mean(),
wholesale.loc[wholesale["HighGrocery"] == True, "Milk"].mean()
], index=["Below 10000", "At least 10000"]).round(1))Expected output
Below 10000 3268.6
At least 10000 12774.2
dtype: float64Another feature helps reveal what the simple grouping contains.
7. Change the rule
import pandas as pd
import numpy as np
from sklearn.preprocessing import StandardScaler
from sklearn.cluster import KMeans
wholesale = pd.read_csv(
"https://busanalytics-book.pages.dev/data/chapter4-wholesale.csv")
features = ["Fresh", "Milk", "Grocery", "Frozen", "Detergents_Paper", "Delicassen"]
scaler = StandardScaler()
X = scaler.fit_transform(wholesale[features])
print(int((wholesale["Grocery"] >= 8000).sum()))
print(int((wholesale["Grocery"] >= 12000).sum()))Expected output
148
90A higher cutoff cannot increase the high-group count.
8. Check feature completeness
import pandas as pd
import numpy as np
from sklearn.preprocessing import StandardScaler
from sklearn.cluster import KMeans
wholesale = pd.read_csv(
"https://busanalytics-book.pages.dev/data/chapter4-wholesale.csv")
features = ["Fresh", "Milk", "Grocery", "Frozen", "Detergents_Paper", "Delicassen"]
scaler = StandardScaler()
X = scaler.fit_transform(wholesale[features])
print(int(wholesale[features].isna().sum().sum()))Expected output
0A complete snapshot does not eliminate the need for checks on other datasets.
9. Compare distribution spread
import pandas as pd
import numpy as np
from sklearn.preprocessing import StandardScaler
from sklearn.cluster import KMeans
wholesale = pd.read_csv(
"https://busanalytics-book.pages.dev/data/chapter4-wholesale.csv")
features = ["Fresh", "Milk", "Grocery", "Frozen", "Detergents_Paper", "Delicassen"]
scaler = StandardScaler()
X = scaler.fit_transform(wholesale[features])
print(wholesale[["Fresh", "Grocery"]].agg(["mean", "std"]).round(1))Expected output
Fresh Grocery
mean 12000.3 7951.3
std 12647.3 9503.2Features with the same monetary units can still have very different spreads.
10. Separate inputs from interpretation
import pandas as pd
import numpy as np
from sklearn.preprocessing import StandardScaler
from sklearn.cluster import KMeans
wholesale = pd.read_csv(
"https://busanalytics-book.pages.dev/data/chapter4-wholesale.csv")
features = ["Fresh", "Milk", "Grocery", "Frozen", "Detergents_Paper", "Delicassen"]
scaler = StandardScaler()
X = scaler.fit_transform(wholesale[features])
print("Channel" in features, "Region" in features)
print("Use these categories after clustering to interpret the spending-based groups.")Expected output
False False
Use these categories after clustering to interpret the spending-based groups.Do not accidentally feed interpretation labels into this feature matrix.