Chapter 4 · Section 6 practice
Extension: Hierarchical Clustering and Dendrograms
Work from question 1 to 10. Edit your Python box, click Run Code, review the result, then click Submit attempt. Reference answers unlock after all ten attempts are submitted. This records completion, not correctness. Each box runs independently. Progress is saved in this browser.
1. Count merge rows
Warm-up. Print the shape of Z.
Write your answer, then run it.
2. Read the first merge height
Warm-up. Print the first merge height rounded to three decimals.
Write your answer, then run it.
3. Read the final group size
Warm-up. Print the count stored in the final merge row.
Write your answer, then run it.
4. Cut below the first merge
Build your skills. Use cut height 0.5 and print the number of groups.
Write your answer, then run it.
5. Make a two-group cut
Build your skills. Use cut height 3 and print membership labels.
Write your answer, then run it.
6. Cut above the final merge
Build your skills. Use cut height 20 and print the group count.
Write your answer, then run it.
7. Inspect cut group sizes
Build your skills. At cut height 3, print membership counts.
Write your answer, then run it.
8. Calculate an average linkage gap
Challenge. Calculate the mean pairwise distance between {A,B} and {D,E}.
Write your answer, then run it.
9. Compare linkage choices
Challenge. Build a complete-linkage tree and print its final merge height beside Ward’s.
Write your answer, then run it.
10. Compare memberships without label IDs
Challenge. Fit KMeans K=2 and cut Ward at 3; compare the pairwise same-group matrices.
Write your answer, then run it.
Reference answers are locked until all 10 attempts are submitted.
1. Count merge rows
import pandas as pd
import numpy as np
points = pd.DataFrame({"X1": [1, 1, 2, 7, 8, 8],
"X2": [1, 2, 1, 7, 7, 8]},
index=["A", "B", "C", "D", "E", "F"])
from scipy.cluster.hierarchy import linkage, fcluster
Z = linkage(points[["X1", "X2"]], method="ward")
print(Z.shape)Expected output
(5, 4)Six points need five merges to become one group.
2. Read the first merge height
import pandas as pd
import numpy as np
points = pd.DataFrame({"X1": [1, 1, 2, 7, 8, 8],
"X2": [1, 2, 1, 7, 7, 8]},
index=["A", "B", "C", "D", "E", "F"])
from scipy.cluster.hierarchy import linkage, fcluster
Z = linkage(points[["X1", "X2"]], method="ward")
print(round(Z[0, 2], 3))Expected output
1.0The height depends on the selected linkage.
3. Read the final group size
import pandas as pd
import numpy as np
points = pd.DataFrame({"X1": [1, 1, 2, 7, 8, 8],
"X2": [1, 2, 1, 7, 7, 8]},
index=["A", "B", "C", "D", "E", "F"])
from scipy.cluster.hierarchy import linkage, fcluster
Z = linkage(points[["X1", "X2"]], method="ward")
print(int(Z[-1, 3]))Expected output
6The final node contains all original observations.
4. Cut below the first merge
import pandas as pd
import numpy as np
points = pd.DataFrame({"X1": [1, 1, 2, 7, 8, 8],
"X2": [1, 2, 1, 7, 7, 8]},
index=["A", "B", "C", "D", "E", "F"])
from scipy.cluster.hierarchy import linkage, fcluster
Z = linkage(points[["X1", "X2"]], method="ward")
labels = fcluster(Z, t=0.5, criterion="distance")
print(len(np.unique(labels)))Expected output
6Below the first merge, observations remain separate.
5. Make a two-group cut
import pandas as pd
import numpy as np
points = pd.DataFrame({"X1": [1, 1, 2, 7, 8, 8],
"X2": [1, 2, 1, 7, 7, 8]},
index=["A", "B", "C", "D", "E", "F"])
from scipy.cluster.hierarchy import linkage, fcluster
Z = linkage(points[["X1", "X2"]], method="ward")
labels = fcluster(Z, t=3, criterion="distance")
print(pd.Series(labels, index=points.index))Expected output
A 1
B 1
C 1
D 2
E 2
F 2
dtype: int32Membership follows tree connections below the cut.
6. Cut above the final merge
import pandas as pd
import numpy as np
points = pd.DataFrame({"X1": [1, 1, 2, 7, 8, 8],
"X2": [1, 2, 1, 7, 7, 8]},
index=["A", "B", "C", "D", "E", "F"])
from scipy.cluster.hierarchy import linkage, fcluster
Z = linkage(points[["X1", "X2"]], method="ward")
labels = fcluster(Z, t=20, criterion="distance")
print(len(np.unique(labels)))Expected output
1Above the final merge all observations belong together.
7. Inspect cut group sizes
import pandas as pd
import numpy as np
points = pd.DataFrame({"X1": [1, 1, 2, 7, 8, 8],
"X2": [1, 2, 1, 7, 7, 8]},
index=["A", "B", "C", "D", "E", "F"])
from scipy.cluster.hierarchy import linkage, fcluster
Z = linkage(points[["X1", "X2"]], method="ward")
labels = fcluster(Z, t=3, criterion="distance")
print(pd.Series(labels).value_counts().sort_index())Expected output
1 3
2 3
Name: count, dtype: int64Counts describe the partition regardless of label numbers.
8. Calculate an average linkage gap
import pandas as pd
import numpy as np
points = pd.DataFrame({"X1": [1, 1, 2, 7, 8, 8],
"X2": [1, 2, 1, 7, 7, 8]},
index=["A", "B", "C", "D", "E", "F"])
from scipy.cluster.hierarchy import linkage, fcluster
Z = linkage(points[["X1", "X2"]], method="ward")
from scipy.spatial.distance import cdist
d = cdist(points.loc[["A", "B"]], points.loc[["D", "E"]])
print(round(d.mean(), 3))Expected output
8.529Average linkage considers all four cross-group pairs.
9. Compare linkage choices
import pandas as pd
import numpy as np
points = pd.DataFrame({"X1": [1, 1, 2, 7, 8, 8],
"X2": [1, 2, 1, 7, 7, 8]},
index=["A", "B", "C", "D", "E", "F"])
from scipy.cluster.hierarchy import linkage, fcluster
Z = linkage(points[["X1", "X2"]], method="ward")
complete = linkage(points, method="complete")
print(round(complete[-1, 2], 3), round(Z[-1, 2], 3))Expected output
9.899 15.111Different linkage heights have different definitions and should not be ranked as the same objective.
10. Compare memberships without label IDs
import pandas as pd
import numpy as np
points = pd.DataFrame({"X1": [1, 1, 2, 7, 8, 8],
"X2": [1, 2, 1, 7, 7, 8]},
index=["A", "B", "C", "D", "E", "F"])
from scipy.cluster.hierarchy import linkage, fcluster
Z = linkage(points[["X1", "X2"]], method="ward")
from sklearn.cluster import KMeans
a = KMeans(n_clusters=2, random_state=42, n_init=10).fit_predict(points)
b = fcluster(Z, t=3, criterion="distance")
print(np.array_equal(a[:, None] == a[None, :], b[:, None] == b[None, :]))Expected output
TrueMatching membership relations identify the same partition even when numeric labels differ.