Chapter 4 · Section 2 practice
Distance Measures and Standardisation
Work from question 1 to 10. Edit your Python box, click Run Code, review the result, then click Submit attempt. Reference answers unlock after all ten attempts are submitted. This records completion, not correctness. Each box runs independently. Progress is saved in this browser.
1. Calculate one coordinate gap
Warm-up. Print the Income difference A minus B.
Write your answer, then run it.
2. Calculate raw Euclidean distance
Warm-up. Print the raw A–B distance rounded to three decimals.
Write your answer, then run it.
3. Calculate Manhattan distance
Warm-up. Print raw A–B Manhattan distance.
Write your answer, then run it.
4. Check standardised means
Build your skills. Print the three standardised means rounded to six decimals.
Write your answer, then run it.
5. Check population SD
Build your skills. Print standardised population SDs.
Write your answer, then run it.
6. Compare SD conventions
Build your skills. Print sample and population SD of standardised Age.
Write your answer, then run it.
7. Calculate standardised distance
Build your skills. Print A–B Euclidean distance on the standardised features.
Write your answer, then run it.
8. Compare two neighbours
Challenge. Print standardised A–B and A–C distances.
Write your answer, then run it.
9. Transform a new customer
Challenge. Use the existing scaler for Age=30, Income=50000, Score=70 and print the result.
Write your answer, then run it.
10. Reconstruct a scaling calculation
Challenge. Manually standardise Income with ddof=0; print whether it matches the scaled Income column.
Write your answer, then run it.
Reference answers are locked until all 10 attempts are submitted.
1. Calculate one coordinate gap
import pandas as pd
import numpy as np
customers = pd.DataFrame({"Age": [35, 28, 60, 45, 22],
"Income": [60000, 80000, 30000, 55000, 40000],
"Score": [75, 85, 20, 60, 90]}, index=["A", "B", "C", "D", "E"])
from sklearn.preprocessing import StandardScaler
scaler = StandardScaler()
scaled = pd.DataFrame(scaler.fit_transform(customers),
index=customers.index, columns=customers.columns)
print(customers.loc["A", "Income"] - customers.loc["B", "Income"])Expected output
-20000Income gaps use the original currency-unit scale.
2. Calculate raw Euclidean distance
import pandas as pd
import numpy as np
customers = pd.DataFrame({"Age": [35, 28, 60, 45, 22],
"Income": [60000, 80000, 30000, 55000, 40000],
"Score": [75, 85, 20, 60, 90]}, index=["A", "B", "C", "D", "E"])
from sklearn.preprocessing import StandardScaler
scaler = StandardScaler()
scaled = pd.DataFrame(scaler.fit_transform(customers),
index=customers.index, columns=customers.columns)
print(round(np.sqrt(((customers.loc["A"] - customers.loc["B"]) ** 2).sum()), 3))Expected output
20000.004Square all coordinate differences, sum, then take the root.
3. Calculate Manhattan distance
import pandas as pd
import numpy as np
customers = pd.DataFrame({"Age": [35, 28, 60, 45, 22],
"Income": [60000, 80000, 30000, 55000, 40000],
"Score": [75, 85, 20, 60, 90]}, index=["A", "B", "C", "D", "E"])
from sklearn.preprocessing import StandardScaler
scaler = StandardScaler()
scaled = pd.DataFrame(scaler.fit_transform(customers),
index=customers.index, columns=customers.columns)
print(int(np.abs(customers.loc["A"] - customers.loc["B"]).sum()))Expected output
20017This sums absolute gaps rather than squared gaps.
4. Check standardised means
import pandas as pd
import numpy as np
customers = pd.DataFrame({"Age": [35, 28, 60, 45, 22],
"Income": [60000, 80000, 30000, 55000, 40000],
"Score": [75, 85, 20, 60, 90]}, index=["A", "B", "C", "D", "E"])
from sklearn.preprocessing import StandardScaler
scaler = StandardScaler()
scaled = pd.DataFrame(scaler.fit_transform(customers),
index=customers.index, columns=customers.columns)
print(scaled.mean().round(6))Expected output
Age 0.0
Income 0.0
Score -0.0
dtype: float64Each varying feature is centred on zero, allowing floating-point rounding.
5. Check population SD
import pandas as pd
import numpy as np
customers = pd.DataFrame({"Age": [35, 28, 60, 45, 22],
"Income": [60000, 80000, 30000, 55000, 40000],
"Score": [75, 85, 20, 60, 90]}, index=["A", "B", "C", "D", "E"])
from sklearn.preprocessing import StandardScaler
scaler = StandardScaler()
scaled = pd.DataFrame(scaler.fit_transform(customers),
index=customers.index, columns=customers.columns)
print(scaled.std(ddof=0).round(3))Expected output
Age 1.0
Income 1.0
Score 1.0
dtype: float64This matches StandardScaler's convention.
6. Compare SD conventions
import pandas as pd
import numpy as np
customers = pd.DataFrame({"Age": [35, 28, 60, 45, 22],
"Income": [60000, 80000, 30000, 55000, 40000],
"Score": [75, 85, 20, 60, 90]}, index=["A", "B", "C", "D", "E"])
from sklearn.preprocessing import StandardScaler
scaler = StandardScaler()
scaled = pd.DataFrame(scaler.fit_transform(customers),
index=customers.index, columns=customers.columns)
print(round(scaled["Age"].std(), 3), round(scaled["Age"].std(ddof=0), 3))Expected output
1.118 1.0The sample denominator gives sqrt(5/4), not 1.
7. Calculate standardised distance
import pandas as pd
import numpy as np
customers = pd.DataFrame({"Age": [35, 28, 60, 45, 22],
"Income": [60000, 80000, 30000, 55000, 40000],
"Score": [75, 85, 20, 60, 90]}, index=["A", "B", "C", "D", "E"])
from sklearn.preprocessing import StandardScaler
scaler = StandardScaler()
scaled = pd.DataFrame(scaler.fit_transform(customers),
index=customers.index, columns=customers.columns)
print(round(np.sqrt(((scaled.loc["A"] - scaled.loc["B"]) ** 2).sum()), 3))Expected output
1.335The same records now use standard-deviation units.
8. Compare two neighbours
import pandas as pd
import numpy as np
customers = pd.DataFrame({"Age": [35, 28, 60, 45, 22],
"Income": [60000, 80000, 30000, 55000, 40000],
"Score": [75, 85, 20, 60, 90]}, index=["A", "B", "C", "D", "E"])
from sklearn.preprocessing import StandardScaler
scaler = StandardScaler()
scaled = pd.DataFrame(scaler.fit_transform(customers),
index=customers.index, columns=customers.columns)
print(round(np.sqrt(((scaled.loc["A"] - scaled.loc["B"]) ** 2).sum()), 3))
print(round(np.sqrt(((scaled.loc["A"] - scaled.loc["C"]) ** 2).sum()), 3))Expected output
1.335
3.36Similarity is conditional on the selected weighting and features.
9. Transform a new customer
import pandas as pd
import numpy as np
customers = pd.DataFrame({"Age": [35, 28, 60, 45, 22],
"Income": [60000, 80000, 30000, 55000, 40000],
"Score": [75, 85, 20, 60, 90]}, index=["A", "B", "C", "D", "E"])
from sklearn.preprocessing import StandardScaler
scaler = StandardScaler()
scaled = pd.DataFrame(scaler.fit_transform(customers),
index=customers.index, columns=customers.columns)
new = pd.DataFrame({"Age": [30], "Income": [50000], "Score": [70]})
print(scaler.transform(new).round(3))Expected output
[[-0.597 -0.174 0.159]]Do not fit a separate scaler to the new record.
10. Reconstruct a scaling calculation
import pandas as pd
import numpy as np
customers = pd.DataFrame({"Age": [35, 28, 60, 45, 22],
"Income": [60000, 80000, 30000, 55000, 40000],
"Score": [75, 85, 20, 60, 90]}, index=["A", "B", "C", "D", "E"])
from sklearn.preprocessing import StandardScaler
scaler = StandardScaler()
scaled = pd.DataFrame(scaler.fit_transform(customers),
index=customers.index, columns=customers.columns)
manual = (customers["Income"] - customers["Income"].mean()) / customers["Income"].std(
ddof=0
)
print(np.allclose(manual, scaled["Income"]))Expected output
TrueManual and library calculations agree when the SD conventions match.