Chapter 2 · Section 5 practice
Exploratory Data Analysis
Work from question 1 to 10. Edit your Python box, click Run Code, review the result, then click Submit attempt. Reference answers unlock after all ten attempts are submitted. This records completion, not correctness. Each box runs independently. Progress is saved in this browser.
1. Typical order
Warm-up. Print the mean and median of these sales.
Write your answer, then run it.
2. Observed count
Warm-up. Print the count of observed values and the number of rows.
Write your answer, then run it.
3. A quartile
Warm-up. Print the 75th percentile of [100, 200, 300].
Write your answer, then run it.
4. Sample spread
Build your skills. Print the sample standard deviation of [100, 200, 300].
Write your answer, then run it.
5. Category counts
Build your skills. Print category counts for these orders.
Write your answer, then run it.
7. Reduce down rows
Build your skills. Print Monday and Tuesday totals as a list.
Write your answer, then run it.
8. Reduce across columns
Challenge. Print Central and Kowloon totals as a list.
Write your answer, then run it.
9. A large observation
Challenge. Compare the full mean with the mean after removing the 1000 order. Print both.
Write your answer, then run it.
10. Calculate and interpret correlation
Challenge. Print the correlation between AdSpend and Revenue rounded to three decimals, followed by a statement that it does not establish causation.
Write your answer, then run it.
Reference answers are locked until all 10 attempts are submitted.
1. Typical order
import pandas as pd
sales = pd.Series([80, 100, 120, 140, 1000])
print(sales.mean())
print(sales.median())Expected output
288.0
120.0The mean is sensitive to the large order.
2. Observed count
import pandas as pd
sales = pd.Series([100.0, None, 0.0, 300.0])
print(sales.count())
print(len(sales))Expected output
3
4count excludes missing values, while len counts all entries.
3. A quartile
import pandas as pd
sales = pd.Series([100, 200, 300])
print(sales.quantile(0.75))Expected output
250.0With pandas default interpolation, the quartile lies between 200 and 300.
4. Sample spread
import pandas as pd
sales = pd.Series([100.0, 200.0, 300.0])
print(sales.std())Expected output
100.0The default sample denominator is n − 1.
5. Category counts
import pandas as pd
categories = pd.Series(["Gifts", "Gifts", "Stationery", "Gifts"])
print(categories.value_counts())Expected output
Gifts 3
Stationery 1
Name: count, dtype: int64Three of four entries belong to Gifts.
6. Category shares
import pandas as pd
categories = pd.Series(["Gifts", "Gifts", "Stationery", "Gifts"])
shares = categories.value_counts(normalize=True)
print(shares.loc["Gifts"])Expected output
0.75The denominator is the four counted observations.
7. Reduce down rows
import pandas as pd
df = pd.DataFrame({"Monday": [100, 200], "Tuesday": [150, 250]}, index=["Central", "Kowloon"])
print(list(df.sum(axis=0)))Expected output
[300, 400]Reducing rows leaves one total per day column.
8. Reduce across columns
import pandas as pd
df = pd.DataFrame({"Monday": [100, 200], "Tuesday": [150, 250]}, index=["Central", "Kowloon"])
print(list(df.sum(axis=1)))Expected output
[250, 450]Comparable HKD columns can be added across each store row.
9. A large observation
import pandas as pd
sales = pd.Series([80, 100, 120, 140, 1000])
print(sales.mean())
print(sales[sales < 1000].mean())Expected output
288.0
110.0The comparison measures sensitivity; it does not prove the large order is invalid.
10. Calculate and interpret correlation
import pandas as pd
df = pd.DataFrame({"AdSpend": [100, 200, 300, 400, 500], "Revenue": [250, 300, 450, 410, 650]})
print(round(df["AdSpend"].corr(df["Revenue"]), 3))
print("This association does not establish causation.")Expected output
0.925
This association does not establish causation.A numerical association and a causal effect answer different questions.