Chapter 3 quiz
Ten questions to bring the chapter together.
Work from question 1 to 10. Edit your Python box, click Run Code, review the result, then click Submit attempt. Reference answers unlock after all ten attempts are submitted. This records completion, not correctness. Each box runs independently. Progress is saved in this browser.
Question 1
Print the correlation between LocationOffset and TravelCost. Compare it with correlation between squared LocationOffset and TravelCost. Explain whether the first result justifies ignoring location.
Write your answer, then run it.
Question 2
Fit Revenue on Ads with an intercept and print coefficients.
Write your answer, then run it.
Question 3
Print the first week’s observed, fitted, and residual revenues.
Write your answer, then run it.
Question 4
Print CAPM beta and its 95% coefficient interval.
Write your answer, then run it.
Question 5
Test H0: CAPM beta=1 against a two-sided alternative; print its p-value.
Write your answer, then run it.
Question 6
At retail Ads=5, print a mean CI and individual PI.
Write your answer, then run it.
Question 7
Print the retail record with the largest absolute residual.
Write your answer, then run it.
Question 8
At Price=8, compare multiple linear regression predictions at Ads=3 and Ads=4.
Write your answer, then run it.
Question 9
For the retail SLR, print SST, SSR, SSE, R-squared and regression RMSE using n-2; state one limitation of using training fit to select predictors.
Write your answer, then run it.
Question 10
Print CAPM total-return predictions for market excess scenarios -1% and +1% with RF=0.0002. Explain the prediction scope.
Write your answer, then run it.
Reference answers are locked until all 10 attempts are submitted.
Question 1
import pandas as pd
visits = pd.DataFrame({"LocationOffset": [-3, -2, -1, 0, 1, 2, 3]})
visits["TravelCost"] = 5 + visits["LocationOffset"] ** 2
print(round(visits["LocationOffset"].corr(visits["TravelCost"]), 3))
print(round((visits["LocationOffset"] ** 2).corr(visits["TravelCost"]), 3))
print(
"Location matters through a U-shaped cost relationship even though its Pearson correlation is zero."
)Expected output
0.0
1.0
Location matters through a U-shaped cost relationship even though its Pearson correlation is zero.A single correlation summarises linear association and can miss a curved relationship.
Question 2
import pandas as pd
sales = pd.DataFrame({
"Ads": [1, 2, 3, 4, 5, 6, 7, 8],
"Revenue": [12, 15, 14, 20, 19, 24, 25, 27],
"Price": [9, 8, 10, 7, 9, 6, 7, 6]
})
import statsmodels.api as sm
X = sm.add_constant(sales[["Ads"]])
model = sm.OLS(sales["Revenue"], X).fit()
print(model.params.round(3))Expected output
const 9.643
Ads 2.190
dtype: float64The intercept is explicitly included.
Question 3
import pandas as pd
sales = pd.DataFrame({
"Ads": [1, 2, 3, 4, 5, 6, 7, 8],
"Revenue": [12, 15, 14, 20, 19, 24, 25, 27],
"Price": [9, 8, 10, 7, 9, 6, 7, 6]
})
import statsmodels.api as sm
X = sm.add_constant(sales[["Ads"]])
model = sm.OLS(sales["Revenue"], X).fit()
print(
round(sales.loc[0, "Revenue"], 3),
round(model.fittedvalues.loc[0], 3),
round(model.resid.loc[0], 3),
)Expected output
12 11.833 0.167Residual is observed minus fitted.
Question 4
import pandas as pd
prices = pd.read_csv(
"https://busanalytics-book.pages.dev/data/nvda_spy_daily_2023_2024.csv",
index_col="Date", parse_dates=True).sort_index()
factors = pd.read_csv(
"https://busanalytics-book.pages.dev/data/ff5_daily_2023_2024.csv",
index_col="Date", parse_dates=True) / 100
returns = prices.pct_change(fill_method=None)
data = returns.copy()
# Series columns match the Date index, not the row position.
data["Mkt-RF"] = factors["Mkt-RF"]
data["SMB"] = factors["SMB"]
data["HML"] = factors["HML"]
data["RMW"] = factors["RMW"]
data["CMA"] = factors["CMA"]
data["RF"] = factors["RF"]
# Keep complete dates for all model comparisons.
data = data.dropna()
data["Excess"] = data["NVDA"] - data["RF"]
data["Market"] = data["SPY"] - data["RF"]
cut = int(len(data) * 0.8)
train = data.iloc[:cut].copy()
test = data.iloc[cut:].copy()
import statsmodels.api as sm
X = sm.add_constant(train[["Market"]])
model = sm.OLS(train["Excess"], X).fit()
print(round(model.params["Market"], 3))
print(model.conf_int().loc["Market"].round(3).to_list())Expected output
2.332
[2.005, 2.659]Beta is a return sensitivity and the interval concerns that coefficient.
Question 5
import pandas as pd
prices = pd.read_csv(
"https://busanalytics-book.pages.dev/data/nvda_spy_daily_2023_2024.csv",
index_col="Date", parse_dates=True).sort_index()
factors = pd.read_csv(
"https://busanalytics-book.pages.dev/data/ff5_daily_2023_2024.csv",
index_col="Date", parse_dates=True) / 100
returns = prices.pct_change(fill_method=None)
data = returns.copy()
# Series columns match the Date index, not the row position.
data["Mkt-RF"] = factors["Mkt-RF"]
data["SMB"] = factors["SMB"]
data["HML"] = factors["HML"]
data["RMW"] = factors["RMW"]
data["CMA"] = factors["CMA"]
data["RF"] = factors["RF"]
# Keep complete dates for all model comparisons.
data = data.dropna()
data["Excess"] = data["NVDA"] - data["RF"]
data["Market"] = data["SPY"] - data["RF"]
cut = int(len(data) * 0.8)
train = data.iloc[:cut].copy()
test = data.iloc[cut:].copy()
import statsmodels.api as sm
X = sm.add_constant(train[["Market"]])
model = sm.OLS(train["Excess"], X).fit()
from scipy import stats
t_stat = (model.params["Market"] - 1) / model.bse["Market"]
print(round(2 * stats.t.sf(abs(t_stat), model.df_resid), 6))Expected output
0.0Use the null value 1 and count both tails.
Question 6
import pandas as pd
sales = pd.DataFrame({
"Ads": [1, 2, 3, 4, 5, 6, 7, 8],
"Revenue": [12, 15, 14, 20, 19, 24, 25, 27],
"Price": [9, 8, 10, 7, 9, 6, 7, 6]
})
import statsmodels.api as sm
X = sm.add_constant(sales[["Ads"]])
model = sm.OLS(sales["Revenue"], X).fit()
new = pd.DataFrame({"const": [1], "Ads": [5]})
p = model.get_prediction(new).summary_frame()
print(p[["mean_ci_lower", "mean_ci_upper", "obs_ci_lower", "obs_ci_upper"]].round(3))Expected output
mean_ci_lower mean_ci_upper obs_ci_lower obs_ci_upper
0 19.318 21.872 16.843 24.348The individual interval includes additional outcome variation.
Question 7
import pandas as pd
sales = pd.DataFrame({
"Ads": [1, 2, 3, 4, 5, 6, 7, 8],
"Revenue": [12, 15, 14, 20, 19, 24, 25, 27],
"Price": [9, 8, 10, 7, 9, 6, 7, 6]
})
import statsmodels.api as sm
X = sm.add_constant(sales[["Ads"]])
model = sm.OLS(sales["Revenue"], X).fit()
print(sales.loc[model.resid.abs().idxmax()])Expected output
Ads 3
Revenue 14
Price 10
Name: 2, dtype: int64Investigate context before excluding an observation.
Question 8
import pandas as pd
sales = pd.DataFrame({
"Ads": [1, 2, 3, 4, 5, 6, 7, 8],
"Revenue": [12, 15, 14, 20, 19, 24, 25, 27],
"Price": [9, 8, 10, 7, 9, 6, 7, 6]
})
import statsmodels.api as sm
X = sm.add_constant(sales[["Ads", "Price"]])
model = sm.OLS(sales["Revenue"], X).fit()
new = pd.DataFrame({"const": [1, 1], "Ads": [3, 4], "Price": [8, 8]})
p = model.predict(new)
print(p.round(3).to_list())
print(round(p.iloc[1] - p.iloc[0], 3))Expected output
[16.697, 18.361]
1.664A fitted controlled comparison is not necessarily a causal effect.
Question 9
import pandas as pd
sales = pd.DataFrame({
"Ads": [1, 2, 3, 4, 5, 6, 7, 8],
"Revenue": [12, 15, 14, 20, 19, 24, 25, 27],
"Price": [9, 8, 10, 7, 9, 6, 7, 6]
})
import statsmodels.api as sm
X = sm.add_constant(sales[["Ads"]])
model = sm.OLS(sales["Revenue"], X).fit()
import numpy as np
y = sales["Revenue"]
sst = ((y - y.mean()) ** 2).sum()
ssr = ((model.fittedvalues - y.mean()) ** 2).sum()
sse = (model.resid**2).sum()
print(round(sst, 3), round(ssr, 3), round(sse, 3))
print(round(1 - sse / sst, 4), round(np.sqrt(sse / (len(y) - 2)), 3))
print(
"Extra predictors can improve training fit while worsening prediction on new records."
)Expected output
214.0 201.524 12.476
0.9417 1.442
Extra predictors can improve training fit while worsening prediction on new records.The decomposition uses OLS with an intercept. Held-out error is needed to judge transfer to new records.
Question 10
import pandas as pd
prices = pd.read_csv(
"https://busanalytics-book.pages.dev/data/nvda_spy_daily_2023_2024.csv",
index_col="Date", parse_dates=True).sort_index()
factors = pd.read_csv(
"https://busanalytics-book.pages.dev/data/ff5_daily_2023_2024.csv",
index_col="Date", parse_dates=True) / 100
returns = prices.pct_change(fill_method=None)
data = returns.copy()
# Series columns match the Date index, not the row position.
data["Mkt-RF"] = factors["Mkt-RF"]
data["SMB"] = factors["SMB"]
data["HML"] = factors["HML"]
data["RMW"] = factors["RMW"]
data["CMA"] = factors["CMA"]
data["RF"] = factors["RF"]
# Keep complete dates for all model comparisons.
data = data.dropna()
data["Excess"] = data["NVDA"] - data["RF"]
data["Market"] = data["SPY"] - data["RF"]
cut = int(len(data) * 0.8)
train = data.iloc[:cut].copy()
test = data.iloc[cut:].copy()
import statsmodels.api as sm
X = sm.add_constant(train[["Market"]])
model = sm.OLS(train["Excess"], X).fit()
new = pd.DataFrame({"const": [1, 1], "Market": [-0.01, 0.01]})
print(((model.predict(new) + 0.0002) * 100).round(3).to_list())
print("These are conditional market scenarios, not guaranteed tomorrow returns.")Expected output
[-1.95, 2.714]
These are conditional market scenarios, not guaranteed tomorrow returns.Inputs must be supplied or forecast before a genuine future prediction is possible.