Chapter 3 · Section 7 practice
Multiple Linear Regression and Multicollinearity
Work from question 1 to 10. Edit your Python box, click Run Code, review the result, then click Submit attempt. Reference answers unlock after all ten attempts are submitted. This records completion, not correctness. Each box runs independently. Progress is saved in this browser.
1. Read coefficients
Warm-up. Print all three coefficients rounded to three decimals.
Write your answer, then run it.
2. Read model degrees of freedom
Warm-up. Print residual degrees of freedom.
Write your answer, then run it.
3. Predict a complete scenario
Warm-up. Predict Revenue for Ads=4 and Price=8.
Write your answer, then run it.
4. Hold price fixed
Build your skills. Predict Ads=3 and Ads=4 at Price=8, and print the difference.
Write your answer, then run it.
5. Reconstruct adjusted R-squared
Build your skills. Use n=8 and k=2 excluding the intercept to reconstruct adjusted R-squared; compare with statsmodels.
Write your answer, then run it.
6. Inspect predictor correlation
Build your skills. Print the Ads–Price correlation.
Write your answer, then run it.
7. Calculate a predictor VIF
Build your skills. Regress Ads on Price with an intercept; reconstruct Ads VIF.
Write your answer, then run it.
8. Interpret the overall and individual tests
Challenge. State the overall F-test hypotheses for Ads and Price; print its p-value and the individual Ads p-value, then print the decisions at 5%.
Write your answer, then run it.
9. Compare nested fit and errors
Challenge. Fit Ads-only on the same rows. Print R-squared, adjusted R-squared and RMSE s for each model, using n-2 for SLR and n-3 for this MLR. Explain why training improvement alone is insufficient.
Write your answer, then run it.
10. Recognise exact dependence
Challenge. Create AdsTwice=2*Ads, print Ads and AdsTwice, and explain why keeping both prevents separate slope estimation.
Write your answer, then run it.
Reference answers are locked until all 10 attempts are submitted.
1. Read coefficients
import pandas as pd
sales = pd.DataFrame({
"Ads": [1, 2, 3, 4, 5, 6, 7, 8],
"Revenue": [12, 15, 14, 20, 19, 24, 25, 27],
"Price": [9, 8, 10, 7, 9, 6, 7, 6]
})
import statsmodels.api as sm
X = sm.add_constant(sales[["Ads", "Price"]])
model = sm.OLS(sales["Revenue"], X).fit()
print(model.params.round(3))Expected output
const 21.541
Ads 1.664
Price -1.229
dtype: float64There is one intercept and a slope for each predictor.
2. Read model degrees of freedom
import pandas as pd
sales = pd.DataFrame({
"Ads": [1, 2, 3, 4, 5, 6, 7, 8],
"Revenue": [12, 15, 14, 20, 19, 24, 25, 27],
"Price": [9, 8, 10, 7, 9, 6, 7, 6]
})
import statsmodels.api as sm
X = sm.add_constant(sales[["Ads", "Price"]])
model = sm.OLS(sales["Revenue"], X).fit()
print(int(model.df_resid))Expected output
5Eight observations minus three fitted coefficients leaves five.
3. Predict a complete scenario
import pandas as pd
sales = pd.DataFrame({
"Ads": [1, 2, 3, 4, 5, 6, 7, 8],
"Revenue": [12, 15, 14, 20, 19, 24, 25, 27],
"Price": [9, 8, 10, 7, 9, 6, 7, 6]
})
import statsmodels.api as sm
X = sm.add_constant(sales[["Ads", "Price"]])
model = sm.OLS(sales["Revenue"], X).fit()
new = pd.DataFrame({"const": [1], "Ads": [4], "Price": [8]})
print(round(model.predict(new).iloc[0], 3))Expected output
18.361Both features must be supplied.
4. Hold price fixed
import pandas as pd
sales = pd.DataFrame({
"Ads": [1, 2, 3, 4, 5, 6, 7, 8],
"Revenue": [12, 15, 14, 20, 19, 24, 25, 27],
"Price": [9, 8, 10, 7, 9, 6, 7, 6]
})
import statsmodels.api as sm
X = sm.add_constant(sales[["Ads", "Price"]])
model = sm.OLS(sales["Revenue"], X).fit()
new = pd.DataFrame({"const": [1, 1], "Ads": [3, 4], "Price": [8, 8]})
p = model.predict(new)
print(round(p.iloc[1] - p.iloc[0], 4))Expected output
1.6636The one-unit Ads difference at fixed Price equals its fitted slope.
5. Reconstruct adjusted R-squared
import pandas as pd
sales = pd.DataFrame({
"Ads": [1, 2, 3, 4, 5, 6, 7, 8],
"Revenue": [12, 15, 14, 20, 19, 24, 25, 27],
"Price": [9, 8, 10, 7, 9, 6, 7, 6]
})
import statsmodels.api as sm
X = sm.add_constant(sales[["Ads", "Price"]])
model = sm.OLS(sales["Revenue"], X).fit()
n, k = len(sales), 2
adjusted = 1 - (1 - model.rsquared) * (n - 1) / (n - k - 1)
print(round(adjusted, 4), round(model.rsquared_adj, 4))Expected output
0.9954 0.9954Ordinary R-squared does not penalise extra predictors; adjusted R-squared uses residual degrees of freedom.
6. Inspect predictor correlation
import pandas as pd
sales = pd.DataFrame({
"Ads": [1, 2, 3, 4, 5, 6, 7, 8],
"Revenue": [12, 15, 14, 20, 19, 24, 25, 27],
"Price": [9, 8, 10, 7, 9, 6, 7, 6]
})
import statsmodels.api as sm
X = sm.add_constant(sales[["Ads", "Price"]])
model = sm.OLS(sales["Revenue"], X).fit()
print(round(sales["Ads"].corr(sales["Price"]), 4))Expected output
-0.7055This pairwise check describes predictor dependence.
7. Calculate a predictor VIF
import pandas as pd
sales = pd.DataFrame({
"Ads": [1, 2, 3, 4, 5, 6, 7, 8],
"Revenue": [12, 15, 14, 20, 19, 24, 25, 27],
"Price": [9, 8, 10, 7, 9, 6, 7, 6]
})
import statsmodels.api as sm
X = sm.add_constant(sales[["Ads", "Price"]])
model = sm.OLS(sales["Revenue"], X).fit()
aux = sm.OLS(sales["Ads"], sm.add_constant(sales[["Price"]])).fit()
print(round(1 / (1 - aux.rsquared), 3))Expected output
1.991The auxiliary model outcome is Ads, not Revenue.
8. Interpret the overall and individual tests
import pandas as pd
sales = pd.DataFrame({
"Ads": [1, 2, 3, 4, 5, 6, 7, 8],
"Revenue": [12, 15, 14, 20, 19, 24, 25, 27],
"Price": [9, 8, 10, 7, 9, 6, 7, 6]
})
import statsmodels.api as sm
X = sm.add_constant(sales[["Ads", "Price"]])
model = sm.OLS(sales["Revenue"], X).fit()
print("F: H0 both slopes are zero; H1 at least one slope is nonzero.")
print(
"Ads t: H0 Ads slope is zero; H1 Ads slope differs from zero, keeping Price in the model."
)
print(round(model.f_pvalue, 6), round(model.pvalues["Ads"], 6))
print("Reject overall H0:", bool(model.f_pvalue < 0.05))
print("Reject Ads H0:", bool(model.pvalues["Ads"] < 0.05))Expected output
F: H0 both slopes are zero; H1 at least one slope is nonzero.
Ads t: H0 Ads slope is zero; H1 Ads slope differs from zero, keeping Price in the model.
1e-06 5e-06
Reject overall H0: True
Reject Ads H0: TrueThe overall result concerns the predictor set. The Ads test concerns its conditional slope. Neither proves a causal effect.
9. Compare nested fit and errors
import pandas as pd
sales = pd.DataFrame({
"Ads": [1, 2, 3, 4, 5, 6, 7, 8],
"Revenue": [12, 15, 14, 20, 19, 24, 25, 27],
"Price": [9, 8, 10, 7, 9, 6, 7, 6]
})
import statsmodels.api as sm
X = sm.add_constant(sales[["Ads", "Price"]])
model = sm.OLS(sales["Revenue"], X).fit()
import numpy as np
simple = sm.OLS(sales["Revenue"], sm.add_constant(sales[["Ads"]])).fit()
print(
round(simple.rsquared, 4),
round(simple.rsquared_adj, 4),
round(np.sqrt(simple.mse_resid), 3),
)
print(
round(model.rsquared, 4),
round(model.rsquared_adj, 4),
round(np.sqrt(model.mse_resid), 3),
)
print("Compare validation error before claiming improved prediction on new records.")Expected output
0.9417 0.932 1.442
0.9967 0.9954 0.377
Compare validation error before claiming improved prediction on new records.Training R-squared cannot fall. RMSE s accounts for the estimated coefficients and may rise or fall.
10. Recognise exact dependence
import pandas as pd
sales = pd.DataFrame({
"Ads": [1, 2, 3, 4, 5, 6, 7, 8],
"Revenue": [12, 15, 14, 20, 19, 24, 25, 27],
"Price": [9, 8, 10, 7, 9, 6, 7, 6]
})
import statsmodels.api as sm
X = sm.add_constant(sales[["Ads", "Price"]])
model = sm.OLS(sales["Revenue"], X).fit()
sales["AdsTwice"] = 2 * sales["Ads"]
print(sales[["Ads", "AdsTwice"]])
print(
"AdsTwice contains exactly the same information as Ads. Keeping both makes their separate slopes unidentifiable."
)Expected output
Ads AdsTwice
0 1 2
1 2 4
2 3 6
3 4 8
4 5 10
5 6 12
6 7 14
7 8 16
AdsTwice contains exactly the same information as Ads. Keeping both makes their separate slopes unidentifiable.The second predictor is an exact multiple of the first; extra data with the same relation cannot separate their slopes.