Chapter 3 · Section 2 practice
Fitting and Interpreting a Regression Line
Work from question 1 to 10. Edit your Python box, click Run Code, review the result, then click Submit attempt. Reference answers unlock after all ten attempts are submitted. This records completion, not correctness. Each box runs independently. Progress is saved in this browser.
1. Read the intercept
Warm-up. Print the estimated intercept rounded to three decimals.
Write your answer, then run it.
2. Read the slope
Warm-up. Print the Ads coefficient rounded to three decimals.
Write your answer, then run it.
3. Compute regression RMSE
Warm-up. Print the SLR regression RMSE using n-2 and explain why two degrees of freedom are used.
Write your answer, then run it.
4. Calculate one residual
Build your skills. Print observed minus fitted revenue for the first week.
Write your answer, then run it.
5. Calculate SSE
Build your skills. Square and sum residuals; print SSE.
Write your answer, then run it.
6. Predict a new input
Build your skills. Predict Revenue at Ads=4 with a one-row design table.
Write your answer, then run it.
7. Decompose variation
Build your skills. Calculate SST, SSR and SSE for the fitted OLS model; verify SST = SSR + SSE and print R-squared.
Write your answer, then run it.
8. Compare a proposed line
Challenge. Compare fitted SSE against SSE for the line Revenue=10+2*Ads.
Write your answer, then run it.
9. Find the largest deviation
Challenge. Print the index and signed value of the residual with greatest absolute size.
Write your answer, then run it.
10. Change predictor units
Challenge. Fit using Ads in HKD rather than thousands; print the new slope and compare it with the old slope divided by 1000.
Write your answer, then run it.
Reference answers are locked until all 10 attempts are submitted.
1. Read the intercept
import pandas as pd
sales = pd.DataFrame({
"Ads": [1, 2, 3, 4, 5, 6, 7, 8],
"Revenue": [12, 15, 14, 20, 19, 24, 25, 27],
"Price": [9, 8, 10, 7, 9, 6, 7, 6]
})
import statsmodels.api as sm
X = sm.add_constant(sales[["Ads"]])
model = sm.OLS(sales["Revenue"], X).fit()
print(round(model.params["const"], 3))Expected output
9.643This is fitted revenue at zero advertising, in HKD thousands.
2. Read the slope
import pandas as pd
sales = pd.DataFrame({
"Ads": [1, 2, 3, 4, 5, 6, 7, 8],
"Revenue": [12, 15, 14, 20, 19, 24, 25, 27],
"Price": [9, 8, 10, 7, 9, 6, 7, 6]
})
import statsmodels.api as sm
X = sm.add_constant(sales[["Ads"]])
model = sm.OLS(sales["Revenue"], X).fit()
print(round(model.params["Ads"], 3))Expected output
2.19Units are revenue HKD thousands per advertising HKD thousand.
3. Compute regression RMSE
import pandas as pd
sales = pd.DataFrame({
"Ads": [1, 2, 3, 4, 5, 6, 7, 8],
"Revenue": [12, 15, 14, 20, 19, 24, 25, 27],
"Price": [9, 8, 10, 7, 9, 6, 7, 6]
})
import statsmodels.api as sm
X = sm.add_constant(sales[["Ads"]])
model = sm.OLS(sales["Revenue"], X).fit()
import numpy as np
n = len(sales)
sse = (model.resid**2).sum()
print(round(np.sqrt(sse / (n - 2)), 3))
print(
"Two fitted coefficients: intercept and slope. Regression RMSE is residual standard error."
)Expected output
1.442
Two fitted coefficients: intercept and slope. Regression RMSE is residual standard error.Divide SSE by the residual degrees of freedom n-2, then take the square root.
4. Calculate one residual
import pandas as pd
sales = pd.DataFrame({
"Ads": [1, 2, 3, 4, 5, 6, 7, 8],
"Revenue": [12, 15, 14, 20, 19, 24, 25, 27],
"Price": [9, 8, 10, 7, 9, 6, 7, 6]
})
import statsmodels.api as sm
X = sm.add_constant(sales[["Ads"]])
model = sm.OLS(sales["Revenue"], X).fit()
print(round(sales.loc[0, "Revenue"] - model.fittedvalues.loc[0], 3))Expected output
0.167The subtraction order determines the sign.
5. Calculate SSE
import pandas as pd
sales = pd.DataFrame({
"Ads": [1, 2, 3, 4, 5, 6, 7, 8],
"Revenue": [12, 15, 14, 20, 19, 24, 25, 27],
"Price": [9, 8, 10, 7, 9, 6, 7, 6]
})
import statsmodels.api as sm
X = sm.add_constant(sales[["Ads"]])
model = sm.OLS(sales["Revenue"], X).fit()
print(round((model.resid ** 2).sum(), 3))Expected output
12.476SSE is an aggregate squared deviation, not a revenue amount.
6. Predict a new input
import pandas as pd
sales = pd.DataFrame({
"Ads": [1, 2, 3, 4, 5, 6, 7, 8],
"Revenue": [12, 15, 14, 20, 19, 24, 25, 27],
"Price": [9, 8, 10, 7, 9, 6, 7, 6]
})
import statsmodels.api as sm
X = sm.add_constant(sales[["Ads"]])
model = sm.OLS(sales["Revenue"], X).fit()
new = pd.DataFrame({"const": [1], "Ads": [4]})
print(round(model.predict(new).iloc[0], 3))Expected output
18.405New input columns match the fitted constant and Ads columns.
7. Decompose variation
import pandas as pd
sales = pd.DataFrame({
"Ads": [1, 2, 3, 4, 5, 6, 7, 8],
"Revenue": [12, 15, 14, 20, 19, 24, 25, 27],
"Price": [9, 8, 10, 7, 9, 6, 7, 6]
})
import statsmodels.api as sm
X = sm.add_constant(sales[["Ads"]])
model = sm.OLS(sales["Revenue"], X).fit()
import numpy as np
y = sales["Revenue"]
sst = ((y - y.mean()) ** 2).sum()
ssr = ((model.fittedvalues - y.mean()) ** 2).sum()
sse = (model.resid ** 2).sum()
print(round(sst, 3), round(ssr, 3), round(sse, 3))
print(np.isclose(sst, ssr + sse))
print(round(ssr / sst, 4))Expected output
214.0 201.524 12.476
True
0.9417SSR is explained regression variation here; the identity uses the OLS training fit with an intercept.
8. Compare a proposed line
import pandas as pd
sales = pd.DataFrame({
"Ads": [1, 2, 3, 4, 5, 6, 7, 8],
"Revenue": [12, 15, 14, 20, 19, 24, 25, 27],
"Price": [9, 8, 10, 7, 9, 6, 7, 6]
})
import statsmodels.api as sm
X = sm.add_constant(sales[["Ads"]])
model = sm.OLS(sales["Revenue"], X).fit()
other = sales["Revenue"] - (10 + 2 * sales["Ads"])
print(round((model.resid ** 2).sum(), 3))
print(round((other ** 2).sum(), 3))Expected output
12.476
16The OLS SSE is no larger than that of the proposed line.
9. Find the largest deviation
import pandas as pd
sales = pd.DataFrame({
"Ads": [1, 2, 3, 4, 5, 6, 7, 8],
"Revenue": [12, 15, 14, 20, 19, 24, 25, 27],
"Price": [9, 8, 10, 7, 9, 6, 7, 6]
})
import statsmodels.api as sm
X = sm.add_constant(sales[["Ads"]])
model = sm.OLS(sales["Revenue"], X).fit()
position = model.resid.abs().idxmax()
print(position, round(model.resid.loc[position], 3))Expected output
2 -2.214Largest absolute size can refer to a negative residual.
10. Change predictor units
import pandas as pd
sales = pd.DataFrame({
"Ads": [1, 2, 3, 4, 5, 6, 7, 8],
"Revenue": [12, 15, 14, 20, 19, 24, 25, 27],
"Price": [9, 8, 10, 7, 9, 6, 7, 6]
})
import statsmodels.api as sm
X = sm.add_constant(sales[["Ads"]])
model = sm.OLS(sales["Revenue"], X).fit()
sales["AdsHKD"] = sales["Ads"] * 1000
changed = sm.OLS(sales["Revenue"], sm.add_constant(sales[["AdsHKD"]])).fit()
print(round(changed.params["AdsHKD"], 6))
print(round(model.params["Ads"] / 1000, 6))Expected output
0.00219
0.00219Units change the coefficient even though fitted revenue stays the same.