Chapter 3 · Section 6 practice
Residual Analysis and Unusual Observations
Work from question 1 to 10. Edit your Python box, click Run Code, review the result, then click Submit attempt. Reference answers unlock after all ten attempts are submitted. This records completion, not correctness. Each box runs independently. Progress is saved in this browser.
1. Read residuals
Warm-up. Print the first three residuals.
Write your answer, then run it.
2. Check residual mean
Warm-up. Print the mean residual rounded to six decimals.
Write your answer, then run it.
3. Find the largest residual magnitude
Warm-up. Print the week index with greatest absolute residual.
Write your answer, then run it.
4. Read the unusual record
Build your skills. Print the complete row for the largest absolute residual.
Write your answer, then run it.
5. Read leverage
Build your skills. Print all leverage values rounded to three decimals.
Write your answer, then run it.
6. Read Cook’s distance
Build your skills. Print the index with highest Cook’s distance.
Write your answer, then run it.
7. Choose a time-order diagnostic
Build your skills. Print the plot used to investigate error independence, and one warning pattern.
Write your answer, then run it.
8. Compare a leave-one-out fit
Challenge. Remove the row with index 7 (the eighth week), refit, and print the original and revised slopes.
Write your answer, then run it.
9. Match assumptions to plots
Challenge. Print the visual check for each of the four LINE assumptions.
Write your answer, then run it.
10. Separate a diagnostic from a guarantee
Challenge. Print residual mean and a sentence explaining why it cannot verify the population zero-conditional-mean assumption.
Write your answer, then run it.
Reference answers are locked until all 10 attempts are submitted.
1. Read residuals
import pandas as pd
sales = pd.DataFrame({
"Ads": [1, 2, 3, 4, 5, 6, 7, 8],
"Revenue": [12, 15, 14, 20, 19, 24, 25, 27],
"Price": [9, 8, 10, 7, 9, 6, 7, 6]
})
import statsmodels.api as sm
X = sm.add_constant(sales[["Ads"]])
model = sm.OLS(sales["Revenue"], X).fit()
print(model.resid.head(3).round(3))Expected output
0 0.167
1 0.976
2 -2.214
dtype: float64Residuals are observed minus fitted outcomes.
2. Check residual mean
import pandas as pd
sales = pd.DataFrame({
"Ads": [1, 2, 3, 4, 5, 6, 7, 8],
"Revenue": [12, 15, 14, 20, 19, 24, 25, 27],
"Price": [9, 8, 10, 7, 9, 6, 7, 6]
})
import statsmodels.api as sm
X = sm.add_constant(sales[["Ads"]])
model = sm.OLS(sales["Revenue"], X).fit()
print(round(model.resid.mean(), 6))Expected output
0.0OLS with an intercept makes the sample residual mean essentially zero; this does not prove exogeneity.
3. Find the largest residual magnitude
import pandas as pd
sales = pd.DataFrame({
"Ads": [1, 2, 3, 4, 5, 6, 7, 8],
"Revenue": [12, 15, 14, 20, 19, 24, 25, 27],
"Price": [9, 8, 10, 7, 9, 6, 7, 6]
})
import statsmodels.api as sm
X = sm.add_constant(sales[["Ads"]])
model = sm.OLS(sales["Revenue"], X).fit()
print(model.resid.abs().idxmax())Expected output
2Absolute size differs from choosing the largest signed residual.
4. Read the unusual record
import pandas as pd
sales = pd.DataFrame({
"Ads": [1, 2, 3, 4, 5, 6, 7, 8],
"Revenue": [12, 15, 14, 20, 19, 24, 25, 27],
"Price": [9, 8, 10, 7, 9, 6, 7, 6]
})
import statsmodels.api as sm
X = sm.add_constant(sales[["Ads"]])
model = sm.OLS(sales["Revenue"], X).fit()
week = model.resid.abs().idxmax()
print(sales.loc[week])Expected output
Ads 3
Revenue 14
Price 10
Name: 2, dtype: int64Context is needed before deciding that a row is erroneous.
5. Read leverage
import pandas as pd
sales = pd.DataFrame({
"Ads": [1, 2, 3, 4, 5, 6, 7, 8],
"Revenue": [12, 15, 14, 20, 19, 24, 25, 27],
"Price": [9, 8, 10, 7, 9, 6, 7, 6]
})
import statsmodels.api as sm
X = sm.add_constant(sales[["Ads"]])
model = sm.OLS(sales["Revenue"], X).fit()
influence = model.get_influence()
print(pd.Series(influence.hat_matrix_diag).round(3).to_list())Expected output
[0.417, 0.274, 0.179, 0.131, 0.131, 0.179, 0.274, 0.417]Leverage depends on predictor positions, not directly on residual size.
6. Read Cook's distance
import pandas as pd
sales = pd.DataFrame({
"Ads": [1, 2, 3, 4, 5, 6, 7, 8],
"Revenue": [12, 15, 14, 20, 19, 24, 25, 27],
"Price": [9, 8, 10, 7, 9, 6, 7, 6]
})
import statsmodels.api as sm
X = sm.add_constant(sales[["Ads"]])
model = sm.OLS(sales["Revenue"], X).fit()
influence = model.get_influence()
cook = pd.Series(influence.cooks_distance[0])
print(cook.idxmax())Expected output
2Influence combines predictor position and deviation from the fitted line.
7. Choose a time-order diagnostic
import pandas as pd
sales = pd.DataFrame({
"Ads": [1, 2, 3, 4, 5, 6, 7, 8],
"Revenue": [12, 15, 14, 20, 19, 24, 25, 27],
"Price": [9, 8, 10, 7, 9, 6, 7, 6]
})
import statsmodels.api as sm
X = sm.add_constant(sales[["Ads"]])
model = sm.OLS(sales["Revenue"], X).fit()
print("Plot residuals against time or meaningful observation order.")
print(
"Long runs above or below zero can warn of dependence; a random-looking plot does not prove independence."
)Expected output
Plot residuals against time or meaningful observation order.
Long runs above or below zero can warn of dependence; a random-looking plot does not prove independence.The time-order plot has already been introduced; no new method is required.
8. Compare a leave-one-out fit
import pandas as pd
sales = pd.DataFrame({
"Ads": [1, 2, 3, 4, 5, 6, 7, 8],
"Revenue": [12, 15, 14, 20, 19, 24, 25, 27],
"Price": [9, 8, 10, 7, 9, 6, 7, 6]
})
import statsmodels.api as sm
X = sm.add_constant(sales[["Ads"]])
model = sm.OLS(sales["Revenue"], X).fit()
reduced = sales.drop(index=7)
changed = sm.OLS(reduced["Revenue"], sm.add_constant(reduced[["Ads"]])).fit()
print(round(model.params["Ads"], 4), round(changed.params["Ads"], 4))Expected output
2.1905 2.2143A sensitivity comparison is informative; it does not justify deleting the week.
9. Match assumptions to plots
import pandas as pd
sales = pd.DataFrame({
"Ads": [1, 2, 3, 4, 5, 6, 7, 8],
"Revenue": [12, 15, 14, 20, 19, 24, 25, 27],
"Price": [9, 8, 10, 7, 9, 6, 7, 6]
})
import statsmodels.api as sm
X = sm.add_constant(sales[["Ads"]])
model = sm.OLS(sales["Revenue"], X).fit()
print("L: residuals versus fitted values; look for curvature")
print("I: residuals versus time or meaningful order; look for runs or cycles")
print("N: normal Q-Q plot of residuals; look for departures from the reference line")
print("E: residuals or absolute residuals versus fitted values; look for changing spread")Expected output
L: residuals versus fitted values; look for curvature
I: residuals versus time or meaningful order; look for runs or cycles
N: normal Q-Q plot of residuals; look for departures from the reference line
E: residuals or absolute residuals versus fitted values; look for changing spreadPlots can reveal warning patterns but cannot prove that the assumptions hold.
10. Separate a diagnostic from a guarantee
import pandas as pd
sales = pd.DataFrame({
"Ads": [1, 2, 3, 4, 5, 6, 7, 8],
"Revenue": [12, 15, 14, 20, 19, 24, 25, 27],
"Price": [9, 8, 10, 7, 9, 6, 7, 6]
})
import statsmodels.api as sm
X = sm.add_constant(sales[["Ads"]])
model = sm.OLS(sales["Revenue"], X).fit()
print(round(model.resid.mean(), 6))
print("The fitted residual mean is an OLS identity, not proof of E(error | X)=0.")Expected output
0.0
The fitted residual mean is an OLS identity, not proof of E(error | X)=0.A sample fitting identity is not evidence that confounding is absent.