Linear Regression Theory
Optional theory background: fitted lines, coefficient uncertainty, confidence intervals and prediction intervals. Read Appendix A first if these terms are new.
A sample mean estimates one average. A regression estimates an average at each input value.
Our classroom example: advertising spending \(x\) and weekly revenue \(Y\), both in HKD thousands. We will fit a line, describe coefficient uncertainty, then distinguish average revenue from one new week.
\[ Y=\beta_0+\beta_1x+\varepsilon \]
Thus the population mean is \(E(Y\mid x)=\beta_0+\beta_1x\). “\(\mid x\)” means at the given input \(x\).
At the same advertising input, revenue can differ because of weather, customer demand and other unobserved influences.
The straight line describes average revenue, while noise describes the spread around it.
\[ \hat{y}(x)=\hat{\beta}_0+\hat{\beta}_1x \]
A hat means “estimated from our sample”.
For this sample, \(\hat y(x)\approx9.6429+2.1905x\). Displayed coefficients are rounded; calculations use full precision. An extra HKD1,000 of Ads is associated with HKD2,190.5 higher fitted revenue. Association alone does not establish causation.
Each dot is one observed week. The line supplies a fitted value at that week’s input. The vertical gaps are residuals: observed minus fitted.
\[ e_i=y_i-\hat y(x_i),\qquad SSE=\sum_{i=1}^{n}e_i^2 \]
A residual above the line is positive; below is negative. Squaring stops the signs cancelling. Fitting by ordinary least squares (OLS) means choosing the line with the smallest SSE.
Move the intercept and slope. Watch the residual segments and SSE change. Click Animate OLS fit to move to the least-squares line.
The minimum is for the total squared error, not for each residual separately.
\[ \varepsilon_i=y_i-(\beta_0+\beta_1x_i),\qquad e_i=y_i-(\hat\beta_0+\hat\beta_1x_i) \]
The true population line is unknown, so we cannot observe its noise values directly. Residuals use the estimated line. They help us estimate the noise spread and examine assumptions.
| Assumption | Plain-language meaning |
|---|---|
| Linear mean | The population average follows a straight line |
| Independent errors | One observation’s error does not determine another’s |
| Constant variance | Noise has the same SD \(\sigma\) at every input |
| Normal errors | At a fixed input, errors follow a normal distribution |
These give exact small-sample t tests and intervals. The predictor \(x\) itself need not be normal; \(x\) must vary. Normality concerns errors conditional on \(x\).
\[ s=\mathrm{RMSE}=\sqrt{\frac{SSE}{n-2}} \]
This course uses RMSE for the residual standard error. It estimates the unknown noise SD \(\sigma\), in revenue units. We use \(n-2\) because simple linear regression estimates an intercept and a slope.
In our eight-week sample: \(SSE=12.4762\), so \(s=\sqrt{12.4762/6}=1.4420\) HKD thousands.
Take a new sample of weeks and refit: the fitted slope changes. The population slope stays fixed.
\(SE(\hat\beta_1)\) estimates the SD of slope estimates across repeated samples. It measures the precision of the estimated slope, not the spread of individual revenues.
\[ S_{xx}=\sum_{i=1}^{n}(x_i-\bar x)^2,\qquad SE(\hat\beta_1)=\frac{s}{\sqrt{S_{xx}}} \]
Our sample has \(S_{xx}=42\), so \(SE(\hat\beta_1)=1.4420/\sqrt{42}=0.2225\). Across repeated comparable samples, estimated slopes typically vary around the true slope on a scale of about 0.2225 revenue units per Ads unit.
\[ SE(\hat\beta_0)=s\sqrt{\frac1n+\frac{\bar x^2}{S_{xx}}} \]
The intercept is the estimated mean at \(x=0\). If the observed inputs are far from zero, it is harder to estimate reliably.
Here \(n=8\), \(\bar x=4.5\), \(S_{xx}=42\), giving \(SE(\hat\beta_0)=1.1236\). Do not give a business interpretation at zero if zero is outside the meaningful range.
\[ \hat\beta_j\ \pm\ t^*SE(\hat\beta_j),\qquad df=n-2 \]
\(j=0\) means intercept; \(j=1\) means slope. For our sample, \(df=6\) and a 95% interval uses \(t^*=2.4469\).
Slope CI: \(2.1905\pm2.4469(0.2225)=[1.6460,2.7349]\).
This estimates the population slope: the association is roughly HKD1,646 to HKD2,735 per extra HKD1,000 in Ads under the model.
\[ H_0:\beta_j=0,\qquad H_1:\beta_j\ne0 \]
\[ t_{\mathrm{obs}}=\frac{\hat\beta_j-0}{SE(\hat\beta_j)},\qquad df=n-2 \]
For the slope: \(t=2.1905/0.2225\approx9.8446\), with two-sided \(p\approx0.0000633\). Under a zero population slope and our model assumptions, such a large absolute t-statistic is very rare. Reject at 5%; this alone does not prove causation.
Keep \(df=6\) and adjust the estimated slope or its SE. Compare the two-sided p-value, 95% CI and the null value zero. A larger SE means less precise estimation.
These controls illustrate hypothetical regression outputs; they do not refit the eight observed weeks.
For a matching two-sided 5% t test and 95% t interval:
An interval shows both uncertainty and plausible effect sizes.
\[ Y=\beta_0+\beta_1x_1+\cdots+\beta_kx_k+\varepsilon \]
\(k\) is the number of predictors, excluding the intercept. Example: predict revenue using Ads and Price, so \(k=2\). Each slope describes a change in the average outcome holding the other predictors fixed.
The fitted equation replaces every coefficient by its sample estimate. No matrix algebra is needed to interpret it.
\[ s=\mathrm{RMSE}=\sqrt{\frac{SSE}{n-k-1}},\qquad df=n-k-1 \]
We estimate \(k\) slopes plus one intercept. With \(n=20\) observations and \(k=2\) predictors, \(df=17\).
For coefficient \(j\): CI is \(\hat\beta_j\pm t^*SE(\hat\beta_j)\) and the test uses \((\hat\beta_j-0)/SE(\hat\beta_j)\) with the same df. Software calculates the SEs; the simple-linear-regression slope SE formula does not apply unchanged.
| Test | Null | Alternative |
|---|---|---|
| Individual t | \(\beta_j=0\) | \(\beta_j\ne0\) |
| Overall F | \(\beta_1=\cdots=\beta_k=0\) | At least one slope is nonzero |
An individual t test asks about one predictor with the others included. The overall F test asks whether the predictor set helps explain the mean at all. A significant F test does not say that every coefficient is significant.
\[ x_0=4,\qquad\hat y_0=\hat y(4)\approx18.4048 \]
The subscript 0 labels the prediction case; it does not mean \(x_0=0\). Here the input is Ads=HKD4,000.
\(\hat y_0\) estimates average revenue at this input: HKD18,404.8. It is also our point prediction for one new week, but that week has extra uncertainty.
\[ SE(\hat y_0)=s\sqrt{\frac1n+\frac{(x_0-\bar x)^2}{S_{xx}}} \]
This measures how much our estimated average at \(x_0\) varies across repeated training samples. It is smallest near the average observed input \(\bar x\).
At \(x_0=4\): \(SE(\hat y_0)=0.5218\) HKD thousands. This is smaller than \(s=1.4420\), which measures individual noise.
\[ \hat y_0\ \pm\ t^*SE(\hat y_0) \]
For whom? The population average of all comparable weeks with Ads fixed at HKD4,000.
\(18.4048\pm2.4469(0.5218)=[17.1279,19.6816]\) HKD thousands.
This interval measures uncertainty about the mean response, not the variability of one future week.
A prediction for one future week has two sources of uncertainty:
A prediction interval includes both, so it is wider than a confidence interval for the average at the same input and confidence level.
\[ SE_{\mathrm{pred}}=s\sqrt{1+\frac1n+\frac{(x_0-\bar x)^2}{S_{xx}}} \]
\[ \text{95% PI}:\quad \hat y_0\ \pm\ t^*SE_{\mathrm{pred}} \]
The extra 1 accounts for the new week’s noise. At \(x_0=4\), the PI is \([14.6524,22.1571]\) HKD thousands. It targets one independent new response at the specified input, under the model assumptions.
Move \(x_0\) and change the confidence level. Compare the narrow mean CI with the wider individual PI. Both widen away from \(\bar x=4.5\).
Click Animate new weeks: dots vary around the fitted mean. This is a teaching simulation treating the fitted line and RMSE as the generating model.
| Interval | Target | Units |
|---|---|---|
| Coefficient CI | A population slope or intercept | Coefficient units |
| Fitted mean CI | Average revenue at given \(x_0\) | Revenue units |
| Prediction interval | One new week’s revenue at \(x_0\) | Revenue units |
All three involve estimation uncertainty. Only the prediction interval adds the new observation’s individual noise.
Intervals rely on the regression model and sampling assumptions. They do not repair a biased sample, a curved mean fitted as a straight line, dependent errors, or a changed business relationship.
Extrapolation: predicting outside the observed input range can fail even when the displayed formula produces a narrow-looking interval.
More data can reduce estimation uncertainty; it does not eliminate individual noise.
Return to Sections 2, 4 and 5 to connect these ideas with the Python results.
Prof. Xuhu Wan · HKUST ISOM · Introduction to Business Analytics