Chapter 3 · Appendix B

Linear Regression Theory

Prof. Xuhu Wan

Appendix B — Linear Regression Theory

Optional theory background: fitted lines, coefficient uncertainty, confidence intervals and prediction intervals. Read Appendix A first if these terms are new.

B.1 — Regression Uses the Same Statistical Ideas

A sample mean estimates one average. A regression estimates an average at each input value.

Our classroom example: advertising spending \(x\) and weekly revenue \(Y\), both in HKD thousands. We will fit a line, describe coefficient uncertainty, then distinguish average revenue from one new week.

B.2 — The Population Model

\[ Y=\beta_0+\beta_1x+\varepsilon \]

  • \(x\): advertising input.
  • \(Y\): uncertain revenue at that input.
  • \(\beta_0\): population intercept.
  • \(\beta_1\): population slope.
  • \(\varepsilon\) (epsilon): individual deviation from the average; \(E(\varepsilon\mid x)=0\).

Thus the population mean is \(E(Y\mid x)=\beta_0+\beta_1x\). “\(\mid x\)” means at the given input \(x\).

B.3 — One Input, Several Possible Outcomes

At the same advertising input, revenue can differ because of weather, customer demand and other unobserved influences.

The straight line describes average revenue, while noise describes the spread around it.

Same x can produce different y valuesLine = average; vertical spread = noise

B.4 — The Fitted Line

\[ \hat{y}(x)=\hat{\beta}_0+\hat{\beta}_1x \]

A hat means “estimated from our sample”.

  • \(\hat{\beta}_0,\hat{\beta}_1\): estimated intercept and slope.
  • \(\hat{y}(x)\): estimated average of \(Y\) at input \(x\).

For this sample, \(\hat y(x)\approx9.6429+2.1905x\). Displayed coefficients are rounded; calculations use full precision. An extra HKD1,000 of Ads is associated with HKD2,190.5 higher fitted revenue. Association alone does not establish causation.

B.5 — Observations and the Fitted Line

Each dot is one observed week. The line supplies a fitted value at that week’s input. The vertical gaps are residuals: observed minus fitted.

Advertising xRevenue y: observed dotsFitted mean line

B.6 — Residuals and Squared Error

\[ e_i=y_i-\hat y(x_i),\qquad SSE=\sum_{i=1}^{n}e_i^2 \]

  • \(e_i\): residual for observed week \(i\).
  • \(SSE\): sum of squared residuals; it is never negative.

A residual above the line is positive; below is negative. Squaring stops the signs cancelling. Fitting by ordinary least squares (OLS) means choosing the line with the smallest SSE.

B.7 — Animate the Least-Squares Fit

Move the intercept and slope. Watch the residual segments and SSE change. Click Animate OLS fit to move to the least-squares line.

The minimum is for the total squared error, not for each residual separately.

B.8 — Population Noise and Sample Residuals

\[ \varepsilon_i=y_i-(\beta_0+\beta_1x_i),\qquad e_i=y_i-(\hat\beta_0+\hat\beta_1x_i) \]

The true population line is unknown, so we cannot observe its noise values directly. Residuals use the estimated line. They help us estimate the noise spread and examine assumptions.

B.9 — Assumptions Behind the Formulas

Assumption Plain-language meaning
Linear mean The population average follows a straight line
Independent errors One observation’s error does not determine another’s
Constant variance Noise has the same SD \(\sigma\) at every input
Normal errors At a fixed input, errors follow a normal distribution

These give exact small-sample t tests and intervals. The predictor \(x\) itself need not be normal; \(x\) must vary. Normality concerns errors conditional on \(x\).

B.10 — RMSE Estimates the Noise SD

\[ s=\mathrm{RMSE}=\sqrt{\frac{SSE}{n-2}} \]

This course uses RMSE for the residual standard error. It estimates the unknown noise SD \(\sigma\), in revenue units. We use \(n-2\) because simple linear regression estimates an intercept and a slope.

In our eight-week sample: \(SSE=12.4762\), so \(s=\sqrt{12.4762/6}=1.4420\) HKD thousands.

B.11 — A Slope Also Has a Sampling Distribution

Take a new sample of weeks and refit: the fitted slope changes. The population slope stays fixed.

\(SE(\hat\beta_1)\) estimates the SD of slope estimates across repeated samples. It measures the precision of the estimated slope, not the spread of individual revenues.

Same population relationship, three possible fitted linesDifferent samples → different slope estimates

B.12 — Standard Error of the Slope

\[ S_{xx}=\sum_{i=1}^{n}(x_i-\bar x)^2,\qquad SE(\hat\beta_1)=\frac{s}{\sqrt{S_{xx}}} \]

  • \(S_{xx}\): total squared spread of the observed inputs.
  • \(s\): RMSE, the estimated noise SD.

Our sample has \(S_{xx}=42\), so \(SE(\hat\beta_1)=1.4420/\sqrt{42}=0.2225\). Across repeated comparable samples, estimated slopes typically vary around the true slope on a scale of about 0.2225 revenue units per Ads unit.

B.13 — Standard Error of the Intercept

\[ SE(\hat\beta_0)=s\sqrt{\frac1n+\frac{\bar x^2}{S_{xx}}} \]

The intercept is the estimated mean at \(x=0\). If the observed inputs are far from zero, it is harder to estimate reliably.

Here \(n=8\), \(\bar x=4.5\), \(S_{xx}=42\), giving \(SE(\hat\beta_0)=1.1236\). Do not give a business interpretation at zero if zero is outside the meaningful range.

B.14 — A Confidence Interval for a Coefficient

\[ \hat\beta_j\ \pm\ t^*SE(\hat\beta_j),\qquad df=n-2 \]

\(j=0\) means intercept; \(j=1\) means slope. For our sample, \(df=6\) and a 95% interval uses \(t^*=2.4469\).

Slope CI: \(2.1905\pm2.4469(0.2225)=[1.6460,2.7349]\).

This estimates the population slope: the association is roughly HKD1,646 to HKD2,735 per extra HKD1,000 in Ads under the model.

B.15 — A Two-Sided Test for a Coefficient

\[ H_0:\beta_j=0,\qquad H_1:\beta_j\ne0 \]

\[ t_{\mathrm{obs}}=\frac{\hat\beta_j-0}{SE(\hat\beta_j)},\qquad df=n-2 \]

For the slope: \(t=2.1905/0.2225\approx9.8446\), with two-sided \(p\approx0.0000633\). Under a zero population slope and our model assumptions, such a large absolute t-statistic is very rare. Reject at 5%; this alone does not prove causation.

B.16 — Interactive Coefficient Test and Interval

Keep \(df=6\) and adjust the estimated slope or its SE. Compare the two-sided p-value, 95% CI and the null value zero. A larger SE means less precise estimation.

These controls illustrate hypothetical regression outputs; they do not refit the eight observed weeks.

B.17 — A Coefficient CI and Its Test Agree

For a matching two-sided 5% t test and 95% t interval:

  • CI excludes zero → coefficient differs significantly from zero.
  • CI includes zero → insufficient evidence of a nonzero coefficient.

An interval shows both uncertainty and plausible effect sizes.

0Zero excludedZero includedCompare a two-sided 95% coefficient CI with zero

B.18 — Multiple Linear Regression

\[ Y=\beta_0+\beta_1x_1+\cdots+\beta_kx_k+\varepsilon \]

\(k\) is the number of predictors, excluding the intercept. Example: predict revenue using Ads and Price, so \(k=2\). Each slope describes a change in the average outcome holding the other predictors fixed.

The fitted equation replaces every coefficient by its sample estimate. No matrix algebra is needed to interpret it.

B.19 — RMSE and Degrees of Freedom in Multiple Linear Regression

\[ s=\mathrm{RMSE}=\sqrt{\frac{SSE}{n-k-1}},\qquad df=n-k-1 \]

We estimate \(k\) slopes plus one intercept. With \(n=20\) observations and \(k=2\) predictors, \(df=17\).

For coefficient \(j\): CI is \(\hat\beta_j\pm t^*SE(\hat\beta_j)\) and the test uses \((\hat\beta_j-0)/SE(\hat\beta_j)\) with the same df. Software calculates the SEs; the simple-linear-regression slope SE formula does not apply unchanged.

B.20 — Individual t Tests and the Overall F Test

Test Null Alternative
Individual t \(\beta_j=0\) \(\beta_j\ne0\)
Overall F \(\beta_1=\cdots=\beta_k=0\) At least one slope is nonzero

An individual t test asks about one predictor with the others included. The overall F test asks whether the predictor set helps explain the mean at all. A significant F test does not say that every coefficient is significant.

B.21 — A New Input and a Fitted Mean

\[ x_0=4,\qquad\hat y_0=\hat y(4)\approx18.4048 \]

The subscript 0 labels the prediction case; it does not mean \(x_0=0\). Here the input is Ads=HKD4,000.

\(\hat y_0\) estimates average revenue at this input: HKD18,404.8. It is also our point prediction for one new week, but that week has extra uncertainty.

B.22 — Standard Error of the Fitted Mean

\[ SE(\hat y_0)=s\sqrt{\frac1n+\frac{(x_0-\bar x)^2}{S_{xx}}} \]

This measures how much our estimated average at \(x_0\) varies across repeated training samples. It is smallest near the average observed input \(\bar x\).

At \(x_0=4\): \(SE(\hat y_0)=0.5218\) HKD thousands. This is smaller than \(s=1.4420\), which measures individual noise.

B.23 — Confidence Interval for Average Revenue

\[ \hat y_0\ \pm\ t^*SE(\hat y_0) \]

For whom? The population average of all comparable weeks with Ads fixed at HKD4,000.

\(18.4048\pm2.4469(0.5218)=[17.1279,19.6816]\) HKD thousands.

This interval measures uncertainty about the mean response, not the variability of one future week.

B.24 — One New Week Has Extra Uncertainty

A prediction for one future week has two sources of uncertainty:

  1. We estimated the population mean line using a sample.
  2. The future week has its own unpredictable noise.

A prediction interval includes both, so it is wider than a confidence interval for the average at the same input and confidence level.

At exactly the same input x₀Prediction interval: one new weekConfidence interval: average revenueBoth centred on ŷ₀; different targets

B.25 — Prediction Interval for One New Week

\[ SE_{\mathrm{pred}}=s\sqrt{1+\frac1n+\frac{(x_0-\bar x)^2}{S_{xx}}} \]

\[ \text{95% PI}:\quad \hat y_0\ \pm\ t^*SE_{\mathrm{pred}} \]

The extra 1 accounts for the new week’s noise. At \(x_0=4\), the PI is \([14.6524,22.1571]\) HKD thousands. It targets one independent new response at the specified input, under the model assumptions.

B.26 — Interactive Confidence and Prediction Intervals

Move \(x_0\) and change the confidence level. Compare the narrow mean CI with the wider individual PI. Both widen away from \(\bar x=4.5\).

Click Animate new weeks: dots vary around the fitted mean. This is a teaching simulation treating the fitted line and RMSE as the generating model.

B.27 — Three Different Interval Targets

Interval Target Units
Coefficient CI A population slope or intercept Coefficient units
Fitted mean CI Average revenue at given \(x_0\) Revenue units
Prediction interval One new week’s revenue at \(x_0\) Revenue units

All three involve estimation uncertainty. Only the prediction interval adds the new observation’s individual noise.

B.28 — What the Intervals Cannot Fix

Intervals rely on the regression model and sampling assumptions. They do not repair a biased sample, a curved mean fitted as a straight line, dependent errors, or a changed business relationship.

Extrapolation: predicting outside the observed input range can fail even when the displayed formula produces a narrow-looking interval.

More data can reduce estimation uncertainty; it does not eliminate individual noise.

B.29 — Regression Theory: Check Your Understanding

  1. The fitted line estimates the average of \(Y\) at an input.
  2. OLS chooses coefficients that minimise SSE.
  3. RMSE \(s\) estimates underlying noise SD.
  4. Coefficient SEs, CIs and t tests describe estimation uncertainty.
  5. A mean CI concerns an average; a PI concerns one new outcome.

Return to Sections 2, 4 and 5 to connect these ideas with the Python results.