Chapter 3 · Appendix A

Statistical Foundations

Prof. Xuhu Wan

Appendix A — Statistical Foundations

Optional theory background: variables, samples, normal distributions, estimation and testing. High-school algebra is enough.

A.1 — Why We Need Statistics

A delivery company wants to know its average delivery time. Measuring every delivery may be impossible. We use a smaller set of deliveries to learn about the whole group.

Our route: data → sample → estimate → uncertainty → interval → test. No calculus or programming is needed.

A.2 — Variables and Data

A variable is a characteristic that can take different values. Data are the values we actually record.

Delivery Time (minutes) Area Late?
001 26 East No
002 34 West Yes
003 30 East No

Time, Area and Late? are variables. The recorded entries are data. Delivery is an identifier.

A.3 — Numerical and Categorical Variables

Type Example Meaningful calculation
Numerical Delivery time: 26, 34, 30 Mean time
Categorical Area: East, West Number in each area

A category coded 1 or 2 is still a label. Averaging area codes does not produce an average location.

If Late? is coded 1 for Yes and 0 for No, what does its mean measure?

The fraction of deliveries that are late. This works because the coding was explicitly defined as a 0/1 indicator.

A.4 — Population: The Whole Group

A population is the complete group we want to understand. Here it is all deliveries made by this company during the target month.

Define the group and time period before calculating. Next month may have a different population.

Population: all deliveriesSample: 6 selected deliveries→Unknown mean μCalculated mean x̄

A.5 — Sample: The Observations We Use

A sample is the set of observations selected from the population. Its size is \(n\).

If we randomly choose 36 deliveries from the target month, \(n=36\). Their average helps us estimate the population average. A sample is one set of observations, not a probability curve.

A.6 — A Large Sample Can Still Be Biased

Suppose we record only deliveries near the warehouse. Their average may be too low for all deliveries. More nearby deliveries do not fix this selection problem.

A simple random sample gives every delivery an equal chance of selection. The formulas below assume independent observations from the same population; a small sampling fraction makes sampling without replacement approximately independent.

A.7 — Population Parameters and Sample Statistics

Quantity Population: usually unknown Sample: calculated
Mean \(\mu\) (mu) \(\bar{x}\) (x-bar)
Standard deviation \(\sigma\) (sigma) \(s_x\)

A parameter describes the population. A statistic is calculated from a sample. Different samples can give different statistics.

A.8 — The Sample Mean

\[ \bar{x}=\frac{x_1+x_2+\cdots+x_n}{n} \]

  • \(x_i\): the recorded value for observation \(i\); \(i\) is just a row number.
  • \(n\): sample size.
  • \(\bar{x}\): add the values and divide by their count.

For 26, 30 and 34 minutes: \(\bar{x}=90/3=30\) minutes.

A.9 — Standard Deviation Measures Spread

Two samples can have the same mean and different spreads. A deviation is \(x_i-\bar{x}\). Squaring prevents positive and negative deviations cancelling.

Data: 26, 30, 34 minutes — mean 30263034−4+4Squared deviations: 16 + 0 + 16 = 32

A.10 — Sample Variance and Standard Deviation

\[ s_x^2=\frac{\sum_{i=1}^{n}(x_i-\bar{x})^2}{n-1},\qquad s_x=\sqrt{s_x^2} \]

The symbol \(\sum\) means “add these terms”. Variance has squared units; SD returns to minutes.

For 26, 30, 34: \(s_x^2=32/2=16\) minutes² and \(s_x=4\) minutes. We use \(n-1\) because the sample was also used to estimate its mean.

A.11 — A Random Variable

Before a delivery is selected, let \(X\) be its delivery time. We do not know its value yet. A random variable assigns a number to an uncertain outcome.

After observing it, we might record \(x=34\). Uppercase \(X\) describes the uncertain quantity; lowercase \(x\) is one realised value.

A.12 — A Probability Distribution

A probability distribution describes which values a random variable can take and how likely they are.

For a continuous quantity, probability is area under a curve over an interval. The entire area is 1, or 100%. Curve height is density, not the probability of exactly one value.

Probability is shaded area-202
For a continuous delivery-time model, what is the probability of exactly 30.000… minutes?

Zero. A single point has no area. A range such as 29.5 to 30.5 minutes can have positive probability.

A.13 — A Normal Random Variable

\[ X\sim N(\mu,\sigma^2) \]

This means \(X\) follows a normal distribution: a symmetric bell-shaped curve.

  • \(\mu\) is its centre (mean).
  • \(\sigma>0\) is its standard deviation (spread).
  • The second entry, \(\sigma^2\), is variance.

We use \(N(30,6^2)\) as an imagined delivery-time model. Real delivery times need not be normal.

A.14 — Interactive Normal Curve

Move the mean: the bell shifts. Increase SD: the bell spreads and becomes lower. The shaded area within one SD of the mean stays about 68.27%.

Predict the change before moving a slider.

A.15 — Standardise with a Z-Score

\[ Z=\frac{X-\mu}{\sigma}\sim N(0,1) \]

A z-score says how many population SDs a value lies above or below the mean. With \(\mu=30\) and \(\sigma=6\):

  • \(x=36\) gives \(z=1\): one SD above the mean.
  • \(x=18\) gives \(z=-2\): two SDs below the mean.

About 95% of a normal distribution lies between \(z=-1.96\) and \(z=1.96\).

A.16 — Different Samples Give Different Means

Imagine choosing three deliveries, recording their mean, then starting again with a new sample. The population stays the same, but the selected values change.

Same population → different samples → different means28, 30, 32↓x̄ = 3025, 27, 29↓x̄ = 2731, 33, 35↓x̄ = 33

A.17 — A Sampling Distribution

Take a sample of size \(n\) → calculate one sample mean → repeat with fresh samples.

The distribution of those means is the sampling distribution of \(\bar{X}\). Each value represents a whole sample, rather than one delivery.

A.18 — Animate Repeated Samples

Click Draw sample or Play / pause. Dots show the current deliveries; the histogram collects one mean per sample. Increase \(n\) and restart: the distribution of means becomes narrower.

A.19 — Standard Error of the Mean

\[ SE(\bar{X})=\frac{\sigma}{\sqrt{n}} \]

Standard error (SE) is the SD of a statistic across repeated samples. It measures sampling uncertainty in the estimate.

If \(\sigma=6\): \(n=9\) gives SE=2 minutes; \(n=36\) gives SE=1 minute. Quadrupling the sample size halves the SE.

A.20 — SD and SE Describe Different Things

Quantity Variation of what? In our imagined population
Population SD \(\sigma\) Individual delivery times 6 minutes
SE of \(\bar{X}\) Means across samples of 36 1 minute

A precise estimated average does not mean individual deliveries are all alike.

Individuals: SD=6; sample means (n=36): SE=1IndividualsSample means618304254
Does sampling more deliveries make the population SD shrink?

No. It reduces uncertainty in the sample mean. It does not make individual delivery times less variable.

A.21 — When Is the Mean Approximately Normal?

If the population is normal, the mean of independent observations is normal for any \(n\).

For a non-normal population with finite variance, the Central Limit Theorem often makes the mean approximately normal when \(n\) is sufficiently large. Strong skewness or outliers may require much larger samples. There is no universal “\(n=30\) fixes everything” rule.

A.22 — A Confidence Interval for the Mean

One estimate is not enough: add a margin of error on each side. When population SD \(\sigma\) is known and the sample mean is normal or approximately normal:

\[ \text{95% CI for }\mu:\quad \bar{x}\ \pm\ 1.96\frac{\sigma}{\sqrt{n}} \]

“\(\pm\)” means calculate a lower and an upper endpoint.

A.23 — A Confidence Interval in Minutes

For \(n=36\), \(\bar{x}=32\) and known \(\sigma=6\):

\[ SE=6/\sqrt{36}=1,\qquad CI=32\pm1.96=[30.04,33.96] \]

The target is population mean delivery time \(\mu\), not the time of one new delivery.

Estimate ± margin of error30.04x̄ = 3233.9695% confidence interval for population mean μ

A.24 — Unknown Population SD: Use Student’s t

Usually \(\sigma\) is unknown. Estimate it by the sample SD \(s_x\). This adds uncertainty, so use a t multiplier:

\[ \bar{x}\ \pm\ t^*\frac{s_x}{\sqrt{n}},\qquad df=n-1 \]

\(t^*\) is a positive cutoff determined by confidence level and degrees of freedom \(df\). For \(n=36\), a 95% t interval uses \(t^*\approx2.030\) instead of 1.96. Exact for independent normal observations; approximate under suitable large-sample conditions.

A.25 — Animate 95% Confidence

Repeat the sampling-and-interval procedure. The vertical line is the fixed population mean; each horizontal interval comes from a new sample.

About 95% of intervals cover the true mean over many repetitions—not necessarily exactly 95 out of every 100.

A.26 — What 95% Confidence Means

The method captures the true parameter in 95% of repeated samples under its assumptions. After a particular interval is calculated, the fixed parameter is either inside it or outside it.

A 95% CI does not contain 95% of individual deliveries and does not assign a 95% probability to the fixed parameter being in this realised interval.

A.27 — Hypothesis Testing

Question: is population mean delivery time different from 30 minutes?

\[ H_0:\mu=30\qquad H_1:\mu\ne30 \]

  • Null hypothesis \(H_0\): the reference claim we test.
  • Alternative \(H_1\): the difference we are looking for.

Assume the null and the sampling model, then ask how unusual our sample result would be.

A.28 — The Test Statistic

\[ z_{\mathrm{obs}}=\frac{\bar{x}-\mu_0}{\sigma/\sqrt{n}}=\frac{32-30}{1}=2 \]

  • \(\mu_0=30\): the mean specified by the null.
  • \(z_{\mathrm{obs}}\): observed distance from the null, measured in SEs.

The sample mean is two SEs above the null mean. We compare it with a standard normal curve under \(H_0\).

A.29 — What a P-Value Means

For a two-sided test, count outcomes at least as far from the null in either direction. For \(z_{\mathrm{obs}}=2\):

\[ p=P(Z\le-2)+P(Z\ge2)\approx0.0455 \]

If \(H_0\) and the model are correct, about 4.55% of repeated samples would give a test statistic at least this extreme. This is not the probability that \(H_0\) is true.

Under H₀: count both tails beyond ±2-2022.275%2.275%

A.30 — Significance and Decisions

Choose significance level \(\alpha_{\mathrm{sig}}\) before inspecting the result. At \(\alpha_{\mathrm{sig}}=0.05\), \(p=0.0455<0.05\), so reject \(H_0\) in this example.

  • \(p\le\alpha_{\mathrm{sig}}\): reject the null.
  • \(p>\alpha_{\mathrm{sig}}\): do not reject; this does not prove the null.

The significance level is a long-run false-rejection rate under \(H_0\), not the chance our particular conclusion is wrong. A small p-value alone does not show a large business effect.

A.31 — One-Sided and Two-Sided Questions

Question chosen in advance Null Alternative Count which tail?
Different from 30? \(\mu=30\) \(\mu\ne30\) Both
Greater than 30? \(\mu\le30\) \(\mu>30\) Right
Less than 30? \(\mu\ge30\) \(\mu<30\) Left

For one-sided tests, calculate the reference distribution at the boundary \(\mu=30\).

A.32 — Interactive P-Value Tails

Choose the research question, then move the observed z-statistic. The shaded probability is the p-value. At \(z=2\): two-sided \(p\approx0.0455\), greater-than \(p\approx0.0228\), less-than \(p\approx0.9772\).

A.33 — Choose the Direction Before Seeing the Data

Use a one-sided test only when the question was directional before observing the data. Do not switch sides or switch to one-sided testing to obtain a smaller p-value.

With a preselected “less than 30” alternative, a mean far above 30 is not evidence for that alternative. A one-sided p-value is not always half the two-sided one.

A.34 — A Confidence Interval and a Test

For the same two-sided procedure, a 95% CI and a test at 5% agree:

  • Null value outside the CI → reject.
  • Null value inside the CI → do not reject.

Our CI \([30.04,33.96]\) excludes 30 and our two-sided p-value is 0.0455. They are two ways to describe the same evidence.

Can we compare a 95% two-sided interval with any one-sided test and expect the same decision?

No. The confidence level, sidedness and statistical procedure must match.

A.35 — Statistical Foundations: Check Your Understanding

Term What it describes
Population / sample Target group / selected observations
Parameter / statistic Population quantity / calculated estimate
SD / SE Spread of observations / sampling spread of an estimate
Confidence interval Uncertainty about a parameter
P-value Tail probability under the null model

Next: the same ideas applied to a regression slope and predicted revenue.