Statistical Foundations
Optional theory background: variables, samples, normal distributions, estimation and testing. High-school algebra is enough.
A delivery company wants to know its average delivery time. Measuring every delivery may be impossible. We use a smaller set of deliveries to learn about the whole group.
Our route: data → sample → estimate → uncertainty → interval → test. No calculus or programming is needed.
A variable is a characteristic that can take different values. Data are the values we actually record.
| Delivery | Time (minutes) | Area | Late? |
|---|---|---|---|
| 001 | 26 | East | No |
| 002 | 34 | West | Yes |
| 003 | 30 | East | No |
Time, Area and Late? are variables. The recorded entries are data. Delivery is an identifier.
| Type | Example | Meaningful calculation |
|---|---|---|
| Numerical | Delivery time: 26, 34, 30 | Mean time |
| Categorical | Area: East, West | Number in each area |
A category coded 1 or 2 is still a label. Averaging area codes does not produce an average location.
The fraction of deliveries that are late. This works because the coding was explicitly defined as a 0/1 indicator.
A population is the complete group we want to understand. Here it is all deliveries made by this company during the target month.
Define the group and time period before calculating. Next month may have a different population.
A sample is the set of observations selected from the population. Its size is \(n\).
If we randomly choose 36 deliveries from the target month, \(n=36\). Their average helps us estimate the population average. A sample is one set of observations, not a probability curve.
Suppose we record only deliveries near the warehouse. Their average may be too low for all deliveries. More nearby deliveries do not fix this selection problem.
A simple random sample gives every delivery an equal chance of selection. The formulas below assume independent observations from the same population; a small sampling fraction makes sampling without replacement approximately independent.
| Quantity | Population: usually unknown | Sample: calculated |
|---|---|---|
| Mean | \(\mu\) (mu) | \(\bar{x}\) (x-bar) |
| Standard deviation | \(\sigma\) (sigma) | \(s_x\) |
A parameter describes the population. A statistic is calculated from a sample. Different samples can give different statistics.
\[ \bar{x}=\frac{x_1+x_2+\cdots+x_n}{n} \]
For 26, 30 and 34 minutes: \(\bar{x}=90/3=30\) minutes.
Two samples can have the same mean and different spreads. A deviation is \(x_i-\bar{x}\). Squaring prevents positive and negative deviations cancelling.
\[ s_x^2=\frac{\sum_{i=1}^{n}(x_i-\bar{x})^2}{n-1},\qquad s_x=\sqrt{s_x^2} \]
The symbol \(\sum\) means “add these terms”. Variance has squared units; SD returns to minutes.
For 26, 30, 34: \(s_x^2=32/2=16\) minutes² and \(s_x=4\) minutes. We use \(n-1\) because the sample was also used to estimate its mean.
Before a delivery is selected, let \(X\) be its delivery time. We do not know its value yet. A random variable assigns a number to an uncertain outcome.
After observing it, we might record \(x=34\). Uppercase \(X\) describes the uncertain quantity; lowercase \(x\) is one realised value.
A probability distribution describes which values a random variable can take and how likely they are.
For a continuous quantity, probability is area under a curve over an interval. The entire area is 1, or 100%. Curve height is density, not the probability of exactly one value.
Zero. A single point has no area. A range such as 29.5 to 30.5 minutes can have positive probability.
\[ X\sim N(\mu,\sigma^2) \]
This means \(X\) follows a normal distribution: a symmetric bell-shaped curve.
We use \(N(30,6^2)\) as an imagined delivery-time model. Real delivery times need not be normal.
Move the mean: the bell shifts. Increase SD: the bell spreads and becomes lower. The shaded area within one SD of the mean stays about 68.27%.
Predict the change before moving a slider.
\[ Z=\frac{X-\mu}{\sigma}\sim N(0,1) \]
A z-score says how many population SDs a value lies above or below the mean. With \(\mu=30\) and \(\sigma=6\):
About 95% of a normal distribution lies between \(z=-1.96\) and \(z=1.96\).
Imagine choosing three deliveries, recording their mean, then starting again with a new sample. The population stays the same, but the selected values change.
Take a sample of size \(n\) → calculate one sample mean → repeat with fresh samples.
The distribution of those means is the sampling distribution of \(\bar{X}\). Each value represents a whole sample, rather than one delivery.
Click Draw sample or Play / pause. Dots show the current deliveries; the histogram collects one mean per sample. Increase \(n\) and restart: the distribution of means becomes narrower.
\[ SE(\bar{X})=\frac{\sigma}{\sqrt{n}} \]
Standard error (SE) is the SD of a statistic across repeated samples. It measures sampling uncertainty in the estimate.
If \(\sigma=6\): \(n=9\) gives SE=2 minutes; \(n=36\) gives SE=1 minute. Quadrupling the sample size halves the SE.
| Quantity | Variation of what? | In our imagined population |
|---|---|---|
| Population SD \(\sigma\) | Individual delivery times | 6 minutes |
| SE of \(\bar{X}\) | Means across samples of 36 | 1 minute |
A precise estimated average does not mean individual deliveries are all alike.
No. It reduces uncertainty in the sample mean. It does not make individual delivery times less variable.
If the population is normal, the mean of independent observations is normal for any \(n\).
For a non-normal population with finite variance, the Central Limit Theorem often makes the mean approximately normal when \(n\) is sufficiently large. Strong skewness or outliers may require much larger samples. There is no universal “\(n=30\) fixes everything” rule.
One estimate is not enough: add a margin of error on each side. When population SD \(\sigma\) is known and the sample mean is normal or approximately normal:
\[ \text{95% CI for }\mu:\quad \bar{x}\ \pm\ 1.96\frac{\sigma}{\sqrt{n}} \]
“\(\pm\)” means calculate a lower and an upper endpoint.
For \(n=36\), \(\bar{x}=32\) and known \(\sigma=6\):
\[ SE=6/\sqrt{36}=1,\qquad CI=32\pm1.96=[30.04,33.96] \]
The target is population mean delivery time \(\mu\), not the time of one new delivery.
Usually \(\sigma\) is unknown. Estimate it by the sample SD \(s_x\). This adds uncertainty, so use a t multiplier:
\[ \bar{x}\ \pm\ t^*\frac{s_x}{\sqrt{n}},\qquad df=n-1 \]
\(t^*\) is a positive cutoff determined by confidence level and degrees of freedom \(df\). For \(n=36\), a 95% t interval uses \(t^*\approx2.030\) instead of 1.96. Exact for independent normal observations; approximate under suitable large-sample conditions.
Repeat the sampling-and-interval procedure. The vertical line is the fixed population mean; each horizontal interval comes from a new sample.
About 95% of intervals cover the true mean over many repetitions—not necessarily exactly 95 out of every 100.
The method captures the true parameter in 95% of repeated samples under its assumptions. After a particular interval is calculated, the fixed parameter is either inside it or outside it.
A 95% CI does not contain 95% of individual deliveries and does not assign a 95% probability to the fixed parameter being in this realised interval.
Question: is population mean delivery time different from 30 minutes?
\[ H_0:\mu=30\qquad H_1:\mu\ne30 \]
Assume the null and the sampling model, then ask how unusual our sample result would be.
\[ z_{\mathrm{obs}}=\frac{\bar{x}-\mu_0}{\sigma/\sqrt{n}}=\frac{32-30}{1}=2 \]
The sample mean is two SEs above the null mean. We compare it with a standard normal curve under \(H_0\).
For a two-sided test, count outcomes at least as far from the null in either direction. For \(z_{\mathrm{obs}}=2\):
\[ p=P(Z\le-2)+P(Z\ge2)\approx0.0455 \]
If \(H_0\) and the model are correct, about 4.55% of repeated samples would give a test statistic at least this extreme. This is not the probability that \(H_0\) is true.
Choose significance level \(\alpha_{\mathrm{sig}}\) before inspecting the result. At \(\alpha_{\mathrm{sig}}=0.05\), \(p=0.0455<0.05\), so reject \(H_0\) in this example.
The significance level is a long-run false-rejection rate under \(H_0\), not the chance our particular conclusion is wrong. A small p-value alone does not show a large business effect.
| Question chosen in advance | Null | Alternative | Count which tail? |
|---|---|---|---|
| Different from 30? | \(\mu=30\) | \(\mu\ne30\) | Both |
| Greater than 30? | \(\mu\le30\) | \(\mu>30\) | Right |
| Less than 30? | \(\mu\ge30\) | \(\mu<30\) | Left |
For one-sided tests, calculate the reference distribution at the boundary \(\mu=30\).
Choose the research question, then move the observed z-statistic. The shaded probability is the p-value. At \(z=2\): two-sided \(p\approx0.0455\), greater-than \(p\approx0.0228\), less-than \(p\approx0.9772\).
Use a one-sided test only when the question was directional before observing the data. Do not switch sides or switch to one-sided testing to obtain a smaller p-value.
With a preselected “less than 30” alternative, a mean far above 30 is not evidence for that alternative. A one-sided p-value is not always half the two-sided one.
For the same two-sided procedure, a 95% CI and a test at 5% agree:
Our CI \([30.04,33.96]\) excludes 30 and our two-sided p-value is 0.0455. They are two ways to describe the same evidence.
No. The confidence level, sidedness and statistical procedure must match.
| Term | What it describes |
|---|---|
| Population / sample | Target group / selected observations |
| Parameter / statistic | Population quantity / calculated estimate |
| SD / SE | Spread of observations / sampling spread of an estimate |
| Confidence interval | Uncertainty about a parameter |
| P-value | Tail probability under the null model |
Next: the same ideas applied to a regression slope and predicted revenue.
Prof. Xuhu Wan · HKUST ISOM · Introduction to Business Analytics