Samples, uncertainty and confidence intervals
See how estimates vary from sample to sample, and put honest error bars on them.
- Explain sampling variability and the standard error
- See the central limit theorem in action
- Build a confidence interval, including with the bootstrap
We rarely measure everyone. We take a sample and use it to estimate something about the population - the average delivery time, the share of satisfied customers. A different sample would give a slightly different answer. That wobble is sampling variability, and a good analysis reports it.
The spread of a sample mean across repeated samples is the standard error: SE = σ / √n. Quadrupling the sample size halves the uncertainty.
Try it
Draw samples from a skewed population
The population of delivery times is skewed: most are short, a few very long. Draw 1,000 samples of size 2, then switch to 20 and 50. Watch the sample means pile up into a narrow bell shape around the true mean - the central limit theorem.
Confidence intervals
A 95% confidence interval is a range built by a method that captures the true value in 95% of repeated samples. For a mean, a common approximation is mean ± 1.96 × SE.
When formulas get awkward (medians, ratios, small skewed samples), use the bootstrap (Efron, 1979): resample your own data with replacement thousands of times, compute the statistic each time, and take the middle 95% of the results. It needs nothing but a loop.
1import random
2random.seed(3)
3data = [12, 15, 18, 20, 21]
4print(random.choices(data, k=5))[15, 18, 15, 20, 20]
Key takeaways
Estimates vary from sample to sample; the standard error measures how much.
Sample means become bell-shaped as n grows, even from skewed populations.
Report intervals, not just point estimates - the bootstrap works for almost any statistic.
Lesson quiz
5 questions · pass with 4 correct · up to 50 XP
Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.
Practice: write Python
Write Python in the editor and run it against sample inputs. Python runs locally in your browser using a WebAssembly runtime.
Mean, standard error and a 95% interval
Read a JSON array of numbers. Print mean=M sd=S se=E (mean 1 decimal, sd and se 2 decimals, using statistics.stdev), then 95% CI: [LOW, HIGH] = mean ± 1.96 × se (1 decimal).
- Ten commutes
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
A bootstrap confidence interval
Read a JSON array of numbers. Call random.seed(1), then 2,000 times draw random.choices(values, k=len(values)) and record its statistics.mean. Sort the 2,000 means and print bootstrap 95% CI: [LOW, HIGH] using the means at positions 50 and 1949 (1 decimal).
- Ten commutes
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Questions about this lesson
Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.
Loading posts…