A/B tests and p-values
Run a fair experiment, measure the lift, and judge whether a difference could be chance - without misreading p-values.
- Explain why randomized experiments support causal claims
- Compute conversion rates and lift
- Run a permutation test and interpret its p-value correctly
An A/B test randomly splits users: group A sees the current page, group B a new design. Because chance alone decides who gets what, the groups are alike in every other way on average - so a real difference in outcomes can be credited to the change. Randomization is what turns correlation into causation.
But even two groups seeing the same page differ a little by chance. The question is: is the observed difference bigger than chance would usually produce?
Permutation tests
A permutation test answers that directly. If the change had no effect, the group labels are arbitrary - so shuffle them thousands of times and see how often a shuffled difference is at least as extreme as the real one. That fraction is the p-value: how surprising the data would be if there were no real effect.
The American Statistical Association’s statement on p-values (2016) warns against common misreadings:
- A p-value is not the probability that the hypothesis is true.
- It does not measure the size or importance of an effect.
- Decisions shouldn’t rest on whether p crosses 0.05 alone; report the effect size and its interval too.
Try it
Read the result correctly
Is each statement about an A/B test a correct reading or a misreading?
“p = 0.03 means there’s a 3% chance the new design doesn’t work.”
“p = 0.03: if the design had no effect, a difference this large would show up about 3% of the time.”
“p = 0.0001, so the improvement is huge.”
“p = 0.2, so the design definitely has no effect.”
“Report the lift (+1.5 points, 95% CI 0.4 to 2.6) along with the p-value.”
“Check the result every hour and stop as soon as p < 0.05.”
control, variant = 120 / 2400, 156 / 2400
print(f"{control:.2%} -> {variant:.2%}")
print(f"relative lift: {(variant - control) / control:+.0%}")5.00% -> 6.50% relative lift: +30%
Key takeaways
Randomization makes A/B tests support causal claims.
A permutation test’s p-value: how often chance alone gives a difference this extreme.
Report effect sizes and intervals; don’t treat p < 0.05 as a verdict, and don’t peek.
Lesson quiz
6 questions · pass with 5 correct · up to 50 XP
Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.
Practice: write Python
Write Python in the editor and run it against sample inputs. Python runs locally in your browser using a WebAssembly runtime.
Conversion rates and lift
Two lines: A conversions visitors and B conversions visitors. Print each rate as A: 5.00%, then lift: +P points (+R%) - absolute lift in percentage points (2 decimals) and relative lift (1 decimal), both with a sign.
- A better design
- A worse design
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Run a permutation test
Line 1 holds group A’s values, line 2 group B’s (space-separated). The observed difference is mean(B) − mean(A).
Call random.seed(0), pool the values, and 5,000 times random.shuffle the pool, treating the first len(A) values as A and the rest as B. Count shuffles whose |difference| is at least the |observed| difference. Print observed difference: D (2 decimals) and p-value: P (3 decimals).
- A clear difference
- Probably chance
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Questions about this lesson
Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.
Loading posts…