Tool/Design and Human Factors/No. 0001
A/B Testing
Your team prefers the new signup page, but you keep the old one live for half the visitors to see what happens.
Also called Split Testing · Bucket Testing
- Evidence
- Well established
- Read
- 6 min
- Links
- 14 connections
01You've seen this when…
- in life
You write a hobby newsletter. Half the subscribers randomly receive a question in the subject line; half receive your usual description. You compare link clicks to put your favorite subject line to the test.
- at work
A designer proposes removing two fields from signup. New visitors are randomly assigned to the old or shorter form, and the team measures completed signups.
- out in the world
A library wants fewer overdue books. It randomly assigns eligible cardholders to receive either the existing reminder or a revised message, then compares return rates.
02The idea
A redesign launches on Monday. Sales rise on Tuesday. The redesign might have helped. Sales could also rise on payday or during an advertising campaign, and a competitor’s outage might send buyers your way. A before-and-after comparison cannot separate those effects.
A/B testing keeps the alternatives running at the same time. A is usually the current version; B is the proposed change. Chance decides which eligible person or unit receives which version. You then compare a defined outcome across the assigned groups.
Random assignment makes the groups comparable on average, including on factors you never thought to measure. Groups can still differ in any particular test. That’s why you need an estimate of uncertainty to interpret the winning number.
This is a practical form of randomized experimentation. Its strength is estimating what a change causes, rather than what happens alongside it. It can show which version works better on your chosen outcome. Explaining why one version works better or settling decisions about what to build next and whether the outcome is worth pursuing usually requires additional evidence or judgment.
03How to use it
- Define the decision before the test. Specify who is eligible, what changes, and what result would justify adopting it. Control other differences between versions so you can test one feature; if you leave several changes in place, treat them as a bundle.
- Choose one main outcome and a few safeguards. For signup, measure the percentage of assigned users who complete signup. Treat clicks on the first button as an intermediate outcome. Also check for harm, such as more errors, complaints, or later cancellations.
- Randomize at the right level. Assign each account consistently if people return. If treatment affects a whole classroom or household, consider assigning that group together. Analyze results in a way that respects those grouped assignments.
- Plan the sample and stopping rule. Use the baseline rate and the smallest improvement worth acting on to estimate how much data you need. Statistical power describes your chance of detecting an effect of a given size. Include relevant weekly or seasonal cycles.
- Check that the experiment runs correctly. Confirm that assignment and exposure work as planned and that tracking records outcomes correctly. An unexplained imbalance in group sizes can signal a bug. Account explicitly for all assigned people in the analysis, including those whose outcomes are inconvenient or missing.
- Compare effects, uncertainty, and costs. Report the size of the difference with a confidence interval. These estimates give context to whether the result crosses a significance threshold. Follow the planned stopping rule unless safety requires intervention, or use a statistical method designed for repeated checking.
04A worked example
Ron Kohavi and Stefan Thomke describe a Microsoft employee’s 2012 proposal to change how Bing displayed advertising headlines. The team initially gave the idea low priority, but months later another engineer implemented the change and tested it in an experiment.
What it looks like A minor display change, unlikely to deserve much engineering attention. The team’s initial judgment gives little reason to expect a large payoff.
What’s actually going on The randomized comparison reveals an unexpectedly large effect. Kohavi and Thomke report a 12% increase in advertising revenue. The result is surprising enough that the team investigates whether it reflects an error before accepting it. Internal enthusiasm turns out to be a poor forecast of this change’s value.
What made it work The team implemented the change cheaply and used a concurrent control group, then checked an implausibly strong result. The experiment lets evidence overturn a low-priority label. It establishes a revenue effect. Claims of improved user satisfaction or a lasting benefit under every future condition require evidence beyond this result.
05When to reach for it
Combine it with usability testing. Watching five people struggle can reveal a broken interaction. An A/B test can estimate how widespread the resulting performance difference is.
06When it misleads
- You stop when the dashboard turns favorable. Repeatedly checking an ordinary fixed-sample test and stopping at the first promising result raises false-positive risk. This is one route into p-hacking.
- You search for a winner everywhere. Testing many metrics, variants, or audience segments creates a multiple comparisons problem. Separate planned conclusions from exploratory findings that need confirmation.
- You mistake uncertainty for equality. A small test with an inconclusive result may still allow a substantial benefit or harm. Read the interval to see which effects remain plausible; equivalence remains unresolved.
- The groups affect each other. In a marketplace, changing prices for some buyers can change availability for everyone. Indirect exposure to treatment changes conditions for the control group too. The assignment scheme must account for these spillovers.
- The measurement changes with the treatment. A new page might lose its tracking event or encourage a click without a completed purchase. With random assignment in place, a reliable comparison still requires a working, meaningful outcome measure.
- A short-term lift hides a longer-term cost. Novelty can fade. More notifications can produce immediate engagement and later annoyance. Choose a duration and follow-up suited to the decision.
- You generalize beyond the test. A result among new users this month has uncertain relevance to longtime customers next year. This is an external validity question; answering it requires evidence beyond random assignment.
07Roots
At Rothamsted, an agricultural research station in England, Ronald Fisher faced a problem that product teams would recognize: a better result could come from the intervention or from conditions around it. A wheat plot might yield more because of fertilizer, but also because its soil was different or the weather favored that season.
Fisher helped make random assignment and replicated comparisons into a coherent experimental method. Chance determined which plots received treatments, while the design allowed researchers to estimate how much variation could occur without a treatment effect. His 1935 book, The Design of Experiments, helped establish that framework. Building on existing ingredients of experimentation, he showed how design and statistical inference should work together.
The web made comparable experiments cheap to deliver and fast to measure. Software could assign visitors to versions, preserve those assignments, and record outcomes without arranging a new field trial each time. Online practitioners also encountered new problems: returning users, broken logging, overlapping experiments, and dashboards checked constantly. Kohavi and colleagues documented practical guidance in 2007, and later work developed it further. Today’s A/B test inherits the agricultural logic, but reliable online testing requires much more than a button that splits traffic.
08How solid is this?
Random assignment is a well-established method for estimating causal effects. Individual A/B tests can still mislead through low power, broken measurement, spillovers, selective analysis, or generalizing beyond the tested population and period.
09Connections
- Helps counter Correlation-Causation Fallacy, P-Hacking, Confirmation Bias, Risk Compensation
- Part ofRandomized Experiment, Build-Measure-Learn, Explore-Exploit Trade-Off, Safe-to-Fail Experiment
- See alsoUsability Testing, Multiple Comparisons Problem, Statistical Power, Confidence Interval, External Validity, Two-Way vs. One-Way Doors
+ 4 more in the list
10Origin and sources
Modern randomized experimental design was formalized by Ronald A. Fisher through agricultural research in the 1920s and 1930s. Online controlled-experiment practice adapted these principles to software and web products.
- [1]Fisher, R. A. (1935). The Design of Experiments. Oliver and Boyd.
- [2]Kohavi, R., Henne, R. M., & Sommerfield, D. (2007). Practical guide to controlled experiments on the web. Proceedings of the 13th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 959–967.
- [3]Kohavi, R., & Thomke, S. (2017). The surprising power of online experiments. Harvard Business Review, 95(5), 74–82.
- [4]Kohavi, R., Tang, D., & Xu, Y. (2020). Trustworthy Online Controlled Experiments: A Practical Guide to A/B Testing. Cambridge University Press.
Suggest an edit· Updated 2026-10-02