Tool/Research Method/No. 0221
Counterfactual Reasoning
Counterfactual reasoning asks what would have happened under a different action or condition. In causal research, it compares observed outcomes with unseen alternatives. Donald Rubin, David Lewis and Judea Pearl developed ways to formalize such comparisons and the assumptions they need.
Also called Counterfactual Analysis
- Evidence
- Well established
- Read
- 6 min
- Links
- 14 connections
01You've seen this when…
- in life
You leave home ten minutes late and miss your train. You check the departure time and work backward to whether leaving on time would have changed the outcome.
- at work
A team launches an ad campaign just before the holiday rush. Revenue climbs, and the campaign report credits the ads with the entire increase.
- out in the world
A city installs speed cameras and reports fewer crashes the following year. A council member asks what happens on similar roads without cameras.
02The idea
Every claim that something caused an outcome carries an alternative inside it. Saying an email generated a sale implies that, without the email, the sale would have been less likely. Counterfactual reasoning makes that alternative explicit.
The difficulty is that only one version happens. A customer receives the email or goes without it. Researchers call the outcomes under these different conditions potential outcomes. The difference between them defines the causal effect, but observing both for the same customer at the same moment is impossible.
So the practical task is to build a credible comparison. A randomly selected group that receives no email can show how similar customers behave without it. Historical records, comparable cases and causal models can also help, though each requires assumptions.
The same discipline applies to everyday explanations. When blaming a late departure for a missed train, specify the earlier departure and check whether it would actually have allowed enough time. The alternative needs a plausible path from the changed action to the changed result.
03How to use it
- Specify the outcome and the time window. Decide exactly what needs explaining: purchases during the next seven days, arrival before an interview, or crashes during the following year. A vague outcome lets the comparison drift.
- Name the alternative precisely. Compare sending the email with sending no email, or dispatching a delivery at 8 a.m. with dispatching it at 10 a.m. Different alternatives answer different questions. Comparing a campaign with doing nothing gives a different result from comparing it with another campaign.
- Separate background conditions from consequences. Keep relevant background facts consistent across the two versions, while allowing the effects of the changed action to unfold. For a delivery, retain the actual weather and road restrictions, change the dispatch time, and work through the resulting arrival time. A causal diagram can help distinguish these relationships.
- Find evidence for the missing outcome. A randomized experiment creates comparison groups that are similar on average. When randomization is unavailable, look for a credible natural experiment, comparable cases or repeated observations. Describe why the comparison would behave like the affected group without the intervention.
- Expose the assumptions. List what could break that resemblance: different customers, a seasonal surge, another policy change or people sharing the intervention with the comparison group. These are possible sources of confounding and other errors. Check which assumptions the available evidence can test.
- Carry uncertainty into the decision. Report a range when the data support one. For an imagined personal alternative, identify the links supported by records and those supplied by judgment. Then ask whether the decision changes across plausible versions. Several favorable assumptions stacked together can make an alternative look far more certain than it is.
04A worked example
Consider an illustrative online store testing a promotional email. Before the test begins, it randomly assigns 2,000 eligible customers to two equal groups. One group receives the email; the other receives no promotional email. Both groups face the same prices and shopping window.
During the following week, 60 customers in the email group place an order, compared with 40 in the comparison group.
What it looks like The campaign dashboard displays 60 purchasing customers associated with the email. Crediting all 60 to the campaign makes the email appear responsible for every purchase in that group.
What’s actually going on The comparison group provides an estimate of purchasing without the email. Purchase rates are 6% and 4%, giving an estimated increase of 2 percentage points. Applied to 1,000 recipients, that is about 20 additional purchasing customers. The estimate has sampling uncertainty, and it says nothing certain about which particular customers were persuaded. Profit also depends on order values and campaign costs.
What made it work Random assignment gave the store a credible stand-in for the missing outcome. Choosing the groups before sending the email also avoided comparing enthusiastic email openers with customers who ignored it. Measuring everyone assigned to each group preserved the original comparison. The estimate still depends on comparable measurement and limited spillover between groups.
05When to reach for it
06When it misleads
- An appealing story can outrun the evidence. Imagining that accepting another job would have produced a happier life requires assumptions about colleagues, workload, health and relationships. Label those assumptions. Scenario planning can explore several plausible futures without treating one as the hidden truth.
- Before-and-after comparisons can absorb unrelated changes. Sales may rise with the season, and crashes may fall after an unusually bad year through regression to the mean. An earlier period provides a useful baseline only when it credibly represents the later period without the intervention.
- Similar-looking groups can differ in decisive ways. People who volunteer for training may already be more motivated. Matching their age and job title can leave that difference intact. Selection bias remains possible even when a comparison table looks balanced.
- Changing one condition can change the surrounding system. Customers share promotions, drivers divert onto other roads, and competitors respond to prices. A small experiment can estimate an effect under its own conditions while leaving a large rollout uncertain. State whose outcomes the estimate covers and how the intervention might affect everyone else.
07Roots
In 1923, Jerzy Neyman faced a practical problem in agricultural experiments: each plot could grow only one crop variety at a time. To compare varieties, he represented the yield each plot would produce under each treatment. The harvest revealed one potential outcome; the others remained hidden.
Donald Rubin’s 1974 paper developed this approach for both randomized and nonrandomized studies. It gave researchers a precise way to describe the effect they wanted to estimate before deciding how to estimate it. Paul Holland later called the impossibility of observing both outcomes for one unit the fundamental problem of causal inference.
Philosophy approached the problem through a different door. In his 1973 essay on causation, David Lewis examined alternative possible worlds: worlds close to the actual one in which an event failed to occur. This helped sharpen what people mean when they say an outcome depended on a cause, while raising difficult questions about which alternative worlds count as closest.
Judea Pearl brought counterfactual questions into formal causal models, synthesized in his 2000 book Causality. These models connect facts about what happened with assumptions about how variables respond to interventions. Together, the statistical and philosophical traditions turned an ordinary habit of imagining alternatives into methods whose assumptions can be stated and challenged.
08How solid is this?
Counterfactual comparisons underpin established methods of causal inference. Randomization supports estimates of average effects; observational comparisons and claims about individual outcomes require additional assumptions. Imagined alternatives alone cannot establish a causal effect.
09Connections
- Helps counter Confounding, Correlation-Causation Fallacy, Outcome Bias, Post Hoc Fallacy, Fundamental Attribution Error, Hindsight Bias, Illusion of Control, Illusion of Explanatory Depth, Omission Bias, Self-Serving Bias
- Part of Scenario Planning
- IncludesPotential Outcomes
- See alsoRandomized Experiment, Natural Experiment
+ 4 more in the list
10Origin and sources
A longstanding philosophical idea. Jerzy Neyman described potential outcomes in agricultural experiments in 1923; David Lewis formalized counterfactual causation in 1973, Donald Rubin developed the statistical framework in 1974, and Judea Pearl synthesized structural causal methods in 2000.
- [1]Rubin, D. B. (1974). Estimating causal effects of treatments in randomized and nonrandomized studies. Journal of Educational Psychology, 66(5), 688–701.
- [2]Lewis, D. (1973). Causation. The Journal of Philosophy, 70(17), 556–567.
- [3]Holland, P. W. (1986). Statistics and Causal Inference. Journal of the American Statistical Association, 81(396), 945–960.
- [4]Pearl, J. (2000). Causality: Models, Reasoning, and Inference. Cambridge University Press.
Suggest an edit· Updated 2026-10-02