Pattern/Probability and Statistics/No. 0936

Simpson’s Paradox

Simpson’s paradox is a statistical pattern in which a trend within separate groups disappears or reverses when the groups are combined. Also called the Yule–Simpson paradox, it arises from unequal group weights; pooled and subgroup rates alone do not establish causes.

Also called Yule–Simpson Paradox

a pattern: watch for it

01You've seen this when…

  1. in life

    Your pace improves on both flat and hilly runs. Your monthly average still gets slower because you now run more miles on hills.

  2. at work

    More visitors buy on both mobile and desktop than last month. Overall conversion drops as traffic shifts toward mobile, where fewer visitors buy.

  3. out in the world

    One hospital has better survival rates for both low-risk and high-risk patients. A league table ranks it lower because it treats far more high-risk cases.

02The idea

Two numbers can point in opposite directions without either being calculated incorrectly. A treatment can have a higher success rate for both mild and severe cases, yet a lower success rate across all cases. A business can improve in every customer segment while its overall results deteriorate.

This is Simpson’s paradox: an association within separate groups disappears or reverses when those groups are combined.

The missing piece is the mix. An overall rate is a weighted average of the group rates: each group contributes according to how many observations it contains. If one treatment handles mostly difficult cases and another handles mostly easy ones, their overall rates compare different workloads as well as different treatments.

Both views can describe the data correctly. They answer different questions. The pooled rate describes what happened across the actual mix; the subgroup rates describe what happened within each category. Neither automatically tells you what caused the difference.

03Why it happens

  • Different starting points and group proportions shape the overall result. Large kidney stones are harder to treat than small ones. Mobile visitors may buy less often than desktop visitors. This heterogeneity gives the mix room to affect the overall result. The options contain different proportions of those groups. A treatment with better results in each category can still look worse overall if most of its patients belong to the harder category. The other treatment gets more weight from easier cases.
  • Pooling hides those proportions. A headline percentage compresses the subgroup rates and their denominators into one number. You lose sight of whether an option performs differently or simply receives a different population.

If both options had exactly the same subgroup proportions, an option that did better in every subgroup could not do worse overall. The unequal weights make the reversal possible.

In causal comparisons, the grouping variable may be a confounder: something that influences both which option people receive and how well they do. Simpson reversals can also arise simply from changing population mixes over time.

04A worked example

A 1986 study compared kidney-stone treatments, including open surgery and percutaneous nephrolithotomy, a procedure that removes stones through a small opening. The published counts show the reversal clearly:

  • For small stones, open surgery has the higher success rate. It succeeds in 81 of 87 cases, about 93%. Percutaneous treatment succeeds in 234 of 270, about 87%.
  • For large stones, open surgery also has the higher success rate. It succeeds in 192 of 263 cases, about 73%. Percutaneous treatment succeeds in 55 of 80, about 69%.

What it looks like Combine the categories and percutaneous treatment wins. Open surgery succeeds in 273 of 350 cases, 78%; percutaneous treatment succeeds in 289 of 350, about 83%.

What’s actually going on About 75% of the open-surgery group has large stones, compared with only 23% of the percutaneous group. Open surgery carries much more weight from the harder cases. Its higher success rate within each size category isn’t enough to overcome that imbalance in the pooled comparison.

What would have helped Publishing the subgroup counts alongside the totals, then comparing both treatments using the same stone-size mix. Giving small and large stones equal weight yields roughly 83% for open surgery and 78% for percutaneous treatment. That’s a descriptive adjustment, not proof that open surgery causes better outcomes. The treatments weren’t randomly assigned, and differences beyond stone size could still matter.

05How to spot it

Treat these signs as prompts to investigate. Establishing a paradox requires demonstrating a reversal in the data.

06What to do about it

  • Define the question before choosing the statistic. Describing actual outcomes, comparing performance and estimating a causal effect are different tasks. Write down which one you need.
  • Recover the denominators. Ask for the successes and total cases behind the percentages in each relevant group. Check which groups dominate each option.
  • Inspect meaningful breakdowns. Look at factors that plausibly affect both assignment and outcomes, such as severity before treatment. Don’t search hundreds of arbitrary slices until a preferred story appears.
  • Compare a common mix. Weight each option’s subgroup rates using the same proportions, chosen to represent the population you care about. Show those weights so someone else can reproduce the comparison.
  • Explain why you’re adjusting. A causal diagram can clarify whether the grouping variable belongs in the comparison. Adjustment needs a reason independent of any convenient reversal.
  • Keep both views visible. Present the aggregate together with a breakdown that includes group sizes. A single headline number conceals the information needed to interpret it.

07When it isn’t a reason to split everything

The usefulness of a subgroup result depends on the question you’re asking. For planning next month’s workload, the actual mix of easy and difficult cases may be exactly what matters. Removing that mix would answer a different question.

Adjustment can also mislead. If a training program changes which jobs people obtain, splitting earnings by job type may hide part of the program’s benefit. Grouping by something affected by the intervention differs from grouping by a preexisting difference. Conditioning on a shared consequence can even create an association through collider bias.

Simpson’s paradox is also distinct from the ecological fallacy. That fallacy means drawing conclusions about individuals from group-level data. Simpson’s paradox is a reversal between pooled and subgroup associations. They can occur together, but neither requires the other.

Use totals with a clear understanding of what they combine and whether that combination answers your question.

08Roots

In 1951, Edward H. Simpson examined how to interpret relationships in contingency tables: counts arranged by categories. With three yes-or-no characteristics, there are only eight cells, yet combining categories can change the apparent relationship. His paper made clear how much interpretation could depend on whether a third characteristic remained visible.

Simpson wasn’t the first to notice the trouble. Karl Pearson and colleagues had described related pooling effects in 1899, and George Udny Yule discussed association between categorical attributes in 1903. Statisticians were learning that a combined population could behave differently from its constituent groups, even without an arithmetic mistake.

Colin Blyth introduced the name Simpson’s paradox in a 1972 paper connecting it with principles of decision-making. The label helped the pattern travel beyond technical statistics. Comparisons of medical treatments and other institutional outcomes made it concrete: the same records could support opposite rankings. Later work on causal inference sharpened the practical lesson. The table alone doesn’t tell you whether combining or separating the groups answers the question you meant to ask.

09How solid is this?

ContestedMixedUsefulEstablished

The reversal is a mathematical consequence of unequal weighting, with documented examples in observed data. Determining which option works better requires further evidence about causes beyond the pooled or subgroup rates.

10Connections

confused withcountered bycountered bypart ofpart ofSimpson’sParadoxNot written yetEcologicalFallacyNot written yetDirectedAcyclic GraphNot written yetRandomizedExperimentConfoundingHeterogeneityCorrelationSelection BiasConditionalProbabilityNot written yetCollider Bias

11Origin and sources

Earlier descriptions by Karl Pearson and colleagues (1899) and George Udny Yule (1903). Edward H. Simpson analyzed the pattern in 1951; Colin R. Blyth introduced the name Simpson’s paradox in 1972.

  1. [1]Simpson, E. H. (1951). The Interpretation of Interaction in Contingency Tables. Journal of the Royal Statistical Society. Series B (Methodological), 13(2), 238–241.
  2. [2]Yule, G. U. (1903). Notes on the Theory of Association of Attributes in Statistics. Biometrika, 2(2), 121–134.
  3. [3]Blyth, C. R. (1972). On Simpson's Paradox and the Sure-Thing Principle. Journal of the American Statistical Association, 67(338), 364–366.
  4. [4]Charig, C. R., Webb, D. R., Payne, S. R., & Wickham, J. E. (1986). Comparison of treatment of renal calculi by open surgery, percutaneous nephrolithotomy, and extracorporeal shockwave lithotripsy. British Medical Journal (Clinical Research Edition), 292(6524), 879–882.
  5. [5]Pearl, J. (2014). Understanding Simpson's Paradox. The American Statistician, 68(1), 8–13.

Suggest an edit· Updated 2026-10-02