Concept/Probability and Statistics/No. 0460

Heterogeneity

Heterogeneity is variation in outcomes, effects, or causes across people, places, or times. In statistics and causal inference, it describes how groups can differ even when an overall mean is the same. It can limit how well a result applies to a different group.

a concept: name it

01You've seen this when…

  1. in life

    You and a friend follow the same running plan. After six weeks, your pace improves while your friend keeps losing training days to knee pain.

  2. at work

    The support dashboard shows shorter waits. Enterprise customers now get an answer within minutes, while everyone else waits longer than before.

  3. out in the world

    A home-insulation program cuts energy use across the city. Savings are substantial in detached houses and barely detectable in apartments.

02The idea

An average folds different cases into one number. Heterogeneity is the variation among those cases: who benefits, how much, under which conditions, and through which process.

Two classes can both average 70 on a test. In one, nearly everyone scores between 65 and 75. In the other, scores cluster around 40 and 100. The same average describes very different teaching problems.

The idea reaches beyond the spread of outcomes. A tutoring program might improve scores more for beginners than for advanced students. A pricing change might raise sales in one region and reduce them in another. The effect itself varies.

That distinction matters. Students having different final scores establishes variation in outcomes. Establishing variation in tutoring’s effect requires estimating how much tutoring changes scores for different students.

Variance measures spread numerically. Heterogeneity is the broader idea that people, settings, or processes differ. An interaction effect is one way to represent it: the relationship between two variables changes with a third.

Observed differences also contain sampling noise. A small subgroup can look exceptional simply because its estimate is imprecise. Analysis tries to distinguish that uncertainty from underlying variation.

03Why it matters

Heterogeneity changes what evidence supports and what action makes sense.

  • Choose whom an intervention serves. An average benefit can coexist with harm to a subgroup. Knowing the pattern can support a different rollout, an alternative design, or a safeguard for the people affected.
  • Match evidence to the next setting. A successful trial among experienced users gives limited guidance about first-time users. External validity concerns how well findings travel. Differences in skills, infrastructure, or baseline risk can change the result.
  • Explain a shifting headline number. The average can move when the mix of people changes, even if each group’s outcome stays steady. This is one route to Simpson’s paradox, where a pooled relationship reverses the relationships within groups.
  • Budget for unequal needs. A standard service allowance may cover most customers while leaving a smaller group with much higher needs. Planning from the average alone can produce predictable shortages.

Before applying a result, name the population, outcome, and time period it describes. Then identify the differences between that population and the people affected by the next decision. This makes the question concrete enough to investigate.

04A worked example

Imagine a college testing a guided-practice app with 2,000 students. Half have reliable access to a laptop; half mainly use a shared phone. Within each group, the college randomly assigns 500 students to the app and 500 to the existing worksheets.

Among laptop users, 76% of the app group pass the course, compared with 60% of the worksheet group: a gain of 16 percentage points. Among shared-phone users, the figures are 52% and 60%: a drop of 8 percentage points.

What it looks like The app improves the pass rate from 60% to 64% overall. A four-point gain supports the proposal to replace worksheets across the college.

What’s actually going on The pooled result combines two sharply different estimated effects. The groups have equal size, so their effects average to four points. If the differences persist, a college with mostly shared-phone users could get a negative overall result. The experiment also leaves the mechanism open: device access, study conditions, and other characteristics differ together.

What would have helped Specifying device-access groups before the trial, estimating the difference between their treatment effects with an uncertainty interval, and testing a mobile-friendly version. Interviews and usage records could help explain the pattern. A follow-up trial would test whether the apparent harm persists.

These numbers are illustrative. The experiment separates the app’s effect within each group; it does not establish that supplying a laptop would remove the difference.

05Where people trip up

  • Compare effects directly. A statistically significant result in one group and an inconclusive result in another can arise from different sample sizes. Estimate the difference between the effects and its uncertainty. Separate significance labels provide weak evidence for heterogeneity.
  • Resist endless subgroup searches. Split results by age, region, device, income, and dozens of combinations, and some striking patterns will appear by chance. The multiple comparisons problem grows with the search. Specify a few plausible comparisons in advance and check unexpected findings in fresh data.
  • Separate baseline differences from treatment differences. Higher-income customers may spend more under every pricing plan. That establishes a difference in spending levels. Estimating how much each plan changes spending requires a credible comparison within each group. Confounding can distort that comparison when treatment assignment follows existing differences.
  • State the scale of the effect. A treatment that doubles a success rate takes 2% to 4% in one group and 20% to 40% in another. Both have the same relative effect. Their absolute gains are two and twenty percentage points. Costs, benefits, and risks often depend on the absolute change.

A subgroup average also leaves variation within that subgroup. An age band or customer tier is a rough description of individuals. Treating its estimate as a precise prediction for every member adds certainty the data cannot supply.

06Where it doesn’t change the decision

Some decisions can reasonably use a representative average. If a low-cost program will be offered to everyone, the population mix matches the study, and the goal is total benefit, the overall effect may answer the question. Harms, limited places, unequal costs, or fairness requirements make subgroup differences more consequential.

Across studies, apparent heterogeneity can also reflect different measurements or follow-up periods. Before a meta-analysis combines results, researchers need to establish that the studies address comparable questions. More detailed grouping helps only when the groups are meaningful and the estimates have enough precision.

07Roots

At Rothamsted Experimental Station in England, Ronald Fisher faced a practical problem after arriving in 1919: crop yields varied across a field before any fertilizer treatment began. Soil differences could overwhelm the comparison researchers wanted to make. Decades of agricultural records made that problem hard to ignore.

Fisher helped develop experimental designs that accommodated those differences. Grouping similar plots into blocks made treatment comparisons more precise, while randomizing treatments within blocks protected them against bias. His 1935 book, The Design of Experiments, helped carry these methods beyond agriculture. Variation among experimental units became something researchers could explicitly build into a study.

Later, causal inference made variation in treatment response easier to describe. Work on potential outcomes, including Donald Rubin’s 1974 paper, represented a treatment effect as the difference between what would happen to a person under each treatment. Those differences could vary from person to person. As clinical trials accumulated, meta-analysis brought another question into focus: how much did effects differ across studies and settings?

Heterogeneity has no single inventor or discovery date. It grew from repeated attempts to compare unlike cases fairly. Modern applications ask the same practical question: which differences matter enough to change the conclusion?

08How solid is this?

ContestedMixedUsefulEstablished

Variation across populations and settings is widely documented, and methods for estimating it are established. Evidence for any particular subgroup difference depends on the study design, sample size, measurement scale, and whether the pattern survives fresh data.

09Connections

confused withincludesincludesHeterogeneityConfoundingNot written yetVarianceSimpson’sParadoxNot written yetMultipleComparisons ProblemNot written yetExternalValidityNot written yetInteractionEffectNot written yetMeta-AnalysisMean vs. MedianWisdomof Crowds

10Origin and sources

A longstanding concept in statistics with no single inventor. Agricultural experimental design, including Ronald Fisher’s work in the 1920s and 1930s, and later causal inference and meta-analysis developed ways to account for it.

  1. [1]Fisher, R. A. (1935). The Design of Experiments. Oliver and Boyd.
  2. [2]Rubin, D. B. (1974). Estimating causal effects of treatments in randomized and nonrandomized studies. Journal of Educational Psychology, 66(5), 688–701.
  3. [3]Higgins, J. P. T., Thompson, S. G., Deeks, J. J., & Altman, D. G. (2003). Measuring inconsistency in meta-analyses. BMJ, 327(7414), 557–560.
  4. [4]Rothwell, P. M. (2005). Subgroup analysis in randomised controlled trials: importance, indications, and interpretation. The Lancet, 365(9454), 176–186.

Suggest an edit· Updated 2026-10-02