Concept/Probability and Statistics/No. 0193

Confounding

Confounding is the mixing of a possible causal effect with other differences between groups that affect the outcome. In causal inference and epidemiology, a shared cause can influence both exposure and outcome, making an observed association exaggerate, hide or reverse the effect.

Also called Confounding Bias

a concept: name it

01You've seen this when…

  1. in life

    People who take vitamins seem healthier, so you buy a bottle. You haven’t compared their diets, smoking or exercise.

  2. at work

    Employees who volunteer for a training course earn higher performance ratings. Management credits the course, without checking whether its volunteers were already the most motivated employees.

  3. out in the world

    A school tops the district’s exam rankings. Families rush to enroll, but its students also come from households that can afford tutoring and stable housing.

02The idea

The training course and the performance ratings move together. But that leaves two explanations tangled: perhaps the course improves performance, or perhaps motivated employees both choose the course and perform better. Both could be true.

Confounding is this mixing of explanations. The comparison you observe combines the relationship you want to understand with other differences that affect the outcome.

The simplest pattern has a shared cause. Motivation influences both course attendance and job performance. Disease severity influences both whether someone receives treatment and whether they recover. A causal diagram draws two arrows out of that shared cause, making the alternative explanation visible.

The causal question is what would happen to comparable people under different choices. The observed comparison may instead put different kinds of people on each side. Counterfactual reasoning keeps that distinction clear: would these employees have performed better because they attended, compared with how they would have performed without attending?

A confounder can exaggerate an effect, hide it or reverse its apparent direction. An association can still reflect a causal effect even when a confounder is present, but the association alone cannot settle the causal question.

03Why it matters

Confounding can make a sensible decision look foolish, or a useless intervention look effective. Hospitals that treat the sickest patients and provide better care may still have worse survival rates. A successful investment strategy may owe its returns to taking more risk, so those returns alone leave its skill at picking better assets unclear.

This matters whenever you plan to act on a comparison. Buying the supplement, copying the investment strategy or funding the training course assumes that changing the apparent cause will change the outcome. Confounding weakens that assumption.

Sometimes the aggregate comparison even points in the opposite direction from comparisons within relevant groups. That’s one way Simpson’s paradox can arise. Confounding can also quietly shift an estimate while leaving its sign unchanged.

Evaluate associations case by case: identify why the groups differ, then ask whether the evidence separates those differences from the effect you care about.

04A worked example

In 1948, Britain’s Medical Research Council published a trial of streptomycin for pulmonary tuberculosis. Patients received either streptomycin plus bed rest or bed rest alone. The study used random allocation, with assignments concealed from clinicians until a patient had been accepted into the trial.

What it looks like An uncontrolled comparison would seem straightforward: compare survival among patients receiving the new drug with survival among patients not receiving it. Better survival in the first group would seem to show that the drug works.

What’s actually going on Treatment decisions can themselves carry information about a patient’s prospects. Doctors might prioritize the sickest patients, making an effective drug look worse. Or they might prioritize patients thought most likely to recover, making it look better. These are hypothetical selection mechanisms, not descriptions of what happened in the trial. They show why disease severity and prognosis could confound a comparison based on doctors’ choices.

What would have helped The investigators paired two safeguards: random allocation prevented prognosis from systematically determining treatment assignment, and concealing the assignments prevented clinicians from steering particular patients into a preferred group. Over six months, fewer patients died in the streptomycin group, and chest X-rays showed greater improvement.

The lesson isn’t that every medical question needs this exact design. It’s that a credible control group must represent what would have happened without treatment, rather than merely consist of people who happened not to receive it.

05Where people trip up

  • They mistake a large sample for a fair comparison. More observations can make a confounded estimate more precise without making it more accurate. Before celebrating a tiny margin of error, examine how people entered each group.
  • They adjust for everything available. Not every measured variable belongs in the analysis. A consequence of treatment can sit on the pathway through which treatment works. Adjusting for it may remove part of the very effect you’re trying to estimate. Use a causal account to choose among the variables in your spreadsheet.
  • They control for a shared consequence. Restricting a study to hospitalized patients, for example, can connect otherwise unrelated causes of hospitalization. This shared-consequence pattern is collider bias. Adjustment can create bias as well as remove it.
  • They treat adjustment as a certificate. Accounting for age and income doesn’t establish that motivation, illness or other relevant differences are handled. Check whether the important factors were measured well, whether comparable people exist in both groups, and what remains unmeasured.
  • They assume matching makes people identical. Two patients of the same age can have very different prognoses. Matching and statistical adjustment can address measured differences, but they do not automatically address hidden ones.
  • They invoke confounding without naming a mechanism. Any inconvenient result can be dismissed by vaguely mentioning other factors. A useful challenge names a plausible factor, explains how it affects both the comparison and the outcome, and asks what evidence would test that explanation.

For an important decision, write down how exposure or treatment was chosen before looking at the result. Then draw the likely causes of both that choice and the outcome. If the comparison remains doubtful, look for a randomized experiment or another design with a credible reason the groups are comparable.

06When it isn’t confounding

The causal role of a difference between groups determines whether it is a confounder. A variable associated with treatment can be irrelevant to the outcome and leave the comparison undistorted. Other study errors include selection bias, poor measurement and reverse causation. Each raises a different problem.

Randomization addresses confounding by balancing background causes on average. A particular trial can still have chance imbalances, and problems after assignment can undermine its comparison.

Observational evidence isn’t automatically unusable. Carefully chosen comparisons and adjustment can support causal conclusions. But the conclusion depends on assumptions about which relevant differences have been accounted for. A regression output doesn’t verify those assumptions for you.

07Roots

When Ronald Fisher arrived at Rothamsted Experimental Station in 1919, he inherited decades of agricultural records. Its long-running field experiments included wheat plots receiving different fertilizers. The difficulty was familiar to any gardener: neighboring patches of ground can differ in fertility. A larger harvest could reflect the fertilizer, the soil or both.

Fisher helped turn that difficulty into principles of experimental design. Blocking grouped comparable plots; randomization stopped investigators from systematically giving one treatment the better ground. His 1935 book, The Design of Experiments, helped spread these methods beyond agriculture. The recognition of confounding developed through multiple contributions. Experimental design made the problem explicit and offered ways to prevent it.

Epidemiology faced a harder version. Researchers could not assign people to smoke for decades or choose their childhood living conditions. They had to compare existing groups while separating an exposure from the other conditions accompanying it.

Later causal frameworks sharpened the question. Donald Rubin described treatment effects through outcomes under alternative treatments. Judea Pearl used causal diagrams to show which adjustments separate explanations and which introduce new bias. James Robins developed methods for settings where treatment and health change together over time; his work with Miguel Hernán brought these tools into a practical framework for causal inference. The recurring problem stayed the same: distinguish what an intervention does from why particular people receive it.

08How solid is this?

ContestedMixedUsefulEstablished

Confounding is a foundational problem in causal inference, supported by mathematical results and experimental practice. Whether a particular study has adequately addressed it depends on its design, measurements and causal assumptions.

09Connections

confused withconfused withconfused withcountered bycountered bycountered bycountered bypart ofincludesConfoundingSelection BiasNot written yetCollider BiasHeterogeneityNot written yetRandomizedExperimentCounterfactualReasoningNot written yetControl GroupNot written yetDirectedAcyclic GraphCorrelation-CausationFallacySimpson’sParadoxConditionalProbability

+ 4 more in the list

10Origin and sources

Developed through experimental design and epidemiology, with no single discoverer. Fisher’s work in the 1920s–1930s established key design safeguards; Rubin, Pearl, Robins and Hernán developed modern causal frameworks.

  1. [1]Fisher, R. A. (1935). The Design of Experiments. Oliver and Boyd.
  2. [2]Medical Research Council (1948). Streptomycin Treatment of Pulmonary Tuberculosis. British Medical Journal, 2(4582), 769–782.
  3. [3]Rubin, D. B. (1974). Estimating causal effects of treatments in randomized and nonrandomized studies. Journal of Educational Psychology, 66(5), 688–701.
  4. [4]Pearl, J. (1995). Causal diagrams for empirical research. Biometrika, 82(4), 669–710.
  5. [5]Hernán, M. A., & Robins, J. M. (2020). Causal Inference: What If. Chapman & Hall/CRC.

Suggest an edit· Updated 2026-10-02