Trap/Logical Fallacy/No. 0217
Correlation-Causation Fallacy
The correlation-causation fallacy is treating an association as proof that one thing causes another. Also called cum hoc ergo propter hoc, it is an error in logic and statistics: the link may instead reflect a shared cause, reverse causality, or selection.
Also called Cum Hoc Ergo Propter Hoc
- Evidence
- Well established
- Read
- 6 min
- Links
- 11 connections
01You've seen this when…
- in life
Your fitness app shows you sleep longer on days you walk more. You credit the steps, without noticing those are also your days off.
- at work
Teams that hold more planning meetings ship more features. Management adds meetings to every team’s calendar without comparing team sizes.
- out in the world
A city report shows neighborhoods with more police have more reported crime. A commentator treats the chart as proof that sending officers causes crime.
02The idea
Two things move together. You give one credit for producing the other. Then you act as though changing the first will change the second.
That last step needs more evidence. Correlation tells you that knowing one thing gives you information about another. Causation tells you something stronger: changing one would make a difference to the other, under the conditions being considered.
Larger teams may hold more meetings and ship more features. Adding meetings to a small team doesn’t give it the extra engineers responsible for the relationship.
The fallacy is not noticing an association or proposing a causal explanation. Those are often useful first steps. It is treating the explanation as established without addressing competing explanations.
This differs from the post hoc fallacy, which assigns causation because one event follows another. Here, the mistake can happen without any clear sequence at all. The evidence may simply be a chart showing two things together.
03Why it happens
An association is easy to see. The process behind it usually isn’t. We fill that gap with a story, often the one that fits our expectations or offers an easy intervention.
Several different processes can produce the same-looking relationship:
- A third factor drives both. Days off allow more walking and more sleep. This is confounding: something else helps produce both the supposed cause and its supposed effect.
- The direction runs backward. A struggling student gets more tutoring. Students receiving tutoring may therefore have lower grades, even if tutoring helps them. The difficulty prompts the response.
- The sample creates the pattern. Restricting a study to people admitted to a selective college can distort the relationship between grades and other admission strengths. Selection depends on both. That’s a form of selection bias.
- Chance supplies a convincing match. Search enough charts, time periods, or subgroups and some variables will line up accidentally. Choosing the best-looking match afterward conceals how many opportunities there were to find it.
These explanations can overlap. More police may follow more crime, and more policing may also increase how much crime gets recorded. A single chart cannot separate those processes.
04A worked example
Observational research, including the Nurses’ Health Study, found lower coronary heart disease rates among postmenopausal women using hormone therapy. The findings helped build a case for using hormone therapy to prevent heart disease.
What it looks like Women taking the treatment have healthier hearts, so the treatment protects the heart. Prescribing it should give other women the same benefit.
What’s actually going on Treatment users and nonusers were not interchangeable groups. They could differ in health behaviors, access to care, and other factors affecting heart disease. In 2002, the Women’s Health Initiative randomized trial of estrogen plus progestin found no coronary protection and instead found increased coronary heart disease risk in the population studied.
The contrast needs care. How much of the earlier association came from healthier women choosing treatment remains unresolved by this comparison. Treatment regimens, age, time since menopause, and the way studies defined treatment starts and follow-up also mattered. Later analyses showed that making observational comparisons more like the randomized comparison helped reconcile results.
What would have helped Treating the observational finding as a prevention claim that still needed testing. Researchers needed comparisons of similar treatments, with comparable timing for treatment starts and follow-up. They also needed randomization where feasible. The lesson concerns the inference about heart disease prevention. The broader judgment on hormone therapy for menopausal symptoms remains separate.
05How to spot it
06What to do instead
- Specify the intervention. Replace a vague association with a precise claim: adding one planning meeting per week will increase completed features for teams of this size.
- Draw the competing explanations. List what could influence both variables, what might run backward, and how cases entered the sample. A simple causal diagram can expose assumptions that prose hides.
- Ask for the missing comparison. What would happen to otherwise comparable people or teams without the change? This is counterfactual reasoning: estimating what would happen under each option while accounting for differences between those who chose them.
- Improve the study design. Use a randomized experiment when practical and ethical. Otherwise, look for a credible natural experiment or another design that addresses the main alternatives. Statistical adjustment helps only with differences adequately measured and modeled.
- Look for independent lines of evidence. Combine designs with different weaknesses. This triangulation is stronger than repeating the same flawed comparison in several datasets.
- Match your wording to your evidence. Report an association when that is what you measured. State what assumptions support any stronger causal conclusion.
07When it isn’t a causal fallacy
An association can be useful for prediction regardless of its causal explanation. Search activity might help predict product demand even if increasing searches leaves demand unchanged or reduces it. Prediction and intervention are different jobs.
Evidence beyond randomized trials can support causal conclusions. Smoking’s role in lung cancer rests heavily on observational and mechanistic evidence. Consistency across studies and exposure preceding illness can support a causal case. Dose patterns and credible tests of alternatives can add support. None is an automatic proof by itself.
Sometimes several causal directions operate together. Income can affect health, and health can affect income. Finding reverse causality leaves open whether a forward effect also exists.
The right standard allows uncertainty while requiring evidence strong enough for the claim and the decision, with the remaining assumptions made visible.
08Roots
In 1965, British medical statistician Austin Bradford Hill addressed a problem with lives at stake: when should an observed link between an exposure and illness justify action? Smoking and lung cancer made the problem urgent. Researchers could not responsibly settle it by assigning people to years of cigarette smoking.
Hill described ways to weigh a causal case. He considered timing and consistency and examined whether larger exposures brought larger effects. He presented these considerations as guides to judgment. Using them required weighing the evidence in each case. His aim was practical: avoiding both easy certainty and demands for evidence so perfect that action would never follow.
The warning predates Hill and belongs to longstanding statistical and logical traditions; the modern fallacy label has no single well-established inventor. Later causal-inference researchers, including Judea Pearl, made the distinction more explicit through diagrams and assumptions about how variables influence one another. That work turned the familiar warning about inferring causation from correlation into a sharper task: identify which comparisons could separate the causal explanations.
09How solid is this?
The logical limitation is established: different causal processes can produce the same association. Research documents these problems across fields; identifying a particular cause still depends on study design and explicit assumptions.
10Connections
- Often confused with Correlation
- Countered by A/B Testing, Triangulation, Counterfactual Reasoning, Directed Acyclic Graph, Randomized Experiment, Natural Experiment
- Includes Confounding, Post Hoc Fallacy, Selection Bias
- See also Necessary vs. Sufficient Conditions
+ 1 more in the list
11Origin and sources
A longstanding distinction in statistics and logic, without a single established coiner. Austin Bradford Hill (1965) described considerations for causal judgment; later causal-inference traditions formalized the distinction.
- [1]Hill, A. B. (1965). The Environment and Disease: Association or Causation? Proceedings of the Royal Society of Medicine, 58(5), 295–300.
- [2]Grodstein, F., Stampfer, M. J., Manson, J. E., Colditz, G. A., Willett, W. C., Rosner, B., Speizer, F. E., & Hennekens, C. H. (1996). Postmenopausal Estrogen and Progestin Use and the Risk of Cardiovascular Disease. New England Journal of Medicine, 335(7), 453–461.
- [3]Writing Group for the Women's Health Initiative Investigators. (2002). Risks and Benefits of Estrogen Plus Progestin in Healthy Postmenopausal Women: Principal Results From the Women's Health Initiative Randomized Controlled Trial. JAMA, 288(3), 321–333.
- [4]Hernán, M. A., Alonso, Á., Logan, R., Grodstein, F., Michels, K. B., Willett, W. C., Manson, J. E., & Robins, J. M. (2008). Observational Studies Analyzed Like Randomized Experiments: An Application to Postmenopausal Hormone Therapy and Coronary Heart Disease. Epidemiology, 19(6), 766–779.
- [5]Pearl, J. (2009). Causality: Models, Reasoning, and Inference (2nd ed.). Cambridge University Press.
Suggest an edit· Updated 2026-10-02