Concept/Probability and Statistics/No. 0216
Correlation
Correlation is a measure of how two variables vary together. Developed by Francis Galton and Karl Pearson, its common coefficients describe linear or rank patterns. These summarize the direction and strength of association but cannot, on their own, establish a cause.
- Evidence
- Well established
- Read
- 6 min
- Links
- 14 connections
- Useful when
- Evaluating a claim · Forecasting · Money and investing · Reading data and statistics
01You've seen this when…
- in life
Your mood tracker shows better days when you walk more. Those days also tend to be sunny, with fewer hours at your desk.
- at work
Advertising spending and orders rise together for six months. The marketing team puts the two lines in its budget presentation.
- out in the world
A city report shows that fires attended by more trucks cause more property damage. The biggest fires bring both the largest response and the greatest losses.
02The idea
Picture a scatterplot: one dot per observation, with one variable on each axis. A dot might represent a person’s height and arm span, a day’s temperature and electricity use, or a customer’s spending and number of visits. Correlation compresses part of that picture into a number.
The most familiar version is Pearson’s correlation coefficient, usually written r. It measures how closely the dots follow a straight-line pattern. Its values run from −1 to +1:
- Positive values describe an upward pattern. Larger values of one variable tend to accompany larger values of the other.
- Negative values describe a downward pattern. Larger values of one tend to accompany smaller values of the other.
- Values near zero describe little linear association. A curved relationship can still be strong.
- Values of +1 or −1 describe a perfect straight line. Every observation lies on that line, provided both variables vary.
Pearson’s coefficient compares each value with its variable’s average and scales those differences by the variable’s spread. That makes it independent of ordinary unit changes: measuring height in inches rather than centimeters leaves the correlation unchanged. The number also leaves out the slope. A high correlation can accompany a small change in the outcome.
Spearman’s rank correlation applies the calculation to ranked values. It captures whether higher positions on one measure tend to accompany higher or lower positions on another, including some steadily rising or falling curves.
03Why it matters
Correlation helps identify relationships that deserve attention. A retailer can use the association between weather and demand to improve stock planning. A clinician can investigate a measurement that travels with disease severity. A product team can look for behaviors associated with customers returning.
Prediction can be useful even when the cause remains uncertain. An association gives a forecaster information about what tends to accompany what, as long as the relationship holds in the setting where the forecast will be used.
Correlation also matters when combining risks. Two investments that rise and fall together offer less diversification than their different labels might suggest. Two backup systems exposed to the same conditions can fail together.
The coefficient is one kind of effect size. Its practical meaning depends on the decision: a modest association may improve a forecast across thousands of cases, while a much tighter relationship may be necessary for a safety-critical measurement.
04A worked example
In 1973, statistician Francis Anscombe published four small datasets that became known as Anscombe’s quartet. Each contains 11 pairs of observations. Each has a Pearson correlation of about 0.82, with nearly identical averages and fitted straight lines.
What it looks like Four datasets with much the same relationship. An analyst reading only the summary statistics could give each one the same explanation and use the same forecasting method.
What’s actually going on Their plots tell four different stories. The first shows a roughly linear pattern with scatter. The second follows a curve. The third has most points close to a line, plus one outlying observation. In the fourth, every point except one has the same horizontal value; that single unusual point determines the fitted line and correlation.
The same reported coefficient therefore represents very different evidence about the relationship. In the fourth dataset, there is almost no information about what happens across a range of horizontal values.
What would have helped Drawing all four scatterplots before interpreting the numbers. The curve calls for a different description. The unusual points call for investigation of how they arose and how much the conclusion depends on them. Deleting them automatically would discard information that might explain the process.
05Where people trip up
- Give the number a picture. Pearson’s coefficient summarizes straight-line association. A U-shaped pattern can produce a value near zero even when one variable strongly constrains the other. Plot the observations and inspect nonlinearity, clusters and unusual points before drawing a conclusion.
- Separate association from causal explanation. Advertising spending and orders may rise together because spending increases sales, because strong sales fund larger budgets, or because a holiday season raises both. A shared cause creates confounding. Treating the coefficient alone as proof of an effect is the correlation-causation fallacy.
- Keep magnitude separate from certainty. A correlation of 0.7 based on a handful of observations can be fragile. A correlation of 0.1 in a very large sample can be estimated precisely. Statistical significance addresses evidence against a specified null model; practical importance requires judging the size and consequences of the relationship. Check the sample size and uncertainty interval.
- Check which observations were allowed in. Studying only top-performing applicants compresses the range of ability and can weaken its observed association with later performance. This is restricted range. Conclusions about the selected group may transfer poorly to the full applicant pool.
- Examine groups and time periods. An association across all customers can change or reverse within customer segments, as in Simpson’s paradox. Two monthly series can also correlate because both follow a long-term trend. Compare observations within relevant groups and investigate what the timing contributes.
- Inspect how the variables were recorded. Measurement error can distort correlation. Independent random noise often weakens it; shared recording errors can inflate it. Two survey measures collected through the same response format may share some of that method’s quirks.
06Where it doesn’t settle causation
Correlation alone leaves several causal stories open. A well-designed randomized experiment can help distinguish them by assigning an intervention independently of participants’ existing characteristics.
Observational research can also support causal conclusions when its design and assumptions justify them. Timing, natural experiments and knowledge of the process can strengthen an explanation. The evidential strength comes from that combined argument. A large coefficient by itself still cannot identify causal direction or exclude shared causes.
07Roots
At London’s International Health Exhibition in 1884, Francis Galton ran a laboratory where visitors had their height, strength and other traits measured. He was collecting evidence about human variation and heredity, interests tied to the eugenics movement he helped found. One statistical problem was straightforward: how could he measure the tendency for different bodily dimensions to go together?
Someone tall tended to have larger arm measurements, but the correspondence was incomplete. Galton needed a way to express that tendency across measurements with different units and spreads. In an 1888 paper, he described “co-relations” using anthropometric data. His approach helped turn the observation that traits accompany one another into a measurable relationship.
Karl Pearson gave the product-moment coefficient its familiar mathematical form in the 1890s, developing it in work on regression and heredity. Charles Spearman extended the approach to ranks in 1904, making it possible to study ordered measurements without relying on their original numerical scales. Correlation traveled from these studies into psychology, economics and routine data analysis, where its compactness made it useful—and easy to interpret without looking at the underlying observations.
08How solid is this?
The mathematical properties of Pearson’s and Spearman’s coefficients are well established. Whether a particular correlation supports prediction or a causal explanation depends on sampling, measurement, study design and whether the relationship persists.
09Connections
- Often confused with Correlation-Causation Fallacy, Feedback Loops, Independence
- Countered byRandomized Experiment
- Part of Effect Size
- See also Measurement Error, Restricted Range, Diversification, Confounding, Simpson’s Paradox, Nonlinearity, Statistical Significance, Regression to the Mean, Wisdom of Crowds
+ 4 more in the list
10Origin and sources
Francis Galton developed an approach to measuring correlation in the 1880s and described it in 1888. Karl Pearson formalized the product-moment coefficient in the 1890s; Charles Spearman introduced rank correlation in 1904.
- [1]Galton, F. (1888). Co-relations and their measurement, chiefly from anthropometric data. Proceedings of the Royal Society of London, 45, 135–145.
- [2]Pearson, K. (1896). Mathematical contributions to the theory of evolution.—III. Regression, heredity, and panmixia. Philosophical Transactions of the Royal Society of London. Series A, Containing Papers of a Mathematical or Physical Character, 187, 253–318.
- [3]Spearman, C. (1904). The Proof and Measurement of Association between Two Things. The American Journal of Psychology, 15(1), 72–101.
- [4]Anscombe, F. J. (1973). Graphs in Statistical Analysis. The American Statistician, 27(1), 17–21.
Suggest an edit· Updated 2026-10-02