Concept/Probability and Statistics/No. 0307
Effect Size
Effect size is a statistical measure of the size of a difference or the strength of a relationship. Measures include Cohen’s d and correlation. Its practical importance depends on units, uncertainty, baseline risk, and context, rather than statistical significance alone.
- Evidence
- Well established
- Read
- 6 min
- Links
- 9 connections
01You've seen this when…
- in life
You consider paying for a sleep app. Its website says users sleep significantly longer, but the study reports an average gain of four minutes a night.
- at work
A new onboarding flow raises satisfaction from 3.9 to 4.0 out of 5. The dashboard turns green, and your team has to decide whether that gain justifies a month of engineering work.
- out in the world
A headline says a safety measure cuts the risk of an accident in half. The report puts the risk at two accidents per thousand journeys before the change and one afterward.
02The idea
A result can be real and still barely change anything. It can also be large enough to matter while remaining too uncertain to trust. Effect size addresses the first question: how much difference is there, or how strong is the relationship? Uncertainty addresses the second: how precisely have we measured it?
Effect size is a family of statistics for expressing magnitude. The useful choice depends on what you’re comparing:
- Keep original units when they help, and show both absolute and relative changes for rates. A bus route saves six minutes per trip. A treatment lowers blood pressure by five units on the scale clinicians use. These differences are easy to connect to a decision. A risk falling from 2% to 1% falls by one percentage point and by 50% relative to its starting level. Both descriptions are correct. They answer different questions.
- Standardize when the scales differ. Cohen’s d expresses a difference between group means in units of a standard deviation, usually based on the spread within the groups. A d of 0.5 means the means differ by half that spread. A 50% improvement in performance uses starting performance as its reference.
For a relationship between two quantities, a correlation can serve as an effect size. It describes how strongly they move together in a linear pattern. Establishing that one causes the other requires evidence beyond that relationship.
The number needs its label. An effect of 0.5 could mean half a minute or half a percentage point in the original units, whereas a difference of 0.5 standard deviations means half the spread. Each label identifies a different quantity.
03Why it matters
Statistical significance concerns compatibility with a statistical model. Judging whether an effect is worth acting on requires its magnitude and context. A p-value measures how incompatible the data are with a specified statistical model, often one with no difference—not the size of a benefit.
With enough observations, a tiny difference can clear a significance threshold. With few observations, a worthwhile difference can remain uncertain. Reporting only the threshold verdict hides both possibilities.
Effect size makes the decision concrete. Six minutes saved might be negligible for an occasional passenger but valuable across millions of journeys. A small reduction in a serious side effect might justify a change; the same reduction in a minor inconvenience might not.
Magnitude also connects research to cost-benefit analysis. You can compare the expected improvement with money, time, disruption, and harms. That comparison requires choosing the outcome you actually care about, even when another outcome is easier to measure.
Read an effect estimate alongside a confidence interval. A claimed benefit of six minutes means something different when the interval runs from five to seven minutes than when it runs from a two-minute loss to a fourteen-minute gain.
04A worked example
Consider an illustrative checkout experiment. An online store randomly assigns 80,000 independent shoppers to two groups. Half see the existing checkout; half see a version offering a $2 discount. The existing version produces 4,000 orders from 40,000 shoppers. The discount version produces 4,200.
What it looks like The new checkout wins. Conversion rises from 10% to 10.5%, a 5% relative increase. The difference is statistically significant at the conventional 5% level.
What’s actually going on The absolute effect is 0.5 percentage points: five extra orders per thousand shoppers. An approximate 95% confidence interval runs from 0.08 to 0.92 percentage points, so the magnitude remains somewhat uncertain even though the test clears the threshold.
Now consider profit. Suppose each order contributes $20 before the discount, and all buyers using the new checkout receive it. At 100,000 shoppers, the estimates imply 10,000 orders contributing $200,000 under the existing checkout, versus 10,500 orders contributing $189,000 under the discount version. The conversion gain comes with an estimated $11,000 reduction in contribution. A positive effect on one outcome is a negative effect on another.
What would have helped Choosing contribution per shopper as a decision outcome before running the randomized experiment, while keeping conversion as a supporting measure. Reporting absolute and relative changes with their uncertainty would make the tradeoff visible. The profit estimate needs its own interval to inform the business decision, even when the conversion interval is available.
05Where people trip up
- Treating a benchmark as a verdict. Cohen’s familiar d benchmarks of 0.2, 0.5, and 0.8 are often labeled small, medium, and large. Their boundaries come from convention. A small standardized effect can matter greatly for a serious outcome or a widely used intervention. Compare with meaningful changes in the field before using a generic label.
- Losing the denominator. A standardized effect can grow because the average difference grows or because individual scores vary less. Two studies can therefore report different values of d for the same change in original units. Inspect the raw difference and the spread before ranking interventions.
- Letting relative change hide the starting point. Halving a risk of one in ten is very different from halving a risk of one in a million. Ask for the baseline and the absolute change over the same period. Relative changes are useful when paired with that context.
- Calling an uncertain result negligible. A useful effect can remain compatible with the data when a significance test is inconclusive. Check whether the estimate’s interval includes changes large enough to matter. A large point estimate from a small study may be very imprecise. Statistical power and sample size affect what you can distinguish. The outcome’s practical importance guides what you should value.
- Measuring the wrong thing precisely. A sizable change in a questionnaire score can correspond to a small or zero change in daily life. Check construct validity: does this outcome capture what you care about? Also separate magnitude from causation. A strong association can arise from other differences between the groups.
06Roots
In the early 1960s, psychologist Jacob Cohen examined research published in the Journal of Abnormal and Social Psychology. The papers contained sample sizes, statistical tests, and verdicts about significance. He asked a different question: how often would these study designs detect an effect if one were really there?
To answer it, he needed to specify how large that effect might be. A study could have a good chance of detecting a large difference and little chance of detecting a small one. His 1962 review helped expose the limited statistical power of much psychological research. A significance verdict wasn’t enough to understand what a study could teach you.
Means, ratios, and correlations already served to measure differences and relationships before Cohen’s work. His influential contribution was to organize standardized measures around power analysis and give researchers a shared vocabulary for planning studies. His 1969 book, followed by an expanded edition in 1988, carried that approach across the behavioral sciences.
The convenient small, medium, and large benchmarks traveled especially well. Sometimes they traveled too well: rough reference points became universal grading rules. The lasting lesson is broader than those labels. State the magnitude you care about, design a study capable of detecting it, and report what you found in terms readers can interpret.
07How solid is this?
Effect sizes are standard statistical measures with well-understood estimation methods. There is no universal cutoff for practical importance; interpretation depends on the outcome, uncertainty, costs, and context.
08Connections
- Often confused withP-Value, Statistical Significance
- Part of Cost-Benefit Analysis, Statistical Power
- Includes Correlation
- See alsoConfidence Interval, Randomized Experiment, Standard Deviation, Construct Validity
09Origin and sources
Measuring differences and relationships comes from the broader statistical tradition, with no single inventor. Jacob Cohen popularized standardized effect-size measures through his books on statistical power analysis in 1969 and 1988.
- [1]Cohen, J. (1962). The statistical power of abnormal-social psychological research: A review. Journal of Abnormal and Social Psychology, 65(3), 145–153.
- [2]Cohen, J. (1969). Statistical Power Analysis for the Behavioral Sciences. Academic Press.
- [3]Cohen, J. (1988). Statistical Power Analysis for the Behavioral Sciences (2nd ed.). Lawrence Erlbaum Associates.
- [4]Sullivan, G. M., & Feinn, R. (2012). Using Effect Size—or Why the P Value Is Not Enough. Journal of Graduate Medical Education, 4(3), 279–282.
Suggest an edit· Updated 2026-10-02