Tool/Research Method/No. 0374

Falsification Test

A falsification test checks whether evidence violates an implication of a claim or research design. Rooted in Karl Popper’s philosophy, it helps researchers probe causal explanations and assumptions using checks such as placebo tests, negative controls, and outcomes that should remain unchanged.

a tool: pick it up

01You've seen this when…

  1. in life

    You credit an evening tea with better sleep. Then you check your log: the improvement starts four nights before your first cup.

  2. at work

    A product team credits its redesigned checkout with higher sales. An analyst runs the same comparison on transactions completed before the redesign and finds a similar increase.

  3. out in the world

    A city report credits a bus lane with cutting commute times. The same drop appears on routes that never use it.

02The idea

An explanation carries implications beyond the result it was built to explain. If a checkout redesign caused sales to rise, its effect should begin after customers could use it. If a tutoring program caused better grades, it should have no effect on grades earned before enrollment.

A falsification test examines one of those implications. A result that clashes with it gives a reason to question the explanation, the comparison being used, or the way uncertainty was calculated.

In causal research, the check often concerns an assumption that makes a comparison credible. Researchers might examine an earlier period, an unaffected outcome, or a group outside the intervention’s reach. A negative control deliberately supplies such an unaffected comparison. A placebo test applies the analysis where the proposed cause should have no effect.

The force of the test comes from the link between the claim and the expected result. Write that link down explicitly. A failed check challenges the claim together with the assumptions used to test it. A passed check means the account survived that particular challenge.

This gives confirmation bias less room to operate: the analyst actively searches for evidence that would make the favored account harder to defend.

03How to use it

  1. Specify the claim and its supporting assumptions. Identify the proposed cause, the outcome, and the comparison. For a policy study, this might be that adopting states would have followed similar wage trends to comparison states without the policy. That assumption supports a difference-in-differences estimate.
  2. Find an implication that can fail. Look for an earlier period, an unaffected outcome, an unexposed group, or a feature of the data that the design requires. Explain why the expected result follows. If the intervention could affect the proposed control through spillovers, choose another check or account for that pathway.
  3. Set the test before inspecting its result. Choose the measure, comparison, time window, and size of violation that would concern you. Check whether the data can detect a violation that matters. Preregistration can record these choices; a dated analysis plan serves the same purpose in a smaller project. Report the full set of checks attempted.
  4. Let the result change the analysis. When the check fails, investigate shared causes, selection, timing, measurement, and statistical assumptions. When it passes, examine the uncertainty around the estimate. A wide interval leaves substantial violations unresolved. Record exactly what the test rules out and which assumptions remain untested.

For a checkout launch, a useful starting check is to move the launch date backward and repeat the analysis on earlier transactions. An apparent effect before launch demands an explanation. Run several sensible placebo dates according to a plan, and preserve all their results.

04A worked example

In a 2004 study, economists Marianne Bertrand, Esther Duflo, and Sendhil Mullainathan investigated how researchers estimated the effects of state policies. They repeatedly assigned fictitious laws to states and dates, then applied difference-in-differences analyses to data on women’s wages.

The invented laws had no causal effect. This gave the researchers a known target: a procedure using a 5% significance threshold should produce false alarms roughly 5% of the time under the null hypothesis. In some specifications, conventional methods reported statistically significant effects for up to 45% of the placebo interventions.

What it looks like An analysis with many wage observations, an estimated policy effect, and a small p-value. Presented as an ordinary policy evaluation, such a result could appear persuasive.

What’s actually going on Wages within a state tend to move together over time. Some conventional calculations treated the errors as more independent than they were, understating uncertainty. Repeating the analysis with fictitious laws exposed how often the procedure could manufacture apparently strong evidence.

What made it work The researchers tested the whole estimation procedure against interventions whose causal effect was known to be zero. Repetition revealed its false-alarm rate. They also examined methods that accounted for dependence within states. The exercise identified a weakness in statistical inference that a single impressive policy estimate could easily conceal.

05When to reach for it

06When it misleads

  • The proposed implication is weak. A bus lane can change traffic on neighboring roads. Finding an effect there may reveal a spillover. The test has force only when the supposedly unaffected comparison has a defensible reason to remain unaffected.
  • The check has little power. A small sample or noisy measure can miss a substantial violation. Inspect the estimate and its interval, and consider statistical power. Failure to reach a significance threshold alone supplies little reassurance.
  • The analyst shops for passing checks. Trying many dates, outcomes, and groups creates opportunities to select comforting results. It also increases the chance of an accidental failure through the multiple comparisons problem. Specify the checks ahead of time and disclose the complete set.
  • A failure gets treated as a complete diagnosis. A pre-policy difference could reflect selection, anticipation, a measurement change, or chance. It challenges the proposed design but may leave several explanations open. Investigate which assumption failed before discarding every possible causal account.

Falsifiability is the broader requirement that a claim expose itself to possible refutation. A falsification test is a particular check. Sensitivity analysis serves a different purpose: it measures how conclusions change as assumptions or inputs vary. Both can strengthen an evaluation, provided their conclusions stay within what they tested.

07Roots

In Vienna, Karl Popper became troubled by theories that seemed able to accommodate almost any human behavior. Einstein’s account of gravity offered a sharper challenge: it predicted how starlight would bend near the sun, creating an opportunity for astronomers to obtain observations that conflicted with the theory.

Popper made exposure to possible refutation central to his philosophy of science. His book Logik der Forschung, first published in the 1930s, reached English-language readers as The Logic of Scientific Discovery in 1959. The underlying problem was how to distinguish a theory that takes empirical risks from an explanation flexible enough to absorb every outcome.

Modern falsification tests also draw on experimental controls and the practical problems of observational research. Economists and epidemiologists developed checks using outcomes, dates, and groups where a proposed cause should leave no trace. These methods put testable implications around assumptions that cannot always be verified directly. Popper supplied an influential philosophical framework; the concrete procedures developed across research fields.

08How solid is this?

ContestedMixedUsefulEstablished

An established method in experimental and causal research, with documented cases where placebo tests exposed unreliable inference. Its force depends on the validity of the expected implication, the test’s power, and which checks were selected. Passing a test leaves other explanations and untested assumptions open.

09Connections

counterscounterscounterscounterspart ofincludesFalsificationTestConfirmationBiasNot written yetCongruence BiasNormalizationof DevianceProcessingFluencyNot written yetPopper'sFalsifiabilityNot written yetNegativeControlConsider-the-OppositeStrategyNot written yetStatisticalPowerNot written yetInternalValidityNot written yetDifference-in-Differences

+ 6 more in the list

10Origin and sources

Karl Popper developed the philosophical account of falsification in the 1930s. Specific placebo, negative-control, and diagnostic tests developed across experimental, epidemiological, and econometric methodology.

  1. [1]Popper, K. R. (1959). The Logic of Scientific Discovery. Hutchinson.
  2. [2]Bertrand, M., Duflo, E., & Mullainathan, S. (2004). How Much Should We Trust Differences-In-Differences Estimates? The Quarterly Journal of Economics, 119(1), 249–275.
  3. [3]Imbens, G. W., & Lemieux, T. (2008). Regression discontinuity designs: A guide to practice. Journal of Econometrics, 142(2), 615–635.
  4. [4]Lipsitch, M., Tchetgen Tchetgen, E. J., & Cohen, T. (2010). Negative Controls: A Tool for Detecting Confounding and Bias in Observational Studies. Epidemiology, 21(3), 383–388.

Suggest an edit· Updated 2026-10-02