Trap/Cognitive Bias/No. 0077

Base Rate Fallacy

The base rate fallacy, also called base rate neglect, is giving specific clues too much weight and too little to how common an outcome is in the relevant population. Studied by Daniel Kahneman and Amos Tversky in psychology, it can distort probability judgments.

Also called Base Rate Neglect

a trap: easy to walk into

01You've seen this when…

  1. in life

    A home test flags a rare condition. You focus on the advertised detection rate and barely consider how often healthy people get the same result.

  2. at work

    A candidate’s résumé reads like your idea of a star. You predict an exceptional hire without checking how often candidates with similar résumés actually excel.

  3. out in the world

    A city announces hundreds of facial-recognition matches to a watch list. The report doesn’t say how many faces were scanned or how often people outside the list trigger a match.

02The idea

A positive test gives you information about one case, just as an impressive résumé or a suspicious transaction does. To estimate the probability of the outcome, combine the clue with how common that outcome was before the clue arrived.

That starting frequency is the base rate. The base rate fallacy is giving it too little weight compared with the details of the particular case.

A test that catches most cases of a disease can still produce mostly false alarms if the disease is rare and the test sometimes flags healthy people. Catching someone who has the disease and confirming that someone has it are different conditional probabilities.

The remedy is to combine the starting rate with individual evidence, taking into account how strongly the evidence distinguishes one explanation from another. That’s the logic of Bayesian updating.

03Why it happens

  • The individual story feels more informative. A detailed description of one person is easier to picture than a proportion across thousands. The details can take over even when they offer little help in distinguishing the possibilities.
  • Resemblance substitutes for probability. Someone seems like an engineer, so you judge them likely to be an engineer. That’s the representativeness heuristic. Using resemblance as a shortcut can lead you to underweight prevalence, the error known as base rate neglect.
  • The comparison group disappears. You hear how often a test catches disease while its false-alarm rate goes unreported. You need both groups to tell what share of positive results come from people with the disease.
  • General information seems less relevant. A statistic about a population can feel remote from the person in front of you. Sometimes it describes the wrong population. Sometimes it is exactly the evidence you need, but it loses to a vivid detail.

These are tendencies, not a rule that people always ignore statistics. Research finds substantial variation depending on the task, the information’s relevance, and how the numbers are presented.

04A worked example

Consider an invented screening program for a rare disease. Among the people being screened, 1% have it. The test detects 90% of cases and produces a positive result in 5% of people without the disease.

What it looks like A positive result seems close to a diagnosis. The test catches nine out of ten cases, so you assume your positive result means a roughly 90% chance of having the disease.

What’s actually going on Imagine screening 10,000 people. Of the 100 who have the disease, 90 test positive. Of the 9,900 who don’t, 495 also test positive. That makes 585 positive results, only 90 of which come from people with the disease. The probability of disease after a positive result is about 15%, not 90%.

The result matters: it raises the probability from 1% to about 15%. But a large healthy population supplies enough false alarms to outnumber the true detections. The share of positive results that come from people with the disease is the positive predictive value.

What would have helped Laying out both populations before interpreting the result, then seeking appropriate confirmation. These numbers are illustrative, not guidance for any particular medical test.

05How to spot it

06What to do instead

  • Begin with the right comparison group. Before studying the individual case, estimate how often this outcome occurs in similar cases. Reference-class forecasting applies this approach to projects and predictions. Choose the group based on its fit with the case. Evaluate that fit independently of the answer you prefer.
  • Turn percentages into people or cases. Imagine 1,000 or 10,000 observations. Count how many have the condition and how many don’t, then tally the positives in each group. Research shows that this frequency format can make Bayesian problems easier with formulas left out of the instruction.
  • Ask for both kinds of test performance. How often does the signal appear when the condition is present? How often when it’s absent? Sensitivity and specificity describe these two sides. A single accuracy figure can conceal the distinction.
  • Check whether the detail actually separates the possibilities. An impressive founder matters less if failed ventures also commonly have impressive founders. Useful evidence is more expected under one explanation than the other.
  • Separate probability from action. A 15% chance can justify further testing or a precaution if the consequences are serious. You can take the risk seriously while recognizing substantial uncertainty about whether the outcome will occur.

If the base rate is uncertain, try a plausible low and high value. See whether your conclusion changes enough to affect the decision.

07When it isn’t base rate neglect

A population rate deserves weight only if it fits the case. The prevalence among people receiving routine screening may differ greatly from the prevalence among patients referred because of specific symptoms. Sound reasoning favors a statistic that fits the case, even when a broader one is available.

Strong evidence can also overwhelm a low starting probability. A rare event can become plausible when the evidence supports it. The question is whether the evidence justifies the size of the update.

And some apparent neglect reflects uncertainty about where the numbers came from. If a supplied base rate is unreliable or already accounts for the same evidence you’re adding, using it mechanically can make things worse. Choose a defensible starting probability by checking the available statistics for reliability and overlap with the evidence.

08Roots

In a 1973 study, Daniel Kahneman and Amos Tversky gave participants short personality descriptions drawn from groups of engineers and lawyers. One group contained 70 engineers and 30 lawyers; another had the proportions reversed. The descriptions stayed the same. Would changing the composition of the group change people’s judgments?

Changing the group proportions often shifted people’s judgments much less than probability reasoning required. Participants leaned heavily on how much a description sounded like an engineer or a lawyer. When they received only the group proportions, they used those proportions more readily. Participants gave the personal sketch more weight than statistics they could understand.

Maya Bar-Hillel’s 1980 paper examined the base-rate fallacy through the question of perceived relevance: why does particular information so often outrank general information? Later work challenged the sweeping claim that people routinely ignore base rates, while experiments with frequency formats showed that presentation can substantially improve performance.

One proposed adaptive account is that people reason more easily from accumulated cases than from abstract percentages because ordinary experience supplies cases. That evolutionary explanation remains a hypothesis. The demonstrated benefit of counting cases leaves open why that benefit evolved.

09How solid is this?

ContestedMixedUsefulEstablished

Underweighting base rates is well documented in probability-judgment tasks. It is not universal: relevance, task design, and frequency formats affect performance, and some apparent errors depend on questionable assumptions about the correct starting probability.

10Connections

confused withcountered bycountered bycountered bycountered bycountered bycountered byfollows fromfollows fromBase RateFallacyRepresentativenessHeuristicNot written yetReference-ClassForecastingBayes’ TheoremNot written yetSensitivityand SpecificityNot written yetPositivePredictive ValueBayesianUpdatingBayesian PriorAvailabilityHeuristicSalience BiasConditionalProbability

+ 2 more in the list

11Origin and sources

Daniel Kahneman and Amos Tversky documented base rate neglect in 1973. Maya Bar-Hillel analyzed the base-rate fallacy and the role of perceived relevance in 1980.

  1. [1]Kahneman, D., & Tversky, A. (1973). On the psychology of prediction. Psychological Review, 80(4), 237–251.
  2. [2]Bar-Hillel, M. (1980). The base-rate fallacy in probability judgments. Acta Psychologica, 44(3), 211–233.
  3. [3]Gigerenzer, G., & Hoffrage, U. (1995). How to improve Bayesian reasoning without instruction: Frequency formats. Psychological Review, 102(4), 684–704.
  4. [4]Koehler, J. J. (1996). The base rate fallacy reconsidered: Descriptive, normative, and methodological challenges. Behavioral and Brain Sciences, 19(1), 1–17.

Suggest an edit· Updated 2026-10-02