Concept/Probability and Statistics/No. 0084

Bayes’ Theorem

Bayes’ theorem, or Bayes’ rule, is a mathematical identity that updates the probability of a hypothesis given new evidence. Named after Thomas Bayes, it combines the prior probability and the likelihood of the evidence to calculate an updated posterior probability.

Also called Bayes' Rule

a concept: name it

01You've seen this when…

  1. in life

    A screening test flags a rare condition. You read that the test catches most cases and assume the result means you probably have it.

  2. at work

    A fraud detector flags an order. Your colleague wants to cancel it immediately, but you know the store processes thousands of legitimate orders for every fraudulent one.

  3. out in the world

    A news report says a suspect matches a witness’s description. The description sounds compelling until you realize it also fits hundreds of people in the neighborhood.

02The idea

A positive test provides one piece of evidence. Interpreting it requires knowing how common the condition was before testing and how often the test gives the same result in people without it.

Bayes’ theorem connects those pieces. It calculates the probability of a hypothesis after seeing evidence, using its probability beforehand and the probability of that evidence under competing possibilities.

The starting probability is the prior. The probability after incorporating the evidence is the posterior. The likelihood asks how probable the evidence would be if the hypothesis were true. These describe different probabilities: a test can catch most sick people while most positive results come from people without the condition.

The formula is:

P(H | E) = P(E | H) × P(H) / P(E)

Here H is the hypothesis and E is the evidence. The vertical bar means given: P(H | E) is the probability of H given E. The denominator, P(E), counts all the ways the evidence could appear, including when H is false.

Another useful form is:

Updated odds = starting odds × likelihood ratio.

The likelihood ratio compares how probable the evidence is under one possibility versus another. Evidence that is equally probable under both doesn’t change their relative odds. Evidence that is much more probable under one shifts the balance toward it.

03Why it matters

Strong-looking evidence can mean little when there are many opportunities for a false alarm. A rare disease can produce a convincing signal. Security breaches and fraudulent transactions can do the same. Far more harmless cases may produce that signal too.

The theorem prevents the base rate fallacy: forgetting how common each possibility was before the new detail caught your attention. It also prevents the opposite mistake of refusing to change your mind when evidence that distinguishes between possibilities arrives.

Beyond calculating a precise percentage, the theorem gives you a structure for reasoning with uncertain inputs. Assess how plausible the claim was beforehand, then consider separately how expected the evidence is if the claim is true and how expected it is otherwise.

That separation makes disagreements easier to locate. Two people can agree about a test result but disagree about the population it applies to.

04A worked example

Consider an invented screening program for 10,000 people. The condition affects 1% of this population. The test detects 90% of people who have it and incorrectly flags 5% of people who don’t.

What it looks like Your result is positive. You hear that the test detects 90% of cases and take that to mean a 90% chance that you have the condition.

What’s actually going on Follow the people rather than the percentage:

  • Of 10,000 people, 100 have the condition. The test flags 90 of them. Of the other 9,900 people without it, the test flags 495.
  • There are 585 positive results altogether. Only 90 come from people with the condition.

So the probability of having the condition given a positive result is 90 / 585, or about 15.4%. This is the test’s positive predictive value in this population. The result raises the probability substantially, from 1% to roughly 15%. The diagnosis remains uncertain.

What would have helped Ask for the chance of having the condition after a positive result to put the test’s detection rate in context. Use rates from people like you, and discuss appropriate confirmation with a clinician. Counting outcomes in a sample of 10,000 makes the false positives visible; research finds that this kind of frequency format often makes Bayesian problems easier to solve.

05Where people trip up

  • Keep the two directions separate. The chance of a positive test given illness is not the chance of illness given a positive test. These are different conditional probabilities. Reversing them is the central mistake.
  • Choose a relevant starting population. A rate for someone referred because of specific symptoms may differ from the rate for everyone screened. The relevant rate can depend on a person’s age and exposure and on the criteria for referral. Reference-class forecasting shares this problem: the comparison group must fit the case.
  • Count each piece of evidence once. If your starting estimate already incorporates a symptom, updating on that symptom as though it were new information counts it twice.
  • Check whether evidence overlaps. Two positive tests may share an underlying source of error, so assess whether they provide independent confirmations. A second update must use the probability of that result given what you already know, including the first result. Multiplying ordinary likelihood ratios assumes the relevant conditional independence.
  • Compare your explanation with alternatives. Assess how well your theory predicts an observation relative to its rivals. A rival theory might predict it equally well, or better.
  • Keep uncertainty in the inputs visible. A calculated posterior can look exact even when the prior and test performance are rough estimates. Try a plausible range of inputs before treating the last decimal place as meaningful.

06Where it doesn’t settle the decision

A probability is one input to a decision. A 15% chance may justify an inexpensive follow-up test; the threshold for a risky treatment may be higher. What to do depends on the consequences of acting or waiting, including what happens if you’re wrong.

The theorem requires inputs supplied from outside the formula. Choosing a prior or estimating a likelihood can require measurement and judgment about which assumptions to use. Different reasonable inputs can produce different answers.

The theorem is a mathematical identity you can use regardless of your philosophy of probability. Bayesian updating is the broader practice of using it to revise beliefs, sometimes with models whose assumptions deserve scrutiny. Sound conclusions require a suitable model as well as correct arithmetic.

07Roots

Thomas Bayes, a minister and mathematician in eighteenth-century England, imagined a ball landing at an unknown position on a table. Further balls would land on one side or the other. From those observed outcomes, how could someone infer where the first ball had landed?

The setup turned a familiar probability problem backward. Bayes used results already observed to infer an unknown chance, reversing the usual task of predicting results from a known chance. His argument used a particular assumption about the unknown starting position. The general-purpose formula printed in textbooks today came later.

Bayes died in 1761. His friend Richard Price found the essay among his papers and arranged its publication by the Royal Society in 1763. The work might otherwise have remained a private mathematical exercise.

Pierre-Simon Laplace independently developed a more general approach in 1774. He carried this reasoning from an imagined table into problems of scientific inference. The method became known as inverse probability: working from effects back toward possible causes. Its modern name honors Bayes, while its development owes much to Laplace.

08How solid is this?

ContestedMixedUsefulEstablished

A mathematical identity whenever the relevant probabilities are defined and the evidence has nonzero probability. Applications can still be wrong because the priors, likelihoods, or dependence assumptions are wrong; the theorem doesn’t validate its inputs.

09Connections

counterscounterspart ofincludesincludesincludesincludesincludesBayes’ TheoremBase RateFallacyRepresentativenessHeuristicBayesianUpdatingBayesian PriorNot written yetLikelihoodNot written yetPositivePredictive ValueNot written yetLikelihoodRatioConditionalProbabilityIndependenceNot written yetReference-ClassForecasting

10Origin and sources

Thomas Bayes’s essay was published posthumously in 1763 through Richard Price. Pierre-Simon Laplace independently developed a more general treatment in 1774.

  1. [1]Bayes, T. (1763). An essay towards solving a problem in the doctrine of chances. Philosophical Transactions, 53, 370–418.
  2. [2]Gigerenzer, G., & Hoffrage, U. (1995). How to improve Bayesian reasoning without instruction: Frequency formats. Psychological Review, 102(4), 684–704.
  3. [3]McGrayne, S. B. (2011). The Theory That Would Not Die: How Bayes' Rule Cracked the Enigma Code, Hunted Down Russian Submarines, and Emerged Triumphant from Two Centuries of Controversy. Yale University Press.

Suggest an edit· Updated 2026-10-02