Concept/Probability and Statistics/No. 0086

Bayesian Prior

A Bayesian prior is a probability distribution representing uncertainty before the current evidence is added. In Bayesian statistics, it can reflect earlier data, expert knowledge, or model assumptions. Bayes’ theorem combines it with a likelihood to give a posterior distribution.

Also called Prior Distribution

a concept: name it

01You've seen this when…

  1. in life

    You plan an outdoor birthday party for April. Before checking this week’s forecast, you look at how often it rains there in April.

  2. at work

    A colleague reports that a new checkout doubles purchases. You’ve seen several small tests produce jumps that disappear, so you hold off on rewriting the revenue forecast.

  3. out in the world

    A town offers disease screening to everyone. A resident gets a positive result, and the nurse checks how common the disease is among people with the resident’s risk factors before explaining the result.

02The idea

Before fresh evidence arrives, some possibilities already look more plausible than others. A familiar delivery service probably delivers tomorrow’s package. A sales forecast promising a hundredfold increase needs substantial support. Earlier experience gives these expectations a starting shape.

A Bayesian prior represents that starting uncertainty with probabilities. It can describe a single proposition, such as a 5% chance that a machine has a fault, or an unknown quantity, such as the purchase rate among future customers. For a quantity, the prior spreads probability across possible values. Its center expresses what seems likely; its spread expresses how uncertain that judgment is.

Bayesian inference combines the prior with the likelihood: how probable the observed evidence would be under each possible explanation. The result is a posterior, the updated probability distribution. In compact form:

Posterior ∝ prior × likelihood.

The proportionality sign means the result must be rescaled so its probabilities add up to one. This is the machinery of Bayes’ theorem.

The word before refers to the evidence currently being analyzed. A prior can already contain years of observations. Yesterday’s posterior can become today’s prior when new evidence arrives, provided the model still fits the situation and the same observations aren’t counted twice.

03Why it matters

Small samples leave plenty of room for chance. Two purchases among ten customers could come from a business with a low purchase rate having a lucky day, or one with a higher rate having an ordinary day. The prior helps allocate probability among those explanations.

It also makes assumptions available for inspection. A manager’s forecast already depends on expectations about customers, competitors, and similar launches. Writing those expectations as a distribution lets other people challenge both their direction and their strength.

For noisy estimates, a prior can provide regularization: it pulls estimates toward plausible values and restrains extreme conclusions based on thin evidence. This is especially useful when estimating many rates from small groups, such as sales across dozens of stores.

For rare events, using an appropriate starting probability helps counter the base rate fallacy. A positive screening result needs to be interpreted alongside how common the condition was before testing.

04A worked example

Suppose a team is selling an online course. This is an illustrative example. Related launches suggest that about 10% of qualified visitors might buy, with substantial uncertainty. The team represents that belief using a beta distribution, a distribution for rates between zero and one, with parameters 2 and 18.

For the calculation, these parameters act as starting weights of two purchases and eighteen non-purchases. They are model weights chosen to represent the team’s uncertainty. Ten visitors then arrive, and two buy.

What it looks like The observed purchase rate is 20%. A forecast built directly from that sample would expect roughly twenty purchases per hundred similar visitors.

What’s actually going on The Bayesian calculation adds the two purchases and eight non-purchases to the starting weights. The updated distribution has parameters 4 and 26, giving a posterior mean of 4 ÷ 30, or about 13%. The prior receives substantial weight because the new sample is small.

A different starting distribution changes the result. A uniform prior, represented by parameters 1 and 1, assigns equal density to every rate from zero to one. After the same observations, its parameters become 3 and 9, giving a mean of 25%. Both calculations follow the same updating rule; their starting assumptions differ.

What would have helped Documenting why the earlier launches are comparable, examining the whole range of plausible purchase rates, and doing a sensitivity analysis with alternative priors. These forecasts also assume visitors are comparable and their purchase outcomes are independent given the underlying rate. A burst of traffic from enthusiastic existing customers could undermine that model.

05Where people trip up

  • Turning an opinion into certainty. A prior includes both a location and a spread. Putting nearly all probability around a favored value makes it hard for a small sample to change the conclusion. Choose the spread to reflect how much supporting information exists. Possibilities assigned exactly zero probability remain at zero under ordinary Bayesian updating.
  • Calling a uniform prior neutral. Equal density across purchase rates becomes unequal density when those rates are expressed as odds. There is no universally neutral way to spread probability across every possible representation. Explain which scale the prior uses and why that scale suits the problem.
  • Using the current evidence twice. Building a prior from today’s test results and then updating it with those same results overstates what has been learned. Data-informed methods can be valid, but their fitting procedure must account for how the data were used. Keep a clear record of which observations enter where.
  • Carrying old expectations into a changed setting. A prior based on last year’s customers may fit poorly after a price change or expansion into another country. Check for distribution shift, weaken outdated information, and examine whether the model could plausibly generate the outcomes now appearing.

06Where it doesn’t settle the question

A precise posterior can still rest on a poor model. If the likelihood ignores repeat visitors, measurement errors, or changing demand, careful prior selection alone cannot repair the inference. Check both the starting assumptions and the model’s predictions.

People sometimes expect enough data to wash away every prior. Under suitable conditions, informative data reduce the influence of reasonable priors. That reassurance weakens when data are sparse, several explanations fit equally well, or the model has many poorly measured quantities.

A prior also leaves room for confirmation bias. Someone can choose assumptions that favor the desired answer. Showing alternative priors and the conclusions they produce makes that choice easier to scrutinize.

07Roots

In 1763, Richard Price sent the Royal Society a paper left by his late friend Thomas Bayes, a minister and mathematician. Bayes had examined a backward-looking problem: after observing successes and failures, what could someone infer about the unknown chance of success?

He explored it through a thought experiment involving a ball on a level table. Its unseen position determines the chance that later balls land on one side of it. Observing those later outcomes gives information about the first ball’s position. The setup supplied both a starting distribution and a way to revise it through observations.

Pierre-Simon Laplace independently developed related methods from 1774 onward and extended this reasoning to practical scientific problems. Their work laid foundations for what became Bayesian inference, though modern terminology and practice developed later.

Today, priors range from simple distributions for a single rate to models that share information across hospitals, stores, or regions. The recurring question remains recognizable: how should existing knowledge and fresh observations jointly shape uncertainty?

08How solid is this?

ContestedMixedUsefulEstablished

The role of a prior follows mathematically from Bayes’ theorem. The quality of an applied conclusion depends on the prior, the likelihood, and their fit to the situation; the formula alone cannot establish those assumptions.

09Connections

confused withcounterscounterspart ofpart ofBayesian PriorNot written yetLikelihoodBase RateFallacyConfirmationBiasBayes’ TheoremBayesianUpdatingNot written yetRegularizationSensitivityAnalysisDistributionShiftBurden of Proof

10Origin and sources

Thomas Bayes’s posthumous essay, communicated by Richard Price in 1763, and Pierre-Simon Laplace’s independent work from 1774 laid the foundations. Prior distributions became an explicit component of modern Bayesian inference.

  1. [1]Stigler, S. M. (1986). The History of Statistics: The Measurement of Uncertainty before 1900. Harvard University Press.
  2. [2]Gelman, A., Carlin, J. B., Stern, H. S., Dunson, D. B., Vehtari, A., & Rubin, D. B. (2013). Bayesian Data Analysis (3rd ed.). Chapman and Hall/CRC.
  3. [3]Gelman, A., Simpson, D., & Betancourt, M. (2017). The Prior Can Often Only Be Understood in the Context of the Likelihood. Entropy, 19(10), 555.

Suggest an edit· Updated 2026-10-02