Concept/Probability and Statistics/No. 0187
Conditional Probability
Conditional probability is the chance of an event given that another event has occurred or information is known. In classical probability theory, it is found by dividing the probability of both events by the probability of the given event, when that probability is positive.
- Evidence
- Well established
- Read
- 6 min
- Links
- 12 connections
01You've seen this when…
- in life
You usually allow twenty minutes for the drive to the station. Today it’s raining hard, and you reconsider the chance of catching your train.
- at work
Most support tickets close within a day. A customer asks about one already open for three days, and the team replies with the overall figure.
- out in the world
A city introduces a screening program. Its leaflet lists how often the test detects the condition, while residents want to know how often a positive result means they have it.
02The idea
A probability always rests on some information. Conditional probability makes that information explicit. The chance of missing a train given heavy rain can differ from the chance across all journeys. The chance that a ticket closes tomorrow can change once it has already remained open for three days.
Write this as P(A | B): the probability of event A given event B. The vertical bar means given. For events where B has a positive probability, the rule is:
P(A | B) = P(A and B) / P(B)
Start with the cases where B happens. Within that group, find the share where A also happens. The denominator sets the group; the numerator counts the cases that meet both conditions.
Suppose a jar contains twenty marbles. Eight are blue, and two of those blue marbles are large. Given that a drawn marble is blue, its chance of being large is two out of eight, or 25%. The other twelve marbles fall outside the group specified by the condition.
Information sometimes leaves the probability unchanged. A fair coin’s next toss still has a 50% chance of heads after an earlier tail, assuming the tosses are independent. Conditioning changes the question’s reference group; its numerical effect depends on the relationship between the events.
03Why it matters
Decisions usually arrive with information already attached. A patient has symptoms. A shipment is overdue. A borrower has missed a payment. A project has passed its first milestone. Each detail can change which historical cases provide a useful comparison.
An overall average can therefore give a poor answer to a specific decision. Suppose most deliveries arrive on schedule. Once a package has missed its delivery date, estimating its remaining journey calls for information about packages that reached that same situation.
This connects conditional probability to reference-class forecasting: choose past cases that share the details relevant to the outcome. Adding every available detail can leave too few cases for a reliable estimate, so the task also requires judgment about which conditions matter.
It also explains why two apparently conflicting percentages can both be correct. They may describe different groups.
04A worked example
Consider an invented screening test used in a population where 1% of people have a condition. The test returns a positive result for 90% of people who have it. It also returns a positive result for 5% of people who do not. Assume those rates apply to the people being screened.
What it looks like A positive result seems to carry a 90% chance of having the condition. That number is prominent in the test description, and it sounds like the answer someone awaiting a diagnosis needs.
What’s actually going on Imagine screening 10,000 people. Of the 100 people with the condition, 90 test positive. Of the 9,900 people without it, 495 test positive. That gives 585 positive results altogether.
Among those 585 people, 90 have the condition. The probability of having it given a positive result is therefore 90 / 585, or about 15%. This is the test’s positive predictive value in this population.
The 90% figure answers P(positive result | condition). The 15% figure answers P(condition | positive result). Both are conditional probabilities, with different reference groups. Because the condition is uncommon, even a 5% false-positive rate produces many positive results among people without it.
What would have helped Show the expected counts alongside the test rates, and label which group each percentage describes. For an actual medical decision, use estimates appropriate to the patient and the screening population; these invented figures describe no particular test.
05Where people trip up
- Reverse the condition. The probability of a positive test among people with a condition differs from the probability of the condition among people with a positive test. Write the two groups in words before using either percentage. Bayes’ theorem provides a rule for connecting these directions when the necessary probabilities are available.
- Lose the underlying frequency. A rare event starts with relatively few cases. False alarms or misleading clues can outnumber accurate ones even when each is individually uncommon. Ignoring that starting frequency is the base rate fallacy.
- Use a convenient group. A website poll estimates responses among people who chose to answer it. Applying that percentage to every customer requires evidence that respondents represent the wider group. Sampling bias changes the population on which the estimate rests.
- Overlook the information carried by survival. A ticket still open after three days belongs to the set that remained unresolved that long. Likewise, a fund still trading belongs to the set that survived. Excluding closed tickets or failed funds changes the comparison group; in performance comparisons, this can produce survivorship bias.
A useful habit is to finish the phrase among which cases? Whenever a percentage appears, identify the people, events or periods included in its denominator. Then check whether that group matches the question being answered.
06Where it doesn’t establish a cause
A conditional probability describes how often an outcome occurs within a specified group. A causal claim asks what would happen if someone changed the situation.
Suppose students receiving tutoring have lower pass rates than students without tutoring. Students who are already struggling may be more likely to receive help. Their starting difficulties can explain the observed difference through confounding.
The comparison alone doesn’t establish whether tutoring helps or harms. A randomized experiment, or another credible causal design, can address that question. Conditional probabilities still describe the observed groups accurately when the underlying data are sound; assigning a causal meaning requires additional evidence.
07Roots
In 1763, Richard Price presented a paper left by his late friend Thomas Bayes to the Royal Society in London. Bayes’s thought experiment involved balls tossed onto a flat table. The first ball’s position was hidden; later throws supplied evidence by landing to its left or right. The problem was to infer something unknown from observed results.
Conditional reasoning already belonged to the mathematics of games of chance. A card drawn from a deck changes the possibilities for the next draw. Bayes helped develop the reverse problem: given observed outcomes, what could be inferred about the chance that produced them? Pierre-Simon Laplace later developed this approach further, making conditional reasoning central to inference from evidence. The familiar formula associated with Bayes connects a conditional probability in one direction to the probability in the other.
In 1933, Andrey Kolmogorov published a slim monograph that placed probability on a unified system of mathematical axioms. Conditional probability took its place within that framework, alongside rules for combining events and handling independence. Its everyday calculation remains accessible: specify what is known, restrict attention to the corresponding cases, and calculate the share containing the outcome of interest.
08How solid is this?
Conditional probability is a mathematical definition with established rules. Applying it to a practical question depends on reliable data, an appropriate comparison group and justified assumptions about how the data were generated.
09Connections
- Often confused with Conjunction Fallacy
- Part of Bayes’ Theorem
- Includes Bayesian Updating, Independence, Positive Predictive Value
- See also Base Rate Fallacy, Confounding, Sampling Bias, Survivorship Bias, Reference-Class Forecasting, Randomized Experiment, Simpson’s Paradox
+ 2 more in the list
10Origin and sources
Developed within classical probability theory for reasoning about related events. Bayes’s posthumous essay (1763) advanced inverse probability; Kolmogorov’s axiomatic framework (1933) supplied its modern mathematical foundation.
- [1]Bayes, T. (1763). An Essay towards solving a Problem in the Doctrine of Chances. Philosophical Transactions, 53, 370–418.
- [2]Kolmogorov, A. N. (1950). Foundations of the Theory of Probability. Translated by Nathan Morrison. Chelsea Publishing Company.
- [3]Feller, W. (1968). An Introduction to Probability Theory and Its Applications. Volume I, third edition. John Wiley & Sons.
Suggest an edit· Updated 2026-10-02