Concept/Probability and Statistics/No. 0551
Law of Large Numbers
The law of large numbers is a family of probability theorems stating that, under suitable conditions, a sample average approaches its expected value as the sample grows. First proved by Jacob Bernoulli, it does not imply that short-run imbalances must be corrected.
- Evidence
- Well established
- Read
- 6 min
- Links
- 9 connections
- Useful when
- Evaluating a claim · Forecasting · Reading data and statistics · Risk and safety
01You've seen this when…
- in life
You track your commute. One crash makes Tuesday take 70 minutes, but after several weeks that terrible morning no longer sets the average.
- at work
A factory’s first ten inspected parts include three defects. The manager uses a much larger random sample before treating 30% as the line’s usual defect rate.
- out in the world
A pollster randomly samples 1,500 voters. Compared with a sample of 40, individual surprises have less influence on the reported percentage, though the larger sample still needs to represent the electorate.
02The idea
An early average is easy to push around. One expensive repair dominates your first month’s car costs. One unusually slow ticket dominates a support team’s first hour. As observations accumulate, each new result carries less weight.
The law of large numbers gives this intuition a precise foundation. In its familiar form, observations are independent, come from the same stable distribution, and have a finite expected value. As the sample grows, its average becomes increasingly likely to fall close to that expected value.
Expected value means the probability-weighted average of the possible outcomes. For a fair coin, assign heads a value of 1 and tails a value of 0. The expected value is 0.5, and the sample average is simply the proportion of heads. With enough flips, that proportion is likely to be close to one-half.
The crucial word is proportion. The law says that an imbalance can become small relative to the total, not that heads and tails must soon have equal counts. The absolute imbalance can remain or grow. A run of heads leaves the chance of tails unchanged.
Nor does it promise a smooth journey. An average can move away from its expected value before moving closer again. There is no universal sample size at which uncertainty disappears.
03Why it matters
The law explains why repeatable uncertainty can become manageable at scale. For a warranty provider, the fate of any particular appliance remains uncertain. Across many comparable appliances, it can estimate the average repair cost much more reliably. That estimate supports pricing decisions and helps the provider plan staffing and reserves, provided failures are not dominated by a shared problem.
It also explains why tiny samples are so volatile. A restaurant rated by four customers can jump from excellent to mediocre after one bad review. A restaurant with thousands of ratings barely moves. Treating both averages as equally informative invites the law of small numbers bias: expecting a handful of observations to resemble the whole population.
More data improves precision, and a precise estimate can still be inaccurate. If your survey systematically misses night-shift workers, collecting more responses from daytime workers produces a more stable answer to the wrong question. Sampling bias survives scale.
Averaging leaves the underlying odds unchanged. A casino can use repeated wagers to make its average return more predictable. For a gambler, unfavorable odds remain unfavorable as bets accumulate.
04A worked example
Consider a hypothetical experiment with a fair coin and independent flips. After 100 flips, you have 60 heads and 40 tails. You plan another 900 flips.
What it looks like Heads has had its turn, so the next batch should contain extra tails to restore the balance.
What’s actually going on Each remaining flip still has a 50% chance of heads. The expected number of heads in the next 900 flips is 450. Added to the 60 already observed, that gives an expected total of 510 heads out of 1,000 flips, or 51%.
The original ten heads above the halfway mark remain in expectation. Yet their influence shrinks: they put the first average ten percentage points above 50%, but the expanded average only one point above it. Nothing has compensated for the earlier results. A larger denominator has diluted them. The actual next batch could, of course, differ from its expectation.
What would have helped Separate the next outcome from the running average. Predict the next flip using the coin’s unchanged odds. Estimate the eventual proportion using all the observations, without inventing a force that makes the sequence repay its earlier imbalance.
05Where people trip up
- Treating a streak as a debt. Believing that tails is due after several heads is the gambler’s fallacy. Each independent outcome follows its probabilities regardless of the earlier sequence. Relative frequencies can settle down while the odds stay fixed.
- Expecting the average to improve at every step. A new observation can pull an estimate farther from the expected value. Convergence describes the average’s long-run behavior, during which individual updates can move in either direction.
- Overestimating information from a high row count. Ten thousand readings from one faulty sensor share a source of error. That dependence changes their information content compared with ten thousand independent measurements. Shared causes and repeated observations can make a huge dataset much less informative than its row count suggests.
- Assuming a bigger sample fixes a changing process. If customer behavior changes after a price increase, the old and new orders come from different distributions. More historical data may estimate the past very precisely while obscuring the present.
- Confusing convergence with a bell curve. As the sample grows, the law describes where its average tends. Under additional conditions, the central limit theorem describes the approximate shape and scale of the average’s fluctuations. Those conditions support the familiar rule that four times the sample roughly halves the standard error.
- Calling every return toward normal the same thing. The law of large numbers describes averages over growing samples. A less extreme measurement tends to follow one that noise helped make unusually extreme. This tendency is called regression to the mean. In an independent process, future outcomes follow their probabilities regardless of past results.
06Where it doesn’t promise a useful forecast
Independence and an unchanged distribution are sufficient conditions for the familiar version. Other versions can hold under different conditions. Some dependent processes also obey a law of large numbers. You need a reason to believe that the relevant conditions hold, regardless of how large your spreadsheet is.
A finite expected value also matters. Some mathematical distributions have no finite mean at all. For them, the usual claim that the average settles near a fixed expected value does not apply.
Even when a finite mean exists, fat-tailed distributions can make convergence painfully slow. Rare, enormous losses may dominate years of ordinary results. The theorem offers a long-run guarantee under stated conditions. Your budget or patience may run out before you can benefit.
07Roots
In Basel, Jacob Bernoulli wanted probability to reach beyond the gaming table. With dice, the possible outcomes and their chances could be specified in advance. In civil affairs, commerce, and other practical questions, people often had observations but did not know the underlying chances. Could repeated experience justify a reliable estimate?
Bernoulli used an urn containing differently colored balls to make the problem tractable. If a drawing had a fixed probability of success, how many repeated trials would make the observed proportion likely to lie within a chosen distance of that probability? His achievement was a mathematical bound connecting the number of trials, the desired closeness, and the probability of achieving it. Repetition became something more rigorous than an appeal to experience.
The result appeared in Ars conjectandi in 1713, eight years after his death, in a volume published through his nephew Nicolaus Bernoulli. The argument concerned success-or-failure trials, the setting now associated with coin flips and proportions.
Siméon Denis Poisson introduced the name law of large numbers in 1837. Later mathematicians extended the idea to broader kinds of observations and different forms of convergence. The modern theorem is much wider than Bernoulli’s urn, but it still answers his central question: when does accumulated experience give a dependable average?
08How solid is this?
This is a family of mathematical theorems established by proof, independently of empirical replication. Applying it to real data requires checking assumptions about sampling, dependence, stability, and the existence of a finite mean.
09Connections
- Often confused with Gambler’s Fallacy, Regression to the Mean, Central Limit Theorem, Law of Small Numbers Bias
- See also Expected Value, Independence, Standard Error, Sampling Bias, Fat-Tailed Distributions
10Origin and sources
Jacob Bernoulli proved the foundational result for success-or-failure trials in Ars conjectandi, published posthumously in 1713. Siméon Denis Poisson introduced the name in 1837.
- [1]Bernoulli, J. (1713). Ars conjectandi.
- [2]Feller, W. (1968). An Introduction to Probability Theory and Its Applications. Vol. I (3rd ed.). Wiley.
- [3]Durrett, R. (2019). Probability: Theory and Examples (5th ed.). Cambridge University Press.
Suggest an edit· Updated 2026-10-02