Tool/Meta-Concept/No. 0125
Calibration
Calibration, or probability calibration, is agreement between stated probabilities and observed results over many forecasts. In forecasting and judgment research, events assigned a 70% chance occur about 70% of the time. Calibration alone does not show that forecasts are informative.
Also called Probability Calibration
- Evidence
- Well established
- Read
- 6 min
- Links
- 19 connections
01You've seen this when…
- in life
You give yourself a 90% chance of making the 7:10 train. Your calendar shows you’ve missed it on four of the last ten mornings.
- at work
A sales team labels deals likely and builds its hiring plan around them. A review shows that likely means very different odds depending on which salesperson enters it.
- out in the world
Election forecasters give several races a 70% chance of going to one party. After voting ends, an analyst checks the whole group, not just the race that produced a surprise.
02The idea
Confidence becomes useful when it has a dependable meaning. If your 80% predictions succeed about eight times in ten, someone can use that number to plan. If they succeed only half the time, your confidence is overstating what you know.
Calibration is the agreement between stated probabilities and observed results across many judgments. Among events assigned a 70% chance, roughly 70% should happen. Among answers given 90% confidence, roughly 90% should be correct. The same principle applies to low probabilities: events assigned a 10% chance should sometimes happen, not never.
You cannot establish calibration from one outcome. A 90% forecast can fail without being a bad forecast. A 10% forecast can come true without being a brilliant one. The test is the pattern across a sufficiently large, relevant set of judgments.
Calibration measures the match between probabilities and results. Knowledge and the ability to make useful distinctions are separate qualities. Someone who predicts a town’s long-run rain rate every morning might be calibrated while missing every approaching storm. Useful forecasting needs both honest probabilities and information that distinguishes one situation from another.
The practical aim is to match confidence to what your judgments deserve. That calls for selective adjustments: sometimes lowering confidence, sometimes recognizing that your supposedly tentative judgments are usually right.
03How to use it
- Make the outcome judgeable. Write a yes-or-no claim with a deadline and a rule for resolving it. For a delivery, distinguish arriving at the warehouse from being available for use. Set the rule before seeing the result.
- Start with comparable cases. Check how often this outcome has happened in similar circumstances. Reference-class forecasting gives you a starting probability; case-specific evidence can move it.
- Record the probability before the outcome. Put a number beside the claim in a decision journal. Preserve the original forecast. If you update it, save each version and record when and why it changed.
- Review the whole eligible set. Include dull predictions and all failures, including embarrassing misses. Use a consistent cutoff, such as the forecast made one week before each deadline. Keep predictions that turned out badly in the review.
- Compare probabilities with frequencies. Group nearby estimates, such as 60–69%. Compare their average stated probability with the fraction of events that happened. Record the number of cases too: six successes out of ten tell you much less than sixty out of a hundred.
- Adjust, then test again. If a recurring mismatch appears, examine the reasoning behind it. Revise future estimates and check a fresh batch. Keep the old data intact, even when they make your record look poorly calibrated.
A Brier score can summarize forecast performance alongside this check. It rewards probabilities close to outcomes, but it measures more than calibration alone.
04A worked example
In this illustrative example, a purchasing manager records forecasts for 50 supplier deliveries. Each receives an 80% chance of arriving by its agreed date. She defines arrival as the warehouse accepting the shipment before closing time. Only 30 arrive on time.
What it looks like Faced with a run of supplier problems, the manager continues planning around 80% reliability because she treats several late shipments with plausible explanations as exceptions.
What’s actually going on The forecasts collectively implied about 40 on-time deliveries. There were 30, or 60%. That mismatch remains even with plausible explanations for individual delays. It suggests either that the manager’s estimates are too optimistic or that conditions have changed. The observed 60% leaves uncertainty about the supplier’s true future rate, especially if several deliveries share the same disruption.
What made it work The manager had saved probabilities and resolution rules before results arrived. She lowers her working estimate for comparable shipments and investigates whether rush orders need a separate estimate. She leaves more room for lateness in project schedules, then records the next batch. The improvement comes from making confidence accountable to evidence and checking whether the adjustment from 80 to 60 holds up.
05When to reach for it
06When it misleads
- A small sample looks like a verdict. Eight successes from ten 80% predictions are compatible with calibration; establishing it requires more evidence. Ten updates to the same event share one outcome, so they provide dependent tests.
- Averages hide different weaknesses. Someone can be overconfident about hiring and underconfident about technical estimates, with the errors canceling overall. Examine meaningful categories, but avoid slicing the data until every tiny group tells a convenient story. That’s overfitting.
- Matching frequencies becomes the entire goal. Predicting the base rate every time can match outcome frequencies while providing little help with individual cases. Check whether forecasts distinguish more likely cases from less likely ones, and whether they improve on a simple baseline. Proper scoring rules help evaluate overall probability quality.
- The world changes beneath the record. A supplier, market, or organization may behave differently next year. Past calibration provides evidence that needs reassessment over time. New evidence still requires Bayesian updating.
- Unresolvable claims enter the exercise. A claim such as whether a policy was best needs a defined goal and alternatives before its outcome can be agreed on. Calibration can test predictions that bear on that judgment. Choosing the values behind it is a separate task.
07Roots
At the U.S. Weather Bureau, meteorologist Glenn Brier worked on a problem that ordinary right-or-wrong scoring handled poorly: judging forecasts expressed as probabilities. A chance of rain leaves room for both a wet afternoon and a dry one. Simply counting correct predictions throws that information away. In 1950, Brier published a scoring rule that compared the probabilities assigned to possible outcomes with what occurred. It became the Brier score, an important foundation for evaluating probabilistic forecasts. Its scope extends beyond calibration.
The question then moved from weather forecasts to people’s judgments. In the 1970s, psychologist Sarah Lichtenstein and colleagues, including Baruch Fischhoff, studied how confidence matched correctness on tasks such as general-knowledge questions. Participants supplied answers and estimates of how likely those answers were to be right. That extra number exposed a gap an ordinary test score leaves hidden: knowing the answer and knowing how much to trust yourself are different abilities.
Their work also examined whether feedback and training could improve the match. It helped make calibration a practical learning target in addition to a statistical description. Calibration’s origins span multiple fields. Weather verification and psychological research developed different ways to compare stated probabilities with outcomes; later forecasting practice continued that work. Each pursued the same question: does this person’s 70% actually behave like 70%?
08How solid is this?
Calibration is a well-defined, established property of probability judgments. Research shows that feedback and training can improve it, though gains do not transfer reliably across every task. Calibration alone does not establish that forecasts are informative.
09Connections
- Helps counterIllusion of Validity, Overconfidence Effect, Hot Hand Fallacy, Illusory Superiority, Optimism Bias, Outcome Bias
- Can follow fromProper Scoring Rule, Decision Journal, Intellectual Humility
- Part of Metacognition
- Includes Dunning-Kruger Effect
- See also Bayesian Updating, Reference-Class Forecasting, Brier Score, Overfitting, Circle of Competence, Deliberate Practice, Hindsight Bias, Wisdom of Crowds
+ 9 more in the list
10Origin and sources
Developed across probability forecasting and judgment research, with important contributions from Glenn Brier (1950) and Sarah Lichtenstein, Baruch Fischhoff, and colleagues in the 1970s and 1980s.
- [1]Brier, G. W. (1950). Verification of forecasts expressed in terms of probability. Monthly Weather Review, 78(1), 1–3.
- [2]Lichtenstein, S., & Fischhoff, B. (1977). Do those who know more also know more about how much they know? Organizational Behavior and Human Performance, 20(2), 159–183.
- [3]Lichtenstein, S., & Fischhoff, B. (1980). Training for calibration. Organizational Behavior and Human Performance, 26(2), 149–171.
- [4]Gneiting, T., Balabdaoui, F., & Raftery, A. E. (2007). Probabilistic forecasts, calibration and sharpness. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 69(2), 243–268.
Suggest an edit· Updated 2026-10-02