Concept/Probability and Statistics/No. 0600

Measurement Error

Measurement error is the gap between a recorded value and the quantity it aims to measure. Studied in metrology and statistics, it can arise from instruments, reports, or coding. Errors may vary at random, follow a fixed pattern, or depend on other traits.

a concept: name it

01You've seen this when…

  1. in life

    You weigh flour without resetting the kitchen scale after putting the bowl on it. Every recipe looks as though it uses more flour than it does.

  2. at work

    A support dashboard counts a problem as solved when an agent sends a reply. The team reports faster resolution while customers keep writing back.

  3. out in the world

    Two hospitals publish infection rates. One tests every patient; the other tests only patients with symptoms. The first appears to have more infections.

02The idea

A recorded number comes through a process: a sensor, a questionnaire, a person entering data, or software assigning a label. Each step can introduce a gap between the intended quantity and the recorded value.

For a numerical measurement, a simple starting model is recorded value = target value + error. Suppose a reference weight is 72 kilograms and a scale reads 72.8. Its error for that reading is +0.8 kilograms. Usually the target value is less accessible than a reference weight, which makes the error harder to establish.

Three patterns deserve separate attention:

  • Random error makes repeat readings vary. A sensor fluctuates, or someone remembers a slightly different amount each time. Averaging helps when these errors are centered around zero and sufficiently independent.
  • Systematic error shifts readings consistently. A scale has an offset, a clock runs slow, or a coding rule assigns the wrong category. Repeating the same process can reproduce the same mistake.
  • Dependent error changes with other variables. Recall may worsen with time, or a device may work differently across body sizes. These patterns can distort comparisons between groups.

The categories overlap. An instrument can have a fixed offset plus random fluctuations, with both changing across conditions. The useful question is which parts of the measurement process generate which errors.

03Why it matters

Measurement error can change a decision, hide a relationship, or create one. A reading near a treatment threshold can move someone into a different category. A business dashboard can reward teams whose recording practices make their results look better.

In a simple regression, random error in the predictor tends to flatten the estimated relationship when that error is independent of the underlying predictor and outcome. This is regression dilution. Under comparable independence assumptions, random error in the outcome mainly reduces precision. Errors connected to group membership or outcomes can push estimates in either direction.

A larger dataset can make an estimate more precise while preserving its measurement bias. Thousands of readings from the same miscalibrated instrument may tightly cluster around the wrong value. Sample size and measurement quality answer different questions.

That distinction matters whenever small differences drive large consequences: rankings, eligibility rules, medical thresholds, or product experiments.

04A worked example

A 2023 randomized crossover trial measured automated blood pressure in 195 adults using cuffs of different sizes. Measuring the same people with different cuffs allowed researchers to isolate the effect of cuff fit.

What it looks like A standard procedure producing an objective number. A regular-size cuff seems like a reasonable default, particularly when a clinic is busy.

What’s actually going on For participants who needed a large cuff, the regular cuff raised systolic readings by an average of 4.8 mm Hg relative to an appropriately sized cuff. For those who needed an extra-large cuff, the average increase was 19.5 mm Hg. The measurement process introduced an error related to arm size. It could therefore exaggerate apparent blood-pressure differences between people with different-sized arms.

These were mean differences in one trial using an automated device. They do not establish the error for every patient or monitor. An appropriately sized cuff served as the comparison; its reading also has measurement uncertainty.

What would have helped Measuring upper-arm circumference and selecting a cuff within the manufacturer’s specified range. Repeated readings with the same poorly fitting cuff can preserve the distortion. Checking cuff fit addresses its source before clinicians interpret the number.

05Where people trip up

  • Repeated agreement can hide a shared offset. Consistent readings establish repeatability. Accuracy also requires comparison with a suitable reference. Check instruments against known standards, especially after repairs, software changes, or changes in operating conditions.

  • More records can preserve the same bias. Averaging reduces some random errors. It leaves a shared offset intact, and correlated errors provide less benefit from repetition. Ten readings taken through the same faulty process can share the same flaw.

  • A convenient definition can change the quantity. Counting tickets closed is easy; establishing whether a customer’s problem was resolved requires another step. Write down what the target means and how each recorded field represents it. This is operationalization, with construct validity asking how well the measure captures the intended concept.

  • Group comparisons need comparable measurement. Check whether groups face the same devices, questions, thresholds, and opportunities for detection. Different testing intensity can produce detection bias. Where judgment enters the process, blinding can reduce the influence of expectations.

  • Extra digits can imply unsupported precision. A dashboard displaying 7.483% may rest on uncertain labels and incomplete reports. Report detail that the measurement process supports, and describe the main uncertainty alongside the estimate.

  • Corrections need evidence about the errors. A statistical adjustment depends on assumptions about how error behaves. Use repeat measurements, reference comparisons, or a validation sample to test those assumptions. Triangulation helps when methods have different weaknesses; several methods sharing one source of error can agree misleadingly.

Before trusting a consequential comparison, trace a handful of records from the event through collection, coding, and analysis. That small audit often reveals a mismatch the final chart conceals.

06When it isn’t measurement error

A person’s blood pressure can change between readings. A room can warm while a thermometer is being checked. Such differences may reflect changes in the target itself. Specify the person, time, location, and conditions before judging a discrepancy.

Sampling bias concerns who enters the data. Measurement error concerns the value recorded for someone or something included. Both can occur together.

Error and uncertainty also have different meanings. Error is the deviation from a reference or target value, often unknown. Measurement uncertainty describes the range of values reasonably compatible with the available measurement information. An uncertainty estimate expresses the limits of knowledge even after recognized errors have been addressed.

07Roots

In 1904, psychologist Charles Spearman examined a problem facing anyone comparing test scores: unreliable measurements could make a relationship appear weaker. Repeating a test might reshuffle people’s scores, and that instability would carry into the correlation. He developed a correction using information about the measurements’ reliability. It became a foundation of statistical measurement theory.

Spearman inherited a much older problem. Astronomers and surveyors repeatedly observed stars and angles, obtaining slightly different values each time. Their need to combine observations helped drive the development of the statistical theory of observational errors. Metrology, the science of measurement, also had to deal with instruments that shared consistent offsets rather than merely fluctuating between readings.

The idea spread into psychology, economics, medicine, and survey research as measurements increasingly passed through people and coding systems. Modern metrology gave practitioners shared rules for expressing their limits. The international Guide to the Expression of Uncertainty in Measurement, first issued in 1993 and republished in 2008, formalized how uncertainty should be evaluated and reported. Its central practical concern remains familiar: how much confidence can a decision-maker place in the number produced by a measurement process?

08How solid is this?

ContestedMixedUsefulEstablished

Measurement error is a foundational topic in metrology and statistics, with well-established mathematical results and extensive experimental documentation. Its effects depend on the error structure: familiar results such as regression dilution require specific assumptions.

09Connections

confused withcountered bycountered byleads toincludesincludesMeasurementErrorInformationBiasNot written yetTriangulationNot written yetBlindingRegressionto the MeanNot written yetDetection BiasNot written yetRegressionDilutionCorrelationDunning-KrugerEffectSampling BiasNot written yetOperationalization

+ 3 more in the list

10Origin and sources

Developed through metrology and the statistical theory of observational errors, especially in astronomy and surveying. Charles Spearman’s 1904 work established an influential account of how unreliable measurements weaken observed correlations.

  1. [1]Spearman, C. (1904). The Proof and Measurement of Association between Two Things. The American Journal of Psychology, 15(1), 72–101.
  2. [2]Fuller, W. A. (1987). Measurement Error Models. John Wiley & Sons.
  3. [3]Joint Committee for Guides in Metrology. (2008). Evaluation of measurement data — Guide to the expression of uncertainty in measurement. JCGM 100:2008.
  4. [4]Ishigami, J., Charleston, J. B., Miller, E. R., III, Matsushita, K., Appel, L. J., & Brady, T. M. (2023). Effects of Cuff Size on the Accuracy of Blood Pressure Readings: The Cuff(SZ) Randomized Crossover Trial. JAMA Internal Medicine, 183(10), 1061–1068.

Suggest an edit· Updated 2026-10-02