The base-rate fallacy in medical testing
In 1978, researchers put a short problem to physicians, residents, and medical students at Harvard teaching hospitals. A disease affects 1 in 1,000 people. Its test produces a false positive about 5% of the time. Someone picked at random, with no symptoms to go on, tests positive. What is the chance they actually have the disease? The most common answer was 95%. The correct answer is about 2%.
That gap — a factor of nearly fifty — is not a lapse in arithmetic by people who cannot do arithmetic. These were clinicians who order tests like this every day. It is a lapse in which fact the mind reaches for. The 5% describes how often the test misfires. The question asks how often a positive patient is truly sick. Those are different quantities, and the fact that closes the distance between them is the one the problem quietly buries: the disease is rare.
Neglecting how common a condition is when you read a positive result is the base-rate fallacy. The base rate — the prevalence, the share of the tested group who truly have the condition before any test is run — carries most of the weight in what a positive result means. Leave it out and a positive test looks far more damning than it is.
Two questions that are not the same question
Look at what the fallacy swaps. The problem hands you the chance of a positive test given disease and asks for the chance of disease given a positive test. Those are inverse conditional probabilities, and reversing them is so tempting it has a name: the confusion of the inverse. One is a property of the test — sensitivity, the share of sick people it flags — and it holds steady whether a condition is common or rare. The other is the predictive value — the share of flagged people who are actually sick — and it slides up and down with prevalence. How those first two numbers are defined, and why they belong to the test rather than to you, is the subject of sensitivity vs. specificity.
Training does not dissolve the reflex. In a 1982 survey of clinical reasoning, physicians were told that breast cancer runs about 1% in the women being screened, that a mammogram is positive in roughly 80% of women who have cancer and in about 9.6% of women who do not. Asked what a single positive mammogram implied, about 95 of 100 physicians put the chance of cancer near 75%. Worked through properly, the figure is under 8% — off by roughly a factor of ten, and in the frightening direction. Two decades later, when researchers handed 48 physicians realistic diagnostic problems phrased in that same probability language, only about one in ten reached the right predictive value.
Natural frequencies: the format that repairs it
The striking part is how cheaply the error undoes itself. Take the identical facts and state them as counts of whole people instead of percentages and conditional probabilities, and most of the confusion lifts. Among those same 48 physicians, rewording the problems as natural frequencies raised the share who answered correctly from about 10% to 46%. No probability changed; only its dress. Whole people are the format the mind handles well, and counting them keeps the base rate in view instead of leaving it to be remembered at the one moment it matters.
The method is a fixed recipe, and it runs on any test:
- Start with a round group of people — 1,000 is convenient.
- Split them by the base rate into those who have the condition and those who do not.
- Apply the sensitivity to the sick group to get the true positives.
- Apply the false-positive rate — 1 minus specificity — to the healthy group to get the false positives.
- Read the predictive value straight off: true positives divided by everyone who tested positive.
Run it on the scenario this calculator opens with — a 1% base rate and a test that is 90% sensitive and 90% specific, the same numbers behind its own worked example. Among 1,000 people, 10 have the condition and the test catches 9 of them. Among the 990 who are healthy, a 90%-specific test still misfires on one in ten — 99 false alarms. Ninety-nine false positives now sit beside nine true ones, so 108 people carry a positive result and only nine are sick: a predictive value of about 8%, meaning more than nine of every ten positives are false. The false alarms swamp the true cases for one reason — they are drawn from a healthy crowd a hundred times larger than the handful who are ill.
Try it
Watch the recipe run live: 1,000 people, a 1% base rate, a 90% / 90% test — and a positive result only about 8% likely to be real.
Open this scenario in the calculator →Hold the test still and move the base rate
The clearest way to see how much of the answer belongs to prevalence is to freeze the test and change nothing but how common the disease is. Keep the same 90% / 90% test and start each row from 10,000 people so the counts stay whole:
| Base rate | Have it | Test flags it (true +) | False alarms | All positives | Positive that is real (PPV) |
|---|---|---|---|---|---|
| 0.1% | 10 | 9 | 999 | 1,008 | ≈ 0.9% |
| 1% | 100 | 90 | 990 | 1,080 | ≈ 8.3% |
| 5% | 500 | 450 | 950 | 1,400 | ≈ 32% |
| 10% | 1,000 | 900 | 900 | 1,800 | 50% |
One test, four verdicts. At a 0.1% base rate it flags more than a thousand people to find nine, and a positive means real disease less than one time in a hundred. By 10% the sick group has grown enough that true positives finally match false alarms one for one, and a positive is a coin flip. The test never changed; the base rate did all the moving. The calculator draws this as a continuous prevalence-to-PPV curve — slide prevalence and watch predictive value climb from near zero toward certainty while sensitivity and specificity sit still.
Try it
Nudge the same 90% / 90% test up to a 10% base rate and let prevalence do the work — the chance a positive is real jumps to 50%.
Open the 10% scenario →Where this lands in the clinic
Population screening lives in exactly the low-prevalence corner where the fallacy does its worst work, because everyone invited is presumed healthy to begin with. Screen-detected breast cancer runs about 0.51% in the pooled Breast Cancer Surveillance Consortium data, and a screening mammogram is roughly 86.9% sensitive and 88.9% specific. Put those through the recipe. Picture 100,000 women screened: about 510 have a cancer a mammogram can find, and 99,490 do not. The test catches about 443 of the 510. But its specificity leaves an 11.1% false-positive rate, which flags roughly 11,043 of the healthy women. Now 11,486 women hold a positive mammogram and 443 of them have cancer — about 1 in 26, or 3.9%. A test that scores close to 90% on both axes is still a false alarm in the large majority of the people it flags, purely because cancer is rare in the screened group. The full profile of that test, benefits and harms alongside the math, is on mammography, by the numbers.
Try it
Load the screening-mammography operating point — 0.51% prevalence, an 86.9% / 88.9% test — and read the positive predictive value straight off the results.
Open the mammography scenario →What the fix settles, and what it doesn't
Reframing a problem as whole people repairs the arithmetic of a single positive result. It stops there. It says nothing about where to set a test's threshold — whether to buy sensitivity at the cost of specificity is a separate decision, again in sensitivity vs. specificity. It is one step short of the general machinery for turning any starting probability into a post-test probability, and for chaining more than one result together; that runs on likelihood ratios and the odds form of Bayes' rule, in from pre-test to post-test probability. And a low predictive value is not a useless test: whether a positive that is usually a false alarm is worth chasing depends on what the follow-up costs and what catching a true case is worth, which is the ground of screening harms and biases.
What the base rate buys is the right question. Not "how accurate is this test?" but "of the people it flags, how many are actually sick?" — and the Harvard survey is the standing reminder that those two numbers can sit a factor of fifty apart. Set the counts out as whole people, and the trap mostly disappears.
References
- Casscells W, Schoenberger A, Graboys TB. Interpretation by physicians of clinical laboratory results. N Engl J Med. 1978;299(18):999–1001.
- Eddy DM. Probabilistic reasoning in clinical medicine: problems and opportunities. In: Kahneman D, Slovic P, Tversky A, eds. Judgment under Uncertainty: Heuristics and Biases. Cambridge University Press; 1982:249–267.
- Gigerenzer G, Hoffrage U. How to improve Bayesian reasoning without instruction: frequency formats. Psychol Rev. 1995;102(4):684–704.
- Hoffrage U, Gigerenzer G. Using natural frequencies to improve diagnostic inferences. Acad Med. 1998;73(5):538–540.
- Lehman CD, Arao RF, Sprague BL, et al. National performance benchmarks for modern screening digital mammography (BCSC). Radiology. 2017;283(1):49–58.