Mammography, by the numbers

Screen 1,000 women who feel completely well, and the arithmetic is settled before a single film is read. About five of them have a cancer the mammogram can catch. The machine will flag four or five of those five — and roughly 110 of the healthy women as well. So about 115 women are told to come back for a closer look, and only about four of them turn out to have cancer. A positive screening mammogram, at the prevalence of an ordinary screening round, is right about 4% of the time. Here is where every one of those numbers comes from.

The three numbers that drive everything

Screening mammography looks for breast cancer in women who have no lump and no symptom — just an age- or risk-based invitation to a population program. Because everyone is presumed healthy walking in, a positive result carries less weight than it would in someone who arrived with a complaint. Three published figures fix the whole calculation, and all three come from the Breast Cancer Surveillance Consortium (BCSC), the largest audited registry of U.S. screening mammography (Lehman et al., 2017).

Sensitivity and specificity are properties of the test and the radiologists reading it; they barely move with the population. Prevalence is entirely about who is being screened. As the difference between sensitivity and specificity plays out below, that third number is the one that decides what a positive result is actually worth.

Working a positive result by hand

Start with 1,000 women and split them by the 0.51% prevalence: about 5 have cancer and about 995 do not. Send both groups through the test.

That leaves roughly 115 positive mammograms: about 4.4 real cancers hiding among about 110 false alarms. The share of positives that are real is the positive predictive value (PPV), and Bayes' theorem gives it exactly, before anyone is rounded to a whole person:

PPV = (0.869 × 0.0051) ÷ [ (0.869 × 0.0051) + (0.111 × 0.9949) ] = 0.00443 ÷ 0.11487 ≈ 0.0386

Round it and a positive screening mammogram is correct about 3.9% of the time — roughly 1 in 26. Put the other way, about 96% of positive screening mammograms are false alarms: not errors, exactly, but healthy women who need a second look to be cleared.

Try it

Open the exact screening round above — 1,000 women, 0.51% prevalence, an 86.9% / 88.9% test. The treatment sliders are set to the contested Cochrane figures discussed below.

Open this scenario in the calculator →

The same round in whole people

Scale the identical rates up to 10,000 women and every cell of the 2×2 lands on close to a whole person, so the fractions above become countable — 51 women with cancer, 9,949 without — while the predictive value stays put.

Screening 10,000 average-risk women — prevalence 0.51%, sensitivity 86.9%, specificity 88.9%. Counts rounded to whole people.
Test positiveTest negativeTotal
Cancer present44 (true positives)7 (false negatives)51
Cancer absent1,104 (false positives)8,845 (true negatives)9,949
Total1,1488,85210,000

Follow the positive column of the confusion matrix: 1,148 women are recalled and 44 have cancer, so PPV = 44 ÷ 1,148 ≈ 3.8% — the same answer the exact rates gave, off by the single case that rounding shuffles between cells. The negative column is the reassuring one. Of 8,852 women cleared, 8,845 are truly cancer-free:

NPV = 8,845 ÷ 8,852 ≈ 0.999 — about 99.9%.

Seven cancers still sit in that negative column — present but not caught this round, the price of an 87% rather than 100% sensitivity, and about one missed cancer for every 1,300 women who screen negative. A normal screening mammogram is a genuinely strong result; a positive one, at this prevalence, is mostly an instruction to look closer.

Change the base rate, change the answer

Nothing about the mammogram improves when it is pointed at a higher-risk group — but the predictive value climbs anyway, because the base rate does the work. Take the same 86.9% / 88.9% test to a group whose pre-test probability is 2% (used here purely to show the mechanism, not as a figure for any named risk category). Screen 1,000 again: about 20 now have cancer and 980 do not. The test flags 20 × 0.869 ≈ 17 of the cancers and 980 × 0.111 ≈ 109 of the healthy, for about 126 positives:

PPV = 17 ÷ 126 ≈ 0.138 — about 1 in 7.

The identical positive result that meant roughly 1-in-26 odds at screening prevalence now means about 1-in-7 — the mammogram unchanged, only the population moved. That single lever, pre-test probability in, post-test probability out, governs every screening result, and is why the same film is read differently in a 40-year-old with no history and a 70-year-old with a strong one.

Try it

The same test at a higher 2% pre-test probability (illustrative only). Watch the positive predictive value climb while sensitivity and specificity stay fixed.

Open the higher-risk variant →

The false alarms add up over a decade

A single false positive is a call-back and, often, a follow-up image or a biopsy that comes back benign. Spread across years of routine screening, those episodes accumulate. Following women through a decade of annual mammography, Elmore and colleagues (NEJM, 1998) estimated that about 49% experience at least one false-positive result, and about 19% are recommended for at least one biopsy that proves benign. The calculator's repeat-testing view plots how that cumulative chance climbs round by round. The simple curve it draws — 1 − specificity raised to the number of rounds, about 69% after ten years here — is an upper bound; the measured 49% is lower because re-reading the same breasts tends to re-flag the same stable quirks year after year.

Benefit and harm on the same population

Everything to this point is well-measured. Sensitivity, specificity, prevalence, and the predictive values that fall out of them come from large audited datasets and are not seriously disputed. The next question — how many deaths screening prevents, and at what cost — is where equally credible experts diverge, and this site does not adjudicate it. Two careful readings sit far apart:

Those figures are not two measurements of one quantity — one counts an invited cohort over a decade, the other a screened cohort over a lifetime — and that mismatch in denominators is much of why the public argument never resolves. The calculator's outcome panel will not settle it either, but it puts a benefit count (from a number-needed-to-screen) beside a harm count (from overdiagnosis) on the same population, so the shape of the trade-off is visible at once. Drop in whichever estimate you find credible and both counts move together.

A positive screen is a question, not an answer

The low predictive value is not a flaw in mammography; it is the reason a diagnostic work-up exists at all. Screening is built to over-call on purpose — to push borderline findings into more imaging, and a biopsy when needed — because the program judges missing the four real cancers worse than briefly recalling the 110 healthy women. Reading a positive screen as a diagnosis rather than a prompt is precisely the base-rate fallacy that runs through every page here, and the downstream costs of that deliberate over-calling — false positives, overdiagnosis, overtreatment — are the subject of screening harms and biases.

References

  1. Lehman CD, Arao RF, Sprague BL, et al. National Performance Benchmarks for Modern Screening Digital Mammography: Update from the Breast Cancer Surveillance Consortium. Radiology, 2017.
  2. Elmore JG, Barton MB, Moceri VM, et al. Ten-Year Risk of False Positive Screening Mammograms and Clinical Breast Examinations. New England Journal of Medicine, 1998.
  3. Gøtzsche PC, Jørgensen KJ. Screening for breast cancer with mammography. Cochrane Database of Systematic Reviews, 2013.
  4. U.S. Preventive Services Task Force. Breast Cancer: Screening (final recommendation). USPSTF, 2024.

Educational model — not medical advice. It illustrates the statistics of testing and treatment; it does not describe any specific real-world test.