Screening harms and biases
A screening program produces two columns of numbers. One counts the people it helps — cancers caught early, deaths deferred. The other counts the people it burdens — healthy people told they might be sick, disease found that never needed finding, procedures that changed nothing. A fair reading puts both on the same page; this guide walks the harm column and how the calculator quantifies it.
The false positives add up
Sensitivity and specificity describe a single test. But screening repeats that test on the same person year after year, and each round is a fresh chance for a false alarm — so the probability of at least one false positive grows with every round.
If the rounds were independent, the arithmetic compounds the pass rate. A test with 90% specificity clears a healthy person 9 times out of 10, so the chance of clearing them every time across k rounds is 0.9k, and the chance of at least one false alarm is 1 − 0.9k. After 5 rounds that is 1 − 0.95 = 1 − 0.590 = 0.41, about 41%; after 10 rounds it is 1 − 0.910 = 1 − 0.349 = 0.65, roughly 65%. A healthy person can go the whole decade and still, more likely than not, be recalled at least once.
That formula assumes equal specificity, independent errors, and people who remain disease-free. Repeated errors can depend on one another, making the cumulative rate lower or higher than this model. With per-round false-positive risk q, the general bounds are q and min(1, kq), not the independence curve. The repeat-testing view shows the independence result alongside these bounds.
Measured programs land in the same territory. Elmore and colleagues followed women through a decade of annual mammography: 49.1% had at least one false-positive mammogram over 10 rounds, and 18.6% at least one false-positive biopsy recommendation. Croswell and colleagues tracked 14 multimodal PLCO screening tests and found a cumulative false-positive risk of 60.4% in men and 48.8% in women; 28.5% of men and 22.1% of women underwent an invasive procedure prompted by a false alarm.
Try it
Start from the classic teaching scenario — 1,000 people, 1% prevalence, a 90% / 90% test — then open the repeat-testing panel and drag the round count from 1 up to 10.
Open this scenario in the calculator →Overdiagnosis: disease that would never have surfaced
A false positive is a scare that resolves — the follow-up comes back clean. Overdiagnosis is the opposite problem: the disease is real, but it would never have caused symptoms or death in the person's lifetime. Welch and Black define it as the detection of a cancer "that would otherwise not go on to cause symptoms or death" — which requires both a reservoir of silent, indolent disease and a test sensitive enough to find it.
Overdiagnosis is not a testing error; it is a correct diagnosis of something that did not need diagnosing. That is what makes it both hard to see and self-reinforcing: the patient is treated, survives, and is counted as a screening success — inflating the apparent benefit — though the treatment carried only risk and no upside. Welch and Black estimate the overdiagnosed fraction of screen-detected cancers at roughly 25% for mammography, about 50% for lung cancer, and around 60% for PSA-detected prostate cancer, while cautioning that these estimates are themselves uncertain and method-dependent.
Why five-year survival can mislead: lead-time and length-time bias
Two biases make screening look more effective than a mortality count would.
Lead-time bias. Screening moves the moment of diagnosis earlier without necessarily moving death. If a cancer would have surfaced with symptoms at 67 and caused death at 70, catching it at 64 turns a "3-year survivor" into a "6-year survivor" even if the person still dies at 70. Survival measured from diagnosis stretches; the lifespan does not. Welch, Schwartz, and Woloshin showed this directly across tumor types: over four decades the change in 5-year survival had essentially no correlation with the change in mortality (Pearson r = .00), while it did track rising incidence — a sign of broader detection, not better treatment. Five-year survival is a sound way to compare treatments inside a trial, but a poor way to judge whether screening postpones death.
Length-time bias. Screening at fixed intervals preferentially catches slow-growing disease. A fast tumor tends to surface with symptoms between screens; a slow, indolent one waits quietly and is far likelier to still be there at the next scan. So the cancers a screen finds are, on average, the least dangerous — flattering survival statistics before any treatment. Overdiagnosis is length-time bias taken to its limit: disease so slow it never would have mattered.
What a false positive sets in motion
A positive screen is rarely the end of it. It triggers a cascade: a recall for repeat imaging, then often a biopsy or invasive procedure, each with its own rate of complications and cost. Those figures are not just alarms — 18.6% of women reached a biopsy recommendation over 10 mammograms, and in Croswell's multimodal cohort 28.5% of men and 22.1% of women underwent an invasive procedure over 14 tests. Between the positive result and its resolution sits a stretch of anxiety, real even when the answer is benign.
Whole-body and multi-organ screens add a distinct harm: the incidental finding — an unrelated nodule, cyst, or shadow the scan was not looking for. Each one reopens the same follow-up cascade, and most turn out to be nothing, which is precisely the problem.
Reading the ledger: number needed to screen, treat, and harm
The honest way to present a screen puts its benefit and harms on the same denominator. The calculator's outcome panel does this: for a given population it shows how many are helped, harmed, treated with no change, and missed, alongside three summary counts —
- NNT (number needed to treat): how many people must be treated for one to benefit.
- NNH (number needed to harm): how many are treated before one is harmed.
- NNS (number needed to screen): how many must be screened for one person to be helped — the whole chain in one number.
Mammography is the worked case where these numbers are most openly contested. The Cochrane review by Gøtzsche and Jørgensen reads the trials conservatively: for roughly every 2,000 women invited to screening over 10 years, on its reading about 1 avoids dying of breast cancer while about 10 are overdiagnosed and treated for a cancer that would never have harmed them. The U.S. Preventive Services Task Force reaches a more favorable net-benefit judgment, giving biennial screening at ages 40–74 a grade B recommendation; its modeling puts overdiagnosis near 14 per 1,000 women screened over a lifetime (a range of 4 to 37 across models). The two rest on different trials, follow-up windows, and definitions of overdiagnosis — much of why they diverge. Experts genuinely disagree here, and this tool takes no side — read each estimate with its own population, comparator and follow-up. Neither program estimate is a treatment NNT.
Try it
Load mammography's measured accuracy — 87% sensitivity, 89% specificity — at an approximate 0.59% prevalence (detected plus missed cancers). The treatment inputs are generic teaching assumptions, separate from screening-program benefits.
Open this scenario in the calculator →The test column gives about 4.4% PPV at this approximate prevalence. The generic treatment outputs do not estimate the benefits or harms of a mammography program.
The harm column is only half the picture, and it is not this site's claim that harms outweigh benefits — that verdict depends on the numbers you enter, and reasonable numbers point different ways for different tests and ages. For the intuitions this page rests on, see why a rare condition makes most positives false in the base-rate fallacy, and how one result revises risk in pre-test to post-test probability. Both screens argued over most are worked end to end in mammography by the numbers and the PSA test by the numbers; the methodology lays out exactly how the calculator turns these inputs into a helped-and-harmed ledger, and which figures are well-sourced versus weak.
References
- Elmore JG, Barton MB, Moceri VM, et al. Ten-year risk of false positive screening mammograms and clinical breast examinations. New England Journal of Medicine, 1998.
- Croswell JM, Kramer BS, Kreimer AR, et al. Cumulative incidence of false-positive results in repeated, multimodal cancer screening. Annals of Family Medicine, 2009.
- Welch HG, Black WC. Overdiagnosis in cancer. Journal of the National Cancer Institute, 2010.
- Welch HG, Schwartz LM, Woloshin S. Are increasing 5-year survival rates evidence of success against cancer? JAMA, 2000.
- Gøtzsche PC, Jørgensen KJ. Screening for breast cancer with mammography. Cochrane Database of Systematic Reviews, 2013.
- U.S. Preventive Services Task Force. Breast Cancer: Screening. USPSTF Recommendation Statement, 2024.