Screening harms and biases
A screening program produces two columns of numbers. One counts the people it helps — cancers caught early, deaths deferred. The other counts the people it burdens — healthy people told they might be sick, disease found that never needed finding, procedures that changed nothing. A fair reading puts both on the same page; this guide walks the harm column and how the calculator quantifies it.
The false positives add up
Sensitivity and specificity describe a single test. But screening repeats that test on the same person year after year, and each round is a fresh chance for a false alarm — so the probability of at least one false positive grows with every round.
If the rounds were independent, the arithmetic compounds the pass rate. A test with 90% specificity clears a healthy person 9 times out of 10, so the chance of clearing them every time across k rounds is 0.9k, and the chance of at least one false alarm is 1 − 0.9k. After 5 rounds that is 1 − 0.95 = 1 − 0.590 = 0.41, about 41%; after 10 rounds it is 1 − 0.910 = 1 − 0.349 = 0.65, roughly 65%. A healthy person can go the whole decade and still, more likely than not, be recalled at least once.
That formula, P(at least one false positive) = 1 − specificityk, is an idealization: it assumes every round is independent. Real repeat testing is not. Re-imaging the same person tends to re-flag the same benign quirk, so the errors correlate and the true rate usually runs lower than the formula. Pushing the other way, a single "round" that scans many organs at once stacks several independent chances into one visit, which can drive it higher. The calculator draws this curve, labels it an upper bound, and lets you move the round count in the repeat-testing view.
Measured programs land in the same territory. Elmore and colleagues followed women through a decade of annual mammography: 49.1% had at least one false-positive mammogram over 10 rounds, and 18.6% at least one false-positive biopsy recommendation. Croswell and colleagues tracked 14 rounds of multimodal PLCO screening and found a cumulative false-positive risk of 60.4% in men and 48.8% in women; 28.5% of men and 22.1% of women underwent an invasive procedure prompted by a false alarm.
Try it
Start from the classic teaching scenario — 1,000 people, 1% prevalence, a 90% / 90% test — then open the repeat-testing panel and drag the round count from 1 up to 10.
Open this scenario in the calculator →Overdiagnosis: disease that would never have surfaced
A false positive is a scare that resolves — the follow-up comes back clean. Overdiagnosis is the opposite problem: the disease is real, but it would never have caused symptoms or death in the person's lifetime. Welch and Black define it as the detection of a cancer "that would otherwise not go on to cause symptoms or death" — which requires both a reservoir of silent, indolent disease and a test sensitive enough to find it.
Overdiagnosis is not a testing error; it is a correct diagnosis of something that did not need diagnosing. That is what makes it both hard to see and self-reinforcing: the patient is treated, survives, and is counted as a screening success — inflating the apparent benefit — though the treatment carried only risk and no upside. Welch and Black estimate the overdiagnosed fraction of screen-detected cancers at roughly 25% for mammography, about 50% for lung cancer, and around 60% for PSA-detected prostate cancer, while cautioning that these estimates are themselves uncertain and method-dependent.
Why five-year survival can mislead: lead-time and length-time bias
Two biases make screening look more effective than a mortality count would.
Lead-time bias. Screening moves the moment of diagnosis earlier without necessarily moving death. If a cancer would have surfaced with symptoms at 67 and caused death at 70, catching it at 64 turns a "3-year survivor" into a "6-year survivor" even if the person still dies at 70. Survival measured from diagnosis stretches; the lifespan does not. Welch, Schwartz, and Woloshin showed this directly across tumor types: over four decades the change in 5-year survival had essentially no correlation with the change in mortality (Pearson r = .00), while it did track rising incidence — a sign of broader detection, not better treatment. Five-year survival is a sound way to compare treatments inside a trial, but a poor way to judge whether screening postpones death.
Length-time bias. Screening at fixed intervals preferentially catches slow-growing disease. A fast tumor tends to surface with symptoms between screens; a slow, indolent one waits quietly and is far likelier to still be there at the next scan. So the cancers a screen finds are, on average, the least dangerous — flattering survival statistics before any treatment. Overdiagnosis is length-time bias taken to its limit: disease so slow it never would have mattered.
What a false positive sets in motion
A positive screen is rarely the end of it. It triggers a cascade: a recall for repeat imaging, then often a biopsy or invasive procedure, each with its own rate of complications and cost. Those figures are not just alarms — 18.6% of women reached a biopsy recommendation over 10 mammograms, and in Croswell's multimodal cohort 28.5% of men and 22.1% of women underwent an invasive procedure over 14 rounds. Between the positive result and its resolution sits a stretch of anxiety, real even when the answer is benign.
Whole-body and multi-organ screens add a distinct harm: the incidental finding — an unrelated nodule, cyst, or shadow the scan was not looking for. Each one reopens the same follow-up cascade, and most turn out to be nothing, which is precisely the problem.
Reading the ledger: number needed to screen, treat, and harm
The honest way to present a screen puts its benefit and harms on the same denominator. The calculator's outcome panel does this: for a given population it shows how many are helped, harmed, treated with no change, and missed, alongside three summary counts —
- NNT (number needed to treat): how many people must be treated for one to benefit.
- NNH (number needed to harm): how many are treated before one is harmed.
- NNS (number needed to screen): how many must be screened for one person to be helped — the whole chain in one number.
Mammography is the worked case where these numbers are most openly contested. The Cochrane review by Gøtzsche and Jørgensen reads the trials conservatively: for roughly every 2,000 women invited to screening over 10 years, on its reading about 1 avoids dying of breast cancer while about 10 are overdiagnosed and treated for a cancer that would never have harmed them. The U.S. Preventive Services Task Force reaches a more favorable net-benefit judgment, giving biennial screening at ages 40–74 a grade B recommendation; its modeling puts overdiagnosis near 14 per 1,000 women screened over a lifetime (a range of 4 to 37 across models). The two rest on different trials, follow-up windows, and definitions of overdiagnosis — much of why they diverge. Experts genuinely disagree here, and this tool takes no side — enter either set of figures and read off the ledger it produces.
Try it
Load mammography's measured accuracy — 87% sensitivity, 89% specificity — at a screen-detected prevalence near 0.5%, with the contested number-needed-to-screen and overdiagnosis figures pre-filled, and read the positive predictive value and the harm column together.
Open this scenario in the calculator →Those last two inputs are number-needed-to-screen and overdiagnosis proxies rather than per-treatment rates, so read the outcome counts as illustrative. The robust lesson is the test column: at 0.5% prevalence the positive predictive value falls below 4%, so most positives are false alarms.
The harm column is only half the picture, and it is not this site's claim that harms outweigh benefits — that verdict depends on the numbers you enter, and reasonable numbers point different ways for different tests and ages. For the intuitions this page rests on, see why a rare condition makes most positives false in the base-rate fallacy, and how one result revises risk in pre-test to post-test probability. Both screens argued over most are worked end to end in mammography by the numbers and the PSA test by the numbers; the methodology lays out exactly how the calculator turns these inputs into a helped-and-harmed ledger, and which figures are well-sourced versus weak.
References
- Elmore JG, Barton MB, Moceri VM, et al. Ten-year risk of false positive screening mammograms and clinical breast examinations. New England Journal of Medicine, 1998.
- Croswell JM, Kramer BS, Kreimer AR, et al. Cumulative incidence of false-positive results in repeated, multimodal cancer screening. Annals of Family Medicine, 2009.
- Welch HG, Black WC. Overdiagnosis in cancer. Journal of the National Cancer Institute, 2010.
- Welch HG, Schwartz LM, Woloshin S. Are increasing 5-year survival rates evidence of success against cancer? JAMA, 2000.
- Gøtzsche PC, Jørgensen KJ. Screening for breast cancer with mammography. Cochrane Database of Systematic Reviews, 2013.
- U.S. Preventive Services Task Force. Breast Cancer: Screening. USPSTF Recommendation Statement, 2024.