Repeated screening and the arithmetic of false alarms
Every accuracy number on this site describes a single round of testing. Almost no screening program consists of one. Mammography is invited every one to two years for decades; FIT is annual; PSA, low-dose CT, cervical screening — all recur. And repetition quietly transforms the arithmetic: a specificity that sounds reassuring per round compounds into a false-alarm rate per screening career that surprises nearly everyone, including the clinicians doing the inviting. This guide works that compounding in both directions — because the same mechanism that piles up false alarms is also, run the other way, how confirmation works.
The compounding, idealized
Take a test with 90% specificity: each round, a disease-free person has a 90% chance of being correctly cleared. Ask instead: what is the chance of getting through n rounds without a single false alarm? If each round erred independently, it would be 0.9 × 0.9 × … — that is 0.9ⁿ — and the chance of at least one false positive would be
P(≥1 false alarm in n rounds) = 1 − specificityⁿ.
The numbers escalate faster than intuition expects. At 90% specificity: 10% after one round, 27% after three, 41% after five, 65% after ten. Even an excellent 99%-specific test reaches 10% after ten rounds. The per-round error never changed; exposure did. The same person keeps re-rolling the same dice, and screening programs are the re-rolling of dice, decade after decade, on millions of healthy people. The calculator's serial-testing panel draws this curve live for whatever specificity you set — drag the slider and watch the decade-long picture swing.
Try it
Start from the classic 90% / 90% test and open the serial-testing curve: 65% of disease-free people flagged at least once by round ten — from a test that is "90% accurate."
Open the serial-testing view →What real programs actually observe
The formula assumes each round errs independently, and that assumption deserves its flag: it is an idealization, and real screening rounds are correlated. The same person brings the same dense breast tissue or the same benign nodule to every round; radiologists compare against prior films and recall less often once a finding is known to be stable. So the naive curve is an upper bound, and the honest question is what repetition costs in practice. Two landmark studies answer with measured, not modeled, numbers:
- Mammography: Elmore and colleagues followed 2,400 women over ten years in the New England Journal of Medicine. After ten screening mammograms, the cumulative probability of at least one false positive was 49.1% — and about one in five of the women with a false alarm underwent a biopsy for it. (Naive formula at their per-round rates: ~69%.)
- Multi-organ screening: Croswell and colleagues analyzed the PLCO trial, whose participants received up to 14 screening tests across four cancers over three years. The cumulative risk of at least one false positive reached 60.4% for men and 48.8% for women — and roughly a third of the men and a quarter of the women with false positives had an invasive follow-up procedure because of one.
Both observed figures land below the independence bound, exactly as the correlation argument predicts — and both are still enormous. Measured or modeled, the conclusion survives: for a disease-free person, a screening career is closer to a coin flip than a formality. Roughly half of everyone who faithfully attends a long-running screening program will, at some point, be told the test found something — and be worked up, imaged, sometimes biopsied, before being cleared. The anxiety, cost, and occasional complication of those work-ups belong on the harms side of any screening ledger; the screening-harms guide prices them alongside overdiagnosis.
None of this is an argument that repetition is a mistake. Repeating a screen is how programs patch imperfect sensitivity — the cancer a mammogram misses this round gets another chance to be caught next round, and per-round misses are the reason a 79%-sensitive FIT can anchor a respectable program at all. Repetition buys real detection; it pays in accumulated false alarms. Both sides of that trade compound with n, which is why neither can be read off a single-round accuracy table.
The flip side: repetition as confirmation
Now run the same compounding on the positive side, and the villain becomes the hero. If a positive result is followed by a second, independent test — and the second is also positive — the errors multiply against each other. A false alarm now requires being unlucky twice: with 90% specificity per round, one person in ten is falsely flagged once, but only about one in a hundred twice. Meanwhile most true cases keep testing positive. Sequential positives concentrate truth.
The odds arithmetic from pre-test to post-test probability makes it exact. A 90%/90% test has a positive likelihood ratio of 9. At 1% prevalence, one positive lifts the probability of disease from 1% to just 8.3% — the base-rate fallacy in action. But a second independent positive multiplies the odds by 9 again: from 8.3%, the odds 0.083 ÷ 0.917 ≈ 0.09 become 0.82, and the probability lands near 45%. Two cheap, mediocre tests in agreement approach what one excellent test achieves. This is why confirmatory testing is the universal design pattern for screening programs: the two-assay HIV algorithm, the colonoscopy after a positive FIT, the diagnostic work-up after a screening mammogram. No screen is asked to be right; each is asked to be right enough that the next, better test is worth running.
The same independence caveat applies in this direction too — and here it cuts against you. Whatever caused a false positive once (a cross-reacting antibody, a benign calcification) may cause it again, so repeating the identical test confirms less than the formula promises. Good confirmation changes the mechanism: a different assay, a different modality, a tissue diagnosis. The math wants a second opinion, not an echo.
Try it
Walk the two-positives arithmetic yourself in the Bayesian-updating panel: start at 1% prevalence with a 90% / 90% test and apply the update twice — 1% to 8.3% to roughly 45%.
Open the Bayesian-updating view →The asymmetry worth remembering
One mechanism, two morals. Repeated screening of the disease-free compounds false alarms toward certainty-of-a-scare; repeated confirmation of a positive compounds specificity toward certainty-of-the-truth. Which one you get is a matter of design — who gets retested, with what, and why. So when a program is described to you as "a simple yearly test," the single-round accuracy is the least of what you need: ask what ten years of it does to a healthy attendee, and what happens after the first positive. Those two answers — not the per-round percentages — are the program.
References
- Elmore JG, Barton MB, Moceri VM, Polk S, Arena PJ, Fletcher SW. Ten-year risk of false positive screening mammograms and clinical breast examinations (cumulative false-positive risk 49.1% after 10 mammograms). New England Journal of Medicine, 1998.
- Croswell JM, Kramer BS, Kreimer AR, et al. Cumulative incidence of false-positive results in repeated, multimodal cancer screening (≈60.4% of men, 48.8% of women after 14 PLCO tests). Annals of Family Medicine, 2009.