Repeated screening and the arithmetic of false alarms

Every accuracy number on this site describes a single round of testing. Almost no screening program consists of one. Mammography is invited every one to two years for decades; FIT is annual; PSA, low-dose CT, cervical screening — all recur. And repetition quietly transforms the arithmetic: a specificity that sounds reassuring per round compounds into a false-alarm rate per screening career that surprises nearly everyone, including the clinicians doing the inviting. This guide works that compounding in both directions — because the same mechanism that piles up false alarms is also, run the other way, how confirmation works.

The compounding, idealized

Take a test with 90% specificity: each round, a disease-free person has a 90% chance of being correctly cleared. Ask instead: what is the chance of getting through n rounds without a single false alarm? If each round erred independently, it would be 0.9 × 0.9 × … — that is 0.9ⁿ — and the chance of at least one false positive would be

P(≥1 false alarm in n rounds) = 1 − specificityⁿ.

The numbers escalate faster than intuition expects. At 90% specificity: 10% after one round, 27% after three, 41% after five, 65% after ten. Even an excellent 99%-specific test comes close to 10% after ten rounds. The per-round error never changed; exposure did. The same person keeps re-rolling the same dice, and screening programs are the re-rolling of dice, decade after decade, on millions of healthy people. The calculator's serial-testing panel draws this curve live for whatever specificity you set — drag the slider and watch the decade-long picture swing.

Try it

Start from the classic 90% / 90% test and open the serial-testing curve. If every round erred independently, 65% of disease-free people would be flagged at least once by round ten — from a test that is "90% accurate."

Open the serial-testing view →

What real programs actually observe

The formula rests on three assumptions: errors that are independent from round to round, the same specificity at every round, and a person who stays disease-free. Independence is an assumption, not a bound — real cumulative rates can come out below or above the curve. If each round has false-positive probability q, all that can be guaranteed for n ≥ 1 rounds is q ≤ P(at least one false positive) ≤ min(1, nq). At 90% specificity and ten rounds, that range runs from 10% to 100%, and independence's 65.1% is just one point inside it. If the same 10% of people are flagged every time, cumulative risk stays at 10%; if a different, non-overlapping 10% is flagged each round, it reaches 100%.

Observed cumulative rates can sit below or above the independence estimate; only the range from q to min(1, nq) is guaranteed, so neither study confirms or refutes the curve. What both show is scale: for a disease-free person, a screening career is closer to a coin flip than a formality. In each study, the estimated chance of at least one false alarm over a full course of screening was about half or more — and each false alarm means further evaluation, from repeat imaging or testing to, sometimes, a biopsy or other invasive procedure, before the all-clear. The anxiety, cost, and occasional complication of those work-ups belong on the harms side of any screening ledger; the screening-harms guide weighs them alongside overdiagnosis.

None of this is an argument that repetition is a mistake. Repeating a screen is how programs patch imperfect sensitivity — the cancer a mammogram misses this round gets another chance to be caught next round, and repetition is what lets a 79%-sensitive FIT anchor a respectable program despite its per-round misses. Repetition buys real detection; it pays in accumulated false alarms. Both sides of that trade compound with n, which is why neither can be read off a single-round accuracy table.

The flip side: repetition as confirmation

Now run the same compounding on the positive side, and the villain becomes the hero. If a positive result is followed by a second, independent test — and the second is also positive — the errors multiply against each other. A false alarm now requires being unlucky twice: with 90% specificity per round, one person in ten is falsely flagged once, but only about one in a hundred twice. Meanwhile most true cases keep testing positive. Sequential positives concentrate truth.

The odds arithmetic from pre-test to post-test probability makes it exact. A 90%/90% test has a positive likelihood ratio of 9. At 1% prevalence, one positive lifts the probability of disease from 1% to just 8.3% — the base-rate fallacy in action. But if the second test's mistakes are unrelated to the first's — among people with the disease and among people without it — a second positive multiplies the odds by 9 again: from 8.3%, the odds 0.083 ÷ 0.917 ≈ 0.09 become 0.82, and the probability lands near 45%. Two cheap, mediocre tests in agreement approach what one excellent test achieves. This is why confirmatory testing is the universal design pattern for screening programs: the two-assay HIV algorithm, the colonoscopy after a positive FIT, the diagnostic work-up after a screening mammogram. No screen is asked to be right; each is asked to be right enough that the next, better test is worth running.

The same independence caveat applies in this direction too — and here it cuts against you. Whatever caused a false positive once (a cross-reacting antibody, a benign calcification) may cause it again, so repeating the identical test confirms less than the formula promises. Good confirmation changes the mechanism: a different assay, a different modality, a tissue diagnosis. The math wants a second opinion, not an echo.

Try it

This scenario sets the prevalence to 8.33% — the probability of disease after one positive — and keeps the same 90% / 90% test. Its positive predictive value, about 45%, is the probability after a second positive, provided the second test's mistakes are unrelated to the first's.

Open the second positive update →

The asymmetry worth remembering

One mechanism, two morals. Under the independence model, repeated screening of the disease-free compounds false alarms toward near-certainty of a scare, while repeated confirmation of a positive compounds specificity toward near-certainty about the truth. Which one you get is a matter of design — who gets retested, with what, and why. So when a program is described to you as "a simple yearly test," the single-round accuracy is the least of what you need: ask what ten years of it does to a healthy attendee, and what happens after the first positive. Those two answers — not the per-round percentages — are the program.

Correction, September 2026. Elmore's 18.6% biopsy estimate applies to all women without breast cancer, and Croswell's invasive-procedure risks apply to every participant; this guide had described both as shares of people with a false alarm. It had also attributed the 69% comparison to Elmore's own data and presented both studies' lower rates as confirmation that rounds are correlated. The corrections log lists every change.

References

  1. Elmore JG, Barton MB, Moceri VM, Polk S, Arena PJ, Fletcher SW. Ten-year risk of false positive screening mammograms and clinical breast examinations (estimated cumulative false-positive risk 49.1% after 10 mammograms; estimated biopsy risk 18.6% among women without breast cancer). New England Journal of Medicine, 1998.
  2. Croswell JM, Kramer BS, Kreimer AR, et al. Cumulative incidence of false-positive results in repeated, multimodal cancer screening (after 14 PLCO tests, cumulative false-positive risk 60.4% in men and 48.8% in women; invasive diagnostic procedure prompted by a false positive 28.5% and 22.1%). Annals of Family Medicine, 2009.
  3. Lehman CD, Arao RF, Sprague BL, et al. National performance benchmarks for modern screening digital mammography: update from the Breast Cancer Surveillance Consortium (specificity 88.9%). Radiology, 2017.

Educational model — not medical advice. It illustrates the statistics of testing and treatment; it does not describe any specific real-world test.