The cascade
A cohort of 1,000 flows left to right: who has the disease, what the test says, who gets treated, and modeled outcomes. Band height = expected people. The outcome partition shows the minimum possible overlap; benefit and harm can occur together.
What happens to everyone
Rounded illustration with the minimum possible benefit–harm overlap. One square per person.
Compare two scenarios
Save a baseline, then change any input. Both columns use 1,000 screened so population size cannot disguise a change in rates. The initial baseline is the teaching example.
| Measure | Saved | Current | Difference |
|---|---|---|---|
| Prevalence | 1.00% | 1.00% | 0.00 pp |
| Sensitivity | 90.00% | 90.00% | 0.00 pp |
| Specificity | 90.00% | 90.00% | 0.00 pp |
| Positive predictive value | 8.33% | 8.33% | 0.00 pp |
| True positives / 1,000 | 9.00 | 9.00 | 0.00 |
| False positives / 1,000 | 99.00 | 99.00 | 0.00 |
| Missed disease / 1,000 | 1.00 | 1.00 | 0.00 |
| Expected benefit / 1,000 | 0.90 | 0.90 | 0.00 |
| Expected harm / 1,000 | 5.40 | 5.40 | 0.00 |
Saved treatment: NNT 10, NNH 20, uptake 100.0%. Current: NNT 10, NNH 20, uptake 100.0%. Benefit and harm describe different endpoints and may overlap; their difference is not a net clinical benefit. The saved baseline lasts while this page is open and is not included in shared links.
The test
2×2 confusion matrix
Rounded counts for 1,000 people. Probabilities use unrounded expected counts. Sensitivity reads across the disease row; PPV reads down the test-positive column — they look at the table from perpendicular directions.
| Test + | Test − | Total | |
|---|---|---|---|
| Disease + | TP9 | FN1 | 10 |
| Disease − | FP99 | TN891 | 990 |
| Total | 108 | 892 | 1,000 |
- Sensitivity TP ÷ (TP + FN) disease + row →
- Specificity TN ÷ (TN + FP) disease − row →
- PPV TP ÷ (TP + FP) test + column ↓
- NPV TN ÷ (TN + FN) test − column ↓
- LR+ sensitivity ÷ (1 − specificity)
- LR− (1 − sensitivity) ÷ specificity
Probability tree
Split the population by disease, then by test result. Each path multiplies to a joint probability; Bayes just compares the two test-positive leaves.
P(disease | test +) = 0.90% ÷ (0.90% + 9.9%) = 8.3%
How PPV collapses with prevalence
Holding sensitivity and specificity fixed, the value of a positive result depends almost entirely on how common the disease is. The dashed line marks the current prevalence.
The base-rate fallacy: at 1.00% prevalence, even a 90%/90% test makes a positive result correct only 8% of the time. Accuracy isn't the whole story — the base rate is.
Fagan nomogram
The likelihood ratio is the lever that turns a pre-test probability into a post-test one — and it doesn't depend on prevalence. Line shown for a positive result (LR+).
This is Bayes' theorem, in odds form: prior odds × likelihood ratio = posterior odds. A large LR+ or small LR− can shift the odds strongly; neither is a universal diagnostic threshold. The starting probability still matters.
The treatment
Treatment outcomes
Illustrative treatment model. Use NNT and NNH for the same population, outcome definitions, comparison, and follow-up period. A trial’s number needed to screen is not a treatment NNT.
108 interventions performed — 99 on people who never had the disease and so could not benefit.
Benefit and harm above are expected counts and may overlap. Within this illustration, between 0.00 and 0.90 people could experience both. The bands and squares use the minimum (0.00); the inputs do not determine the actual overlap. “Neither” means neither modeled treatment effect, not a prediction of a person’s health.
Per 1,000 screened: about 0.90 helped, 5.40 harmed by treatment, 99.00 false alarms, and 1.00 missed.
Watch the relative-vs-absolute trap: a large “relative risk reduction” can still mean a large NNT when the baseline risk is low. Benefit (NNT) only means something next to its harms — that's why helped and harmed are always shown on the same denominator here.
Repeat testing
Serial testing — the false-alarm pile-up
For someone who remains disease-free, 1 − specificityn gives the chance of at least one false alarm if every round has the same specificity and errors are independent. Dependence can put the result above or below this curve. Hover, drag, or use the round control below.
At 10 rounds: 65.1% under independence. With the same per-round specificity but unknown dependence, the possible range is 10.0%–100.0%. Bounds: 1 − specificity to min(1, n × (1 − specificity)).
Reference points (≥1 false positive): Elmore et al., NEJM 1998 — 49% after 10 mammograms · Croswell et al., Ann Fam Med 2009 — ~60% (men) / 49% (women) after 14 multimodal PLCO tests.
Bayesian updating
Bayesian updating: learning a rate from data
Where do numbers like sensitivity or prevalence come from? You start with a prior belief, observe data, and get a posterior. For a rate, this is exact
and runs right here — no simulation: posterior = Beta(α+k, β+n−k).
The posterior mean sits between your prior mean and the observed rate — and more data gives the observations more weight. Additional data need not narrow the interval on every update. That shrinking uncertainty is what the screening sliders quietly assume away by treating each rate as a fixed number.