Learn the math behind medical tests
Screen a thousand people for a condition that turns up in one of every hundred, using a test that is right ninety percent of the time, and about 108 of them will test positive. Nine of those positives are real. The other 99 are false alarms — healthy people the test flagged by mistake. So a positive result on this "ninety-percent-accurate" test is wrong more than nine times out of ten. The test is not broken. The healthy crowd it draws false alarms from is simply a hundred times larger than the handful who are actually sick, and a small slice of a huge group still outnumbers most of a tiny one.
That single gap — between how accurate a test sounds and what one positive result means for the person holding it — is the idea this whole library is built around. It fools almost everyone, including the clinicians who order the tests, because the mind quietly answers an easier question than the one it was asked. "Given a positive result, how likely is disease?" gets swapped for "given disease, how likely is a positive?" Those are different numbers, and for a rare condition they can sit a factor of forty apart. Untangling them is the subject of the base-rate fallacy.
Predictive value is only the first number worth knowing. A screening program also has to answer for what it does to everyone it flags — the extra scans, the biopsies, the treatments — each carrying a benefit and a harm that both have to be counted rather than assumed. A test that finds real disease can still leave a population worse off if it sends far more healthy people than sick ones into follow-up, or if it catches cancers that never would have caused trouble. None of this is intuitive until you have watched it worked out in whole people instead of percentages.
That is what the calculator does, and what these guides teach. Start with the scenario the tool opens with, then read the guide that explains it:
Try it
The classic case — 1,000 people, a 1% base rate, and a test that is 90% sensitive and 90% specific: 108 positives, only 9 of them real.
Open this scenario in the calculator →The five core concepts below give you the machinery; the four worked tests run it on real, cited numbers from actual screening programs. Every figure on this page can be reproduced in the interactive worked example.
Core concepts
Read start to finish, these build the reasoning that a single positive or negative result actually requires.
- Sensitivity vs. specificity — What each number really measures, read straight off the 2×2 table, why moving a test's cutoff trades one for the other, and why neither one answers "I tested positive — do I have it?"
- The base-rate fallacy — Why a genuinely accurate test produces mostly false positives when a disease is rare, the reversal that trips up even practicing physicians, and the natural-frequency trick that makes it obvious.
- Likelihood ratios & the Fagan nomogram — How one prevalence-independent number captures the strength of a result, and how a Fagan nomogram slides a pre-test probability to a post-test one with a straightedge instead of Bayes' formula.
- From pre-test to post-test probability — Bayes in its odds form: turn a starting probability into the probability after a result, and see why a second test taken right after the first behaves so differently.
- Screening harms & biases — The costs a detection rate hides — false-positive work-ups, overdiagnosis, and the lead-time and length-time biases that flatter almost every screening statistic.
Real screening tests, worked through
The same arithmetic, run on published accuracy and prevalence figures for four programs that sit at opposite corners of the map.
- Mammography, by the numbers — A solid ~87% / 89% test at a ~0.5% base rate, and why most positive screening mammograms still turn out to be false alarms that resolve on further work-up.
- The PSA test, by the numbers — A deliberately low-sensitivity test that catches only about one prostate cancer in five at the usual 4 ng/mL cutoff: what it misses, what a positive is worth, and the number needed to screen behind it.
- HIV testing, by the numbers — The opposite extreme, a near-perfect 99.8% / 99.5% test, showing why even that leaves a first positive needing confirmation when prevalence is very low — and why the two-step algorithm exists.
- NIPT prenatal screening, by the numbers — Why "99% accurate" cell-free-DNA screening for Down syndrome is a screen and not a diagnosis: at a ~0.25% base rate a real share of positives are false, and the confirmatory test carries its own risk.
To feel the low-prevalence problem before you dive in, load a near-perfect test at a realistic base rate and watch its predictive value fall anyway:
Try it
A 99.8% / 99.5% test (HIV 4th-generation) at 0.4% prevalence — accuracy most tests never reach, and a positive result still only about 44% likely to be real.
Open this scenario in the calculator →Go deeper
Two references sit under everything above. The Methodology lays out every formula, assumption, and idealization the calculator uses — including where the independence math for repeated testing breaks down. The Glossary gives plain, one-line definitions of sensitivity and specificity, PPV and NPV, likelihood ratios, and NNT, NNH, and NNS, so no term on any guide goes unexplained.