Learn the math behind medical tests

Screen a thousand people for a condition that turns up in one of every hundred, using a test that is right ninety percent of the time, and about 108 of them will test positive. Nine of those positives are real. The other 99 are false alarms — healthy people the test flagged by mistake. So a positive result on this "ninety-percent-accurate" test is wrong more than nine times out of ten. The test is not broken. The healthy crowd it draws false alarms from is simply a hundred times larger than the handful who are actually sick, and a small slice of a huge group still outnumbers most of a tiny one.

That single gap — between how accurate a test sounds and what one positive result means for the person holding it — is the idea this whole library is built around. It fools almost everyone, including the clinicians who order the tests, because the mind quietly answers an easier question than the one it was asked. "Given a positive result, how likely is disease?" gets swapped for "given disease, how likely is a positive?" Those are different numbers, and for a rare condition they can sit a factor of forty apart. Untangling them is the subject of the base-rate fallacy.

Predictive value is only the first number worth knowing. A screening program also has to answer for what it does to everyone it flags — the extra scans, the biopsies, the treatments — each carrying a benefit and a harm that both have to be counted rather than assumed. A test that finds real disease can still leave a population worse off if it sends far more healthy people than sick ones into follow-up, or if it catches cancers that never would have caused trouble. None of this is intuitive until you have watched it worked out in whole people instead of percentages.

That is what the calculator does, and what these guides teach. Start with the scenario the tool opens with, then read the guide that explains it:

Try it

The classic case — 1,000 people, a 1% base rate, and a test that is 90% sensitive and 90% specific: 108 positives, only 9 of them real.

Open this scenario in the calculator →

The five core concepts below give you the machinery; the four worked tests run it on real, cited numbers from actual screening programs. Every figure on this page can be reproduced in the interactive worked example.

Core concepts

Read start to finish, these build the reasoning that a single positive or negative result actually requires.

Real screening tests, worked through

The same arithmetic, run on published accuracy and prevalence figures for four programs that sit at opposite corners of the map.

To feel the low-prevalence problem before you dive in, load a near-perfect test at a realistic base rate and watch its predictive value fall anyway:

Try it

A 99.8% / 99.5% test (HIV 4th-generation) at 0.4% prevalence — accuracy most tests never reach, and a positive result still only about 44% likely to be real.

Open this scenario in the calculator →

Go deeper

Two references sit under everything above. The Methodology lays out every formula, assumption, and idealization the calculator uses — including where the independence math for repeated testing breaks down. The Glossary gives plain, one-line definitions of sensitivity and specificity, PPV and NPV, likelihood ratios, and NNT, NNH, and NNS, so no term on any guide goes unexplained.

Educational model — not medical advice. It illustrates the statistics of testing and treatment; it does not describe any specific real-world test.