A glossary of screening and diagnostic-test statistics
A screening test that looks 90% accurate can still be wrong about a positive result more than nine times out of ten. Reading a result correctly takes a small, exact vocabulary — each word with a plain meaning and a precise formula. This glossary defines the calculator's terms alphabetically, plain sense first and arithmetic second, linking each concept's fuller guide in the Learn library.
The terms
Formulas follow the site's methodology; rates are decimals from 0 to 1, and cohorts are whole people.
Absolute vs. relative risk reduction (ARR, RRR)
Two ways to report a treatment's benefit. Absolute risk reduction is the plain drop in event rate — 2% to 1% is an ARR of one percentage point. Relative risk reduction divides that by the starting risk, so identical data sounds far larger. ARR ties to number needed to treat: NNT = 1 ÷ ARR; relative figures quoted alone routinely mislead.
Base rate
How common a condition is in the group being tested, before any result — the same quantity as prevalence and pre-test probability. It is what a positive result updates from, and the number intuition most often drops. Neglecting it is the base-rate fallacy: judging a positive by test accuracy alone, when the real chance of disease depends as much on rarity.
Bayes' theorem
The rule for revising a probability when evidence arrives — here, turning a pre-test into a post-test probability. In frequency form, PPV = (sens × prev) ÷ [sens × prev + (1 − spec) × (1 − prev)]. The odds form is lighter: post-test odds = pre-test odds × likelihood ratio. Both give the same number, which the calculator draws as a tree and nomogram.
Fagan nomogram
A three-column chart that does Bayes' theorem with a ruler. Mark the pre-test probability on the left axis and the likelihood ratio in the middle, draw a straight line through them, and it meets the right axis at the post-test probability. Introduced by Fagan in 1975, it makes vivid how much a given result moves the odds. See likelihood ratios.
False positive and false negative
The two ways a test is wrong. A false positive flags someone without the condition; a false negative misses someone who has it. False positives drive needless follow-up, cost, and anxiety, and pile up when a condition is rare. False negatives give false reassurance and can delay care. Every cutoff trades one against the other.
Incidence vs. prevalence
Two ways to count a disease. Prevalence is the share of people who have it at one moment — a snapshot, and the pre-test probability screening uses. Incidence is the rate of new cases arising over a period, such as a year. A long-lasting disease can show high prevalence from modest incidence; a fast, fatal one, low prevalence despite high incidence.
Lead-time bias
An illusion that makes screening look life-extending when it may only start the clock sooner. Detect a cancer three years earlier and survival measured from diagnosis gains three years — even if death comes the same day it always would have. So survival comparisons between screened and unscreened groups mislead; only a mortality difference settles it. See screening harms.
Length-time bias
A sampling quirk that flatters screening. Indolent, slow-growing disease stays detectable and symptom-free longer, so periodic screens catch it more often than fast disease that surfaces between rounds. Screen-detected cases skew toward the mildest end, making screened patients look better regardless of treatment. Its extreme form is overdiagnosis. See screening harms.
Likelihood ratio (LR+ and LR−)
How strongly one result shifts the odds of disease, built only from the test. LR+ = sensitivity ÷ (1 − specificity) grades a positive; LR− = (1 − sensitivity) ÷ specificity grades a negative. An LR+ far above 1 argues for disease; an LR− near 0 argues against it. Using only the test's own numbers, they don't depend on prevalence. See likelihood ratios.
Negative predictive value (NPV)
Given a negative result, the chance the person truly is disease-free — read across the negative row: NPV = true negatives ÷ (true negatives + false negatives). Like PPV it moves with prevalence: when a condition is rare, NPV is high almost automatically, since most people genuinely don't have it. Residual risk after a negative test is 1 − NPV.
Number needed to harm (NNH)
How many people must go through a treatment or screening pathway for one to suffer a specified harm. An NNH of 200 means one extra harmful event per 200 people exposed. It is the honest counterweight to number needed to treat — a benefit count only means something beside a harm count on the same denominator. A smaller NNH is worse.
Number needed to screen (NNS)
How many people must be screened for one to be helped — testing, follow-up, treatment, and outcome folded into one figure. In the calculator, NNS = population ÷ people helped. It usually dwarfs the treatment's own number needed to treat, because screening also processes everyone healthy or falsely flagged. Reported screening NNS values run from the hundreds into the thousands.
Number needed to treat (NNT)
How many patients must receive a treatment for one to benefit who otherwise would not. An NNT of 20 means 19 of every 20 treated gain nothing from it. It is the reciprocal of the absolute risk reduction (NNT = 1 ÷ ARR) and reads best whole, not rounded down. Smaller is better: an NNT of 5 helps far more often than one of 100.
Overdiagnosis
Finding a real abnormality that would never have caused symptoms or shortened life — a true positive by the test's standard that still helps no one and invites needless treatment. It is not a false positive; the finding is genuinely present. It is the hardest screening harm to measure, since you cannot tell in advance who was overdiagnosed. See screening harms.
Positive predictive value (PPV)
Given a positive result, the chance the person truly has the condition — the number patients actually care about. PPV = true positives ÷ all positives = (sens × prev) ÷ [sens × prev + (1 − spec) × (1 − prev)]. It hinges on prevalence: a 90/90 test at 1% prevalence yields a PPV near 8%, so most positives are false. See the worked example.
Try it
The classic case: 1,000 people, 1% prevalence, a test that is 90% sensitive and 90% specific — 108 positives, only about 9 of them real.
Open this scenario in the calculator →Pre-test vs. post-test probability
The probability of disease before and after you know a result. The pre-test probability is the starting estimate — usually the prevalence in the relevant group. The post-test probability is what the test updates it to: after a positive it equals the PPV; after a negative, 1 − NPV. The wider the gap between them, the more the test told you. See pre-test to post-test probability.
Prevalence (pre-test probability)
How common a condition is in the group being tested — the single biggest driver of what a positive result means. It equals the base rate and, before any test, the pre-test probability. Raise it and predictive value climbs; lower it and even an excellent test yields mostly false alarms, because the healthy pool those alarms come from grows. A property of the population, not the test.
Try it
Load a 99.8% / 99.5% HIV assay at its 0.4% population prevalence: even that test leaves a positive result only about 44% likely to be real.
Open this scenario in the calculator →ROC curve
The receiver operating characteristic curve plots sensitivity against the false-positive rate (1 − specificity) as the positivity cutoff sweeps every value. It shows the whole menu of sensitivity/specificity trade-offs a test offers, not one operating point. Area under the curve summarizes discrimination: 0.5 is a coin flip, 1.0 perfect separation. Choosing a cutoff picks one point on the curve.
Sensitivity
Of the people who truly have the condition, the fraction the test correctly flags positive — P(positive | disease). A 90%-sensitive test misses one case in ten. It is a fixed property of the test, measured against a reference standard, not a verdict about any one result; a highly sensitive test earns its keep on its negatives. See sensitivity vs. specificity.
Serial (repeat) testing
Screening the same people round after round. Each round is a fresh chance at a false alarm, so cumulative false-positive risk grows: independent rounds would give 1 − specificityn — about 65% after ten rounds at 90% specificity. Real repeats are correlated, making that an upper bound; still, roughly half of women get a false-positive mammogram over a screening decade.
Specificity
Of the people who do not have the condition, the fraction the test correctly clears as negative — P(negative | no disease). A 90%-specific test wrongly flags one healthy person in ten, and across a large healthy group those false positives mount fast. Like sensitivity, it is fixed regardless of prevalence; a highly specific test earns its keep on its positives. See sensitivity vs. specificity.
True positive and true negative
The two ways a test is right. A true positive correctly flags someone who has the condition; a true negative correctly clears someone who does not. With the two errors, they form the four cells of the 2×2 confusion matrix behind every other measure — sensitivity and specificity read down its columns, predictive values across its rows. Watch them fill in under the test.
References
- Altman DG, Bland JM. Diagnostic tests 1: sensitivity and specificity. BMJ, 1994;308(6943):1552.
- Deeks JJ, Altman DG. Diagnostic tests 4: likelihood ratios. BMJ, 2004;329(7458):168–169.
- Fagan TJ. Nomogram for Bayes's theorem. N Engl J Med, 1975;293(5):257.
- Grimes DA, Schulz KF. Uses and abuses of screening tests. The Lancet, 2002;359(9309):881–884.