HIV testing, by the numbers

The HIV antigen/antibody combination assay is about as good as a screening test gets: roughly 99.8% sensitive and 99.5% specific. Draw someone at random from the U.S. adult population, run the test, get a reactive result — and the chance they actually have HIV is about 44%. Not 99%. Not even a slim majority. Barely better than a coin flip. That gap, between a nearly flawless test and a nearly uninformative-looking positive, is not a paradox or a lab error. It is arithmetic, and it is the reason no one is ever told they have HIV on the strength of a single reactive screen.

The single reactive test, worked by hand

Three numbers fix everything downstream — two describing the test, one describing the population. The U.S. Preventive Services Task Force review puts the fourth-generation Ag/Ab combination assay at about 99.8% sensitivity and 99.5% specificity (the reported ranges run to 100% for both; this page uses the conservative low end for specificity, because that is the hardest case for the number we care about). HIV prevalence among U.S. adults is about 0.40% — roughly 1 in 250. Screen 100,000 people at those figures and every cell of the table is determined.

Split the group first: 100,000 × 0.004 = 400 people with HIV, and 99,600 without. Now apply the test to each group. Of the 400 with HIV, 99.8% are flagged — 399 true positives, with fewer than one case missed. Of the 99,600 without HIV, 0.5% are flagged anyway — 498 false positives — while 99,102 are correctly cleared.

Screening 100,000 U.S. adults for HIV — prevalence 0.40%, sensitivity 99.8%, specificity 99.5%. Counts rounded to whole people.
 Test reactiveTest non-reactiveTotal
HIV present399 (true positives)1 (false negatives)400
HIV absent498 (false positives)99,102 (true negatives)99,600
Total89799,103100,000

Follow the reactive column. 897 people get a positive result, and only 399 of them have HIV. The share of positives that are real is the positive predictive value:

PPV = 399 ÷ 897 ≈ 0.445 — about 44–45%.

So at population prevalence, more than half of reactive fourth-generation screens — 498 of 897 — belong to people who do not have HIV. The test barely put a foot wrong: it correctly cleared 99.5% of the healthy. But the healthy group is 249 times larger than the infected one, and 0.5% of a very large number still outweighs 99.8% of a very small one. The false positives win on sheer volume. This is the base-rate fallacy in its sharpest form — the accuracy of a test and the meaning of a positive are two different quantities, and prevalence is what separates them.

The non-reactive column is the genuinely reassuring one. 99,103 people get a negative result and all but one are truly HIV-free — a negative predictive value of about 99.999%. A non-reactive fourth-generation screen, taken after the window period, is close to decisive; a reactive one, at this prevalence, is a question rather than an answer.

Try it

Screen 100,000 U.S. adults for HIV at 0.40% prevalence with the 99.8% / 99.5% assay, and watch the roughly 500 false positives outnumber the 399 real ones. (The treatment sliders here are generic placeholders — this page is about the test, not the therapy.)

Open this scenario in the calculator →

The same 44% in odds

The 2×2 count is one of two roads to the same number; the other runs in odds. The positive likelihood ratio of this test is its true-positive rate over its false-positive rate: 0.998 ÷ 0.005 ≈ 200 — a reactive result is about 200 times more common in someone with HIV than in someone without. Convert the 0.40% prevalence to pre-test odds (0.004 ÷ 0.996 ≈ 0.00402), multiply by the likelihood ratio (× 200 ≈ 0.80), and convert back to a probability (0.80 ÷ 1.80 ≈ 0.445). The same 44.5%. An LR near 200 still lands on a coin flip because the starting odds were so long. Likelihood ratios and the Fagan nomogram covers that machinery; the point here is only that the odds path and the confusion-matrix path agree to the decimal — one theorem in two costumes.

Why the algorithm keeps testing

A result that lands at 44% is not one you act on — it is one you follow. Because prevalence drags the first PPV down, the U.S. laboratory diagnostic algorithm, issued by the CDC and the Association of Public Health Laboratories, never stops at a single result. A reactive fourth-generation immunoassay is reflexed to a second, different test: an HIV-1/HIV-2 antibody differentiation assay. If that is also reactive, the infection is confirmed. If the two disagree — a reactive screen but a non-reactive or indeterminate differentiation assay — the sample proceeds to a third, independent test: an HIV-1 nucleic acid test (NAT) that looks for viral RNA directly.

Here is the mechanism the arithmetic reveals. Once the first screen is reactive, the pre-test probability for the next test is no longer 0.40% — it is 44.5%, the first test's PPV. Feed that number back in and run an independent test of similar strength, and Bayes multiplies the odds a second time: pre-test odds of 0.445 ÷ 0.555 ≈ 0.80, times an LR near 200, gives post-test odds around 160, which converts back to about 99.4%.

Try it

Take the first test's 44.5% PPV as the new pre-test probability and run an equally strong independent test — the posterior of one test becomes the prior of the next, and the PPV jumps to about 99%.

Open the confirmatory step in the calculator →

The calculator's Bayesian-updating panel animates exactly that hand-off. Two caveats keep it from being magic. The real differentiation assay is not a repeat of the fourth-generation test — it is even more specific — so the equal-strength second test above is an illustration of the principle, not a spec sheet; the actual jump is at least this large. And confirmatory errors are not perfectly independent: a cross-reacting antibody — from pregnancy, an autoimmune condition, or a recent vaccination — can fool more than one antibody-based test the same way. That is why the later steps switch to a different class of test, a NAT that detects viral RNA rather than antibody, so the two are unlikely to fail for the same reason.

A higher-risk group starts further along

Prevalence is a lever, and pulling it changes the first PPV before any confirmation runs. Take the identical 99.8% / 99.5% assay to a group where roughly 1 in 20 has HIV — an illustrative 5% pre-test probability, used to show the mechanism, not a figure for any named population — and the first reactive result already means something very different. Split 100,000 the same way: 5,000 have HIV and 95,000 do not. The test flags 4,990 of the infected and 475 of the healthy, so 5,465 reactive results now contain 4,990 true positives:

PPV = 4,990 ÷ 5,465 ≈ 0.913 — about 91%.

Try it

Same assay, a higher 5% pre-test probability (illustrative only). Watch the positive predictive value climb from ~44% to ~91% while sensitivity and specificity stay fixed.

Open the higher-risk variant →

Nothing about the assay changed; only the population did. The same reactive result that meant 44% odds at 0.40% prevalence means about 91% at 5%. Pre-test probability in, post-test probability out — the engine behind every screening result, and the reason a test's paperwork can never tell you what your own positive means without knowing who was tested.

The mirror image of serial screening

It is worth setting this against the more familiar serial-testing story, which runs the other way. Elsewhere on this site — in mammography, in the base-rate fallacy — repeating a screen is a liability: each annual round is a fresh chance for a false alarm, so the cumulative probability of at least one false positive climbs toward 1 − specificityⁿ over the years, and the calculator's repeat-testing view plots that rise. Serial screening repeats one test against time to catch new disease, and pays for it in accumulating false positives. Serial diagnosis runs several different tests against one reactive sample to settle a single question, and is repaid in a PPV driven toward certainty. Same multiplication of independent tests; opposite purpose, and opposite effect on the false-positive count.

The window period, and what the algorithm is really for

The fourth-generation screen owes its early detection to p24 antigen, which appears before antibodies do — so it can turn reactive roughly two to six weeks after exposure, though a negative taken inside that window reflects the window, not the absence of virus. That same biology shapes the confirmatory chain. In acute infection the screen can be reactive while the antibody differentiation assay is still negative; rather than dismiss the clash as a false alarm, the algorithm reads it as the signature of very recent infection and routes the sample to the NAT, which confirms the acute HIV an antibody-only approach would miss.

So the algorithm does two jobs at once: it drives the predictive value of a true positive toward certainty, and it catches the rare true positive the second test alone would have called negative. None of this is guidance about testing or treatment. What the numbers show is narrow and durable: a superb test can still deal a 44% positive when the condition is rare, and the remedy is not a better single test but a second and third independent one.

References

  1. U.S. Preventive Services Task Force. Human Immunodeficiency Virus (HIV) Infection: Screening (fourth-generation assay sensitivity 99.8% / specificity 99.5%; U.S. adult prevalence ~0.40%). USPSTF, 2019.
  2. Centers for Disease Control and Prevention and Association of Public Health Laboratories. Laboratory Testing for the Diagnosis of HIV Infection: Updated Recommendations (Ag/Ab immunoassay → HIV-1/HIV-2 antibody differentiation assay → HIV-1 nucleic acid test). CDC/APHL, 2014.
  3. Fagan TJ. Nomogram for Bayes's theorem (letter). New England Journal of Medicine, 1975.
  4. Jaeschke R, Guyatt GH, Sackett DL. Users' guides to the medical literature. III. How to use an article about a diagnostic test (likelihood ratios and post-test probability). JAMA, 1994.

Educational model — not medical advice. It illustrates the statistics of testing and treatment; it does not describe any specific real-world test.