Concept

Prevalence — where it appears

The share of a population that has a condition at a given time. It is the prior every screening calculation needs, and when estimated from an imperfect test's positive rate it must be corrected for the test's false positives and missed cases.

Named by 3 essays across one field — each of them below, with the objects they name alongside it.

Estimating a prevalence of 0.10% from 1,000 tests. The positive rate reads 5.09%, which is 50.9 times the truth. The Rogan-Gladen correction averages 0.100% — unbiased — with a standard deviation of 0.820 points against the positive rate's 0.695, and it comes out negative on 48.6% of samples.

The prevalence the test has to estimate

Every predictive value takes a prevalence as given, and the prevalence is usually estimated from the same test's positive rate. At a true prevalence of one in a thousand that rate reads 5.09% — fifty times the truth — and the correction that inverts it is unbiased, 18% more variable, and negative on 48.6% of samples of a thousand.

screening · Baserate
Thirty surveys of a thousand people, true prevalence 0.1%, test 90% sensitive and 95% specific. Each row is one survey. The thin bar is the corrected estimate ± 1.96 standard errors; the heavy bar is the exact interval for the positive rate carried through the correction; the tick is the one-sided 95% upper limit. 30 of the thirty printed intervals reach below zero and 0 lie entirely below it.

An upper limit is the finding

A survey of a thousand people with a 90%-sensitive, 95%-specific test, estimating a prevalence of one in a thousand, prints a corrected 95% interval that covers the truth 95.17% of the time — and reaches below zero in 97.95% of surveys and lies entirely below zero in 3.35%. Coverage is not the problem. The problem is what the interval can say, and the honest summary of a typical survey is one number: the prevalence is below 1.64%. Bringing that limit down to twice the truth takes 184,521 people at this specificity and 5,083 with a test that never gives a false positive.

screening · Baserate
The share of healthy people the harm-minimising threshold flags, against the prevalence, for three spreads of the diseased scores. A miss costs a hundred false alarms. With diseased scores as spread as healthy ones the best threshold flags more healthy people smoothly as the prevalence rises. With them twice as spread it flags 81.1% just below a prevalence of 25.7% and everyone just above it; three times as spread, 45.7% and then everyone at 13.0%.

A threshold that jumps

When a test's scores are equally spread in the healthy and the diseased, the harm-minimising threshold slides smoothly as the prevalence changes. Make the diseased scores twice as spread — the ordinary shape of a group that mixes mild and severe cases — and keep the published 90/95 pair, and the best threshold flags 81.1% of healthy people at a prevalence just under 25.7% and every healthy person just above it. At three times the spread it jumps from 45.7% to everyone at 13.0%. A threshold that jumps cannot be given as a formula; it has to be given with the second-best point beside it.

screening · Baserate

Named alongside it

The objects these essays reach for when they reach for this one.

ScreeningSensitivity and specificityBase rateMisclassificationSample sizeClopper–PearsonConfidence intervalCoverageDecision thresholdEstimator biasExpected lossLikelihood ratio

All concepts