An upper limit is the finding
Worth reading first: What a positive test is worth.
The prevalence the test has to estimate found that the raw positive rate of a 90%-sensitive, 95%-specific test reads fifty times the truth when one person in a thousand has the condition, and that the Rogan–Gladen correction that inverts it is unbiased and comes out negative on nearly half of surveys of a thousand. It left the question of what an interval for that prevalence should be, and named three candidates without measuring them: the corrected estimate plus or minus two standard errors, an exact interval carried through the correction and perhaps cut off at zero, and no interval at all — just an upper limit.
All three can be measured exactly. A survey of people returns a count between 0 and , the count is binomial with the apparent rate, and every property of every report is a sum over the possible counts. This essay does those sums and finds that the question was posed about the wrong property. Every report covers the truth about as often as it claims to. What distinguishes them is what they say when they do.
Four reports of one count
The apparent rate is , and the corrected estimate from positives in tests is . From a single count there are four natural things to report.
The printed interval: standard errors, where the standard error is the apparent rate’s divided by . This is what most papers show.
The mapped exact interval: the Clopper–Pearson interval for the apparent rate, with both limits carried through the correction. The correction is an increasing straight line, so the mapped interval contains exactly when the original contains , and it inherits the exact interval’s guarantee of never covering less than 95%.
The truncated interval: the mapped interval with any negative limit raised to zero. The prevalence cannot be negative, so raising a limit to zero never removes the truth from the interval; coverage is unchanged and the interval is shorter.
The upper limit: the one-sided 95% upper limit alone, from the exact one-sided limit for the apparent rate. It makes one claim — the prevalence is below this — and makes it with 95% confidence.
They all cover
The sums over every possible count, for surveys of a thousand:
| true prevalence | printed interval covers | mapped exact covers | upper limit holds | printed reaches below zero | printed lies below zero |
|---|---|---|---|---|---|
| 0.1% | 95.17% | 95.61% | 95.30% | 97.95% | 3.35% |
| 0.5% | 95.23% | 95.69% | 95.36% | 93.86% | 1.10% |
| 1% | 94.57% | 95.72% | 95.08% | 82.79% | 0.22% |
| 5% | 94.60% | 95.67% | 95.20% | 0.10% | 0.00% |
Coverage is not where anything goes wrong. The printed interval covers within half a point of 95% at every prevalence here, which is what one would expect of a normal interval around a count whose mean is about fifty; the apparent rate is dominated by false positives, so the count is never small even when the prevalence is. The exact interval covers a little above 95%, as it is built to.
What goes wrong is in the last two columns. At a prevalence of 0.1%, 97.95% of printed intervals reach below zero, and 3.35% lie entirely below it — an interval for a proportion that excludes every proportion there is. At 1% the lower limit is still negative in 82.79% of surveys. Those intervals are correct in the only sense coverage measures, and nearly all of them say that the data are consistent with no cases at all.
What each report says
The four reports differ in what they would lead a reader to write down.
The printed interval at a prevalence of 0.1% is typically about 3.2 points wide and centred near 0.1%: something like −1.5% to 1.7%. It states a lower limit that is not a prevalence, and one survey in thirty says, in effect, that fewer than no people are affected. A reader who quotes it has to explain its left half away, and a reader who quotes only its centre is quoting the estimate that is negative half the time.
The mapped exact interval has the same problem with slightly different numbers: it reaches below zero in 97.20% of surveys and lies entirely below it in 1.59%. The second number has a use the printed interval’s does not. When the exact interval lies below zero, the survey found fewer positives than the test’s stated false-positive rate alone would produce, beyond what chance allows at 2.5% — which is evidence that the stated specificity is too pessimistic for this population, not evidence about the prevalence. It is the one report in which an impossible answer carries information, and the information is about the test.
The truncated interval replaces every negative limit with zero, so a typical survey reports and one survey in sixty reports . It is honest in the obvious way — every value in it is possible — and its average width falls from 3.32 points to 1.91. But a reader sees a two-sided interval with a lower limit, and the lower limit is not an estimate of anything; it is the boundary of the parameter space, wearing an interval’s clothes.
The upper limit says the one thing the data support. At a true prevalence of 0.1% the typical survey of a thousand reports “below 1.64%, with 95% confidence”, and that statement is true 95.30% of the time. It does not pretend to a lower limit it cannot have. It is also, plainly, not much of a finding: the data can rule out a prevalence above one and a half per cent, sixteen times the truth, and nothing below.
Ten times the survey
A larger survey is the obvious response, and it helps less than a reader might hope.
At ten thousand people tested and a prevalence of 0.1%, the printed interval still reaches below zero in 94.82% of surveys and lies entirely below it in 1.06%. The typical upper limit falls to 0.536% — from sixteen times the truth to a little over five. At a prevalence of 1% the change is decisive: only 4.11% of intervals reach below zero, and the typical upper limit, 1.47%, is within half the truth of it. Ten times the survey moves the boundary between “a prevalence the survey can estimate” and “a prevalence it can only bound” from somewhere above 1% to somewhere below it — and leaves one in a thousand firmly on the wrong side.
That boundary is the practical content of the whole calculation. For a given test and survey size there is a prevalence below which the honest report is an upper limit and above which it is an interval, and it sits roughly where the true positives expected in the survey become comparable to the standard deviation of the false positives. A planner can compute it before the survey is run, from the four numbers the report should print, and the calculation says which kind of answer the survey is going to give.
The same shape in other places
The difficulty here is not peculiar to screening. It is the general shape of estimating a small quantity by subtracting a large known one from a measurement, and it recurs wherever that is done.
A radioactivity count with a known background is the physics version: the source’s rate is the total minus the background, the estimate is negative when the count falls below the expected background, and the unified intervals physicists use for it were devised exactly because the naive interval can lie entirely in the impossible region. A proportion observed near zero is the version without a background: a hole no sample size fills found an interval’s coverage dipping where one success stops covering, which is the same boundary felt from the other side. And a prior is the Bayesian answer to both — it puts no mass below zero, so its interval cannot go there — at the cost that for a rare condition much of what the interval says comes from the prior rather than the survey.
What unites them is that the parameter has a boundary and the data are noisy on the scale of the parameter’s distance from it. Near such a boundary a two-sided interval is the wrong shape for the information, because the information is one-sided: the data can say “not more than this” and have nothing to say about “not less”. An upper limit is not a weaker report than an interval in that regime. It is the report that has the shape of what was learned, and the odds form of the base-rate arithmetic makes the same point about a single positive result: when the prior is small, the useful statement is how large the posterior could be, not where its centre is.
Why the limit sits so far above the truth
The upper limit is uninformative for a reason the arithmetic makes exact. Of the fifty or so positives a survey of a thousand expects, about fifty are false — the false-positive rate of 5% applied to 999 unaffected people — and 0.9 are true. The one true positive is added to a count whose own sampling noise, a binomial standard deviation of about seven, is several times larger than it. The survey is trying to see one case through a noise floor of seven.
The noise floor belongs to the test, not to the survey. Doubling the survey halves the variance of the false-positive count relative to its size but leaves the ratio of signal to floor unchanged per person, so the upper limit falls only as the square root of the number tested. The specificity sets the floor directly: every percentage point of false-positive rate adds about eleven false positives for each true one at this prevalence.
| specificity | people to test for an upper limit of 0.2% |
|---|---|
| 95% | 184,521 |
| 99% | 39,957 |
| 99.9% | 9,110 |
| 100% | 5,083 |
At 95% specificity, pinning a prevalence of one in a thousand below two in a thousand takes 184,521 people. At 99.9% specificity it takes 9,110, and with no false positives at all, 5,083 — which is the survey the sensitivity alone requires, since a perfect-specificity test still misses a tenth of cases. The ratio between the first and last rows, thirty-six, is the price of the false positives, and no amount of care in the correction recovers it: the correction removes their expected number exactly and cannot remove their noise.
At a prevalence of 1% the same bound takes 2,362 people at 95% specificity. The problem is specific to the rare conditions the predictive value arithmetic is most often applied to, which is where the prevalence is also hardest to know.
What a report should say
The measurements support a narrow recommendation and a broader one.
For a rare condition surveyed with an imperfect test, report the one-sided upper limit as the finding. It is the one report whose only claim is supported, it has its stated confidence, and it says in the plainest form what the survey could and could not establish. “The prevalence is below 1.64%” survives every objection this essay and the one before it raise; “the prevalence is 0.1%, 95% interval −1.5% to 1.7%” survives none of them, and “0.1%, 95% interval 0% to 1.7%” survives most of them while inviting a misreading of its lower limit.
Report the count and the test’s characteristics beside it. The upper limit is a function of the count, the sample size, the sensitivity and the specificity, and a reader who wants a two-sided interval, or wants to redo the arithmetic with a better estimate of specificity, can do so only if all four are printed. A surprising count — fewer positives than the false-positive rate predicts — should be reported as what it is: a signal that the test performs better in this population than its validation said, which bears on every predictive value computed from it.
Size the survey for the bound, not for the estimate. A survey planned to “estimate the prevalence” of a rare condition to a given precision will typically be planned from the apparent rate’s binomial standard error and will be far too small, because the relevant noise is the false positives’. The table above is the plan: the number of people it takes for the upper limit to say something a decision could use. A planner who cannot afford it should know before the survey is run that it will return an upper limit sixteen times the truth, and a better use of the budget may be a more specific confirmatory test, which moves the floor rather than averaging over it — the same lever the second test that is not a second opinion found blunted when the two tests’ errors are correlated.
And do not pool the truncated reports. A review that combines prevalence surveys by averaging their reported estimates, having replaced each negative one with zero, has built a biased average out of unbiased pieces: every survey that happened to fall low has been raised and every survey that fell high has been left alone. At a prevalence of one in a thousand, with half the corrected estimates negative, the truncated average is 3.77 times the truth. The corrected estimates themselves, negative ones included, average to the truth — that was the point of the correction — and they, or better the raw counts with each survey’s test characteristics, are what a pooled analysis should be built from. A negative estimate in a table is not an error to be tidied; it is the half of the sampling distribution that makes the other half honest.
Coverage that holds and limits that say little, summed over every count
At a true prevalence of 0.1% and a thousand people tested, the printed interval covers 95.17% of the time, reaches below zero in 97.95% of surveys and lies entirely below zero in 3.35%. The mapped exact interval covers 95.61%, and the one-sided upper limit holds 95.30% of the time at a typical value of 1.64%.
Truncating the exact interval at zero leaves its coverage unchanged and shortens its average width from 3.32 points to 1.91.
The typical upper limit reaches twice the truth at 184,521 people with a 95%-specific test, 39,957 at 99% and 5,083 with no false positives, all at 90% sensitivity.
Every coverage and share is an exact sum over the binomial distribution of the count, and every limit is an exact binomial limit, so nothing here is simulated except the thirty surveys drawn in the first figure. The survey sizes are found by bisection on the upper limit at the median count.
Not claimed: that the sensitivity and specificity are known. They are estimated in a validation study, and their uncertainty adds to the interval — by an amount the essay on the correction measured for the point estimate and which, for a rare condition, is dominated by the uncertainty in the specificity. Not claimed either that a Bayesian interval with a prior on the prevalence would behave like these; with a prior that puts its mass on non-negative values it cannot go below zero, and whether its upper limit is shorter because it is informative or because the prior is doing the work is a question about the prior.
Still open: the specificity the survey itself measures
The table of survey sizes treats the specificity as fixed and known, and it is the one quantity the whole problem turns on. Every survey that tests a population where the condition is rare is, among other things, a very large sample of unaffected people — a validation study for the specificity that nobody analyses as one.
A survey of 184,521 people with a prevalence of one in a thousand contains about 184,000 people without the condition, and their false-positive count estimates the specificity to a precision far better than most validation studies achieve. The difficulty is that the survey cannot say which of its positives are false, so the prevalence and the specificity are estimated from one count and are confounded exactly. Breaking the confounding needs a second piece of information — a confirmatory test on the positives, a known-negative subsample, or a prior on the prevalence — and which of those buys the most, per person tested, has not been worked out here.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A block size that changes — both name confidence interval, coverage, sample size
- A coverage table with its own error — both name confidence interval, coverage, sample size
- A schedule that reads the mean — both name confidence interval, coverage, sample size
- A tenth as wide, and both of them right — both name confidence interval, coverage, sample size
- An interval that covers and says nothing — both name clopper–pearson, confidence interval, coverage
- Robust is not free — both name confidence interval, coverage, sample size
Named objects
A flat tag is an object no other essay names yet.
Clopper–PearsonConfidence intervalCoverageMisclassificationPrevalenceSample sizeScreeningSensitivity and specificity