What a guaranteed minimum costs
Worth reading first: More data is not monotonically better.
A hole no sample size fills found that Wilson’s interval for a proportion covers as little as 83.82% at an expected count of 0.18, at every sample size. The interval that cannot do that is Clopper–Pearson’s, and it is usually described in one breath as exact and too conservative. The shortest interval is the one that misses priced its width against three alternatives. What it did not do is say what the width buys, or whether the same guarantee could be had for less.
Both questions have exact answers, because for a proportion the coverage of any interval is a finite sum.
Three numbers for an interval, not one
An interval for a proportion has no single coverage. Its coverage is a function of the true proportion, it jumps at every endpoint, and a summary has to choose what to keep. Three summaries answer three different questions:
- the worst coverage over every proportion, which is what a guarantee is about;
- the average coverage over a spread of proportions, which is what a long run of studies of different rates experiences;
- the average width, which is what the interval costs a reader.
At thirty trials, summed exactly over every count and averaged over a thousand proportions:
| interval | worst | average | average width |
|---|---|---|---|
| Wilson | 83.71% | 95.24% | 0.2708 |
| Agresti–Coull | 93.38% | 96.01% | 0.2788 |
| Jeffreys | 88.84% | 95.04% | 0.2687 |
| Clopper–Pearson | 95.05% | 97.34% | 0.2990 |
| Blaker | 95.00% | 96.31% | 0.2833 |
| mid-p | 92.45% | 95.82% | 0.2759 |
Only two rows have a worst coverage at or above 95%, and they are the two that are built by inverting an exact test rather than by approximating one. Every other interval, including the ones recommended in place of the textbook Wald, is below 95% somewhere, and two of them are below 90%.
The two that keep the guarantee pay for it in both of the other columns. Clopper–Pearson averages 2.34 points above its label and is 10.4% wider than Wilson; Blaker averages 1.31 points above and is 4.6% wider. So the guarantee itself costs something — no interval that keeps it averages 95% — and the two ways of keeping it cost very different amounts.
Two exact intervals, one sample space
Both intervals are built from the binomial distribution itself rather than from a normal approximation to it, and both contain every proportion that a 5% test would fail to reject. They differ in the test.
Clopper–Pearson inverts two one-sided tests. Its lower limit is the proportion at which seeing the observed count or more has probability 2.5%, and its upper limit is the proportion at which seeing the observed count or fewer has probability 2.5%. Each end is a one-sided exact bound in its own right.
Blaker inverts one two-sided test. A proportion is inside the interval if the observed count’s two-sided acceptability — its smaller tail probability plus the largest opposite tail that does not exceed it — is above 5%. The test does not split its 5% into two halves in advance; it lets the shape of the binomial decide how much of the 5% each side gets.
The figure shows the consequence. Clopper–Pearson’s curve lives above 97% over most of the range — at 60.6% of proportions at thirty trials — and touches 95% only at isolated points. Blaker’s lives nearer the line, above 97% at only 17.8% of proportions, and touches 95% more often. Wilson’s crosses the line in both directions.
Blaker’s interval is always inside Clopper–Pearson’s, for every count. For three successes in thirty trials the three intervals are:
| interval | lower | upper |
|---|---|---|
| Clopper–Pearson | 0.0211 | 0.2653 |
| Blaker | 0.0278 | 0.2596 |
| Wilson | 0.0346 | 0.2562 |
So a reader who wants the guarantee and less width has an interval to use that is not the one textbooks give. The question is what Clopper–Pearson’s extra width is buying, since it is not needed for the total.
What Clopper–Pearson’s extra width buys
The answer is visible only when the misses are split by side, which is the habit the essay on an interval’s two tails argues for.
| ten trials | worst chance of lying below the truth | worst chance of lying above it |
|---|---|---|
| Clopper–Pearson | 2.50% | 2.50% |
| Blaker | 5.00% | 5.00% |
| Wilson | 16.18% | 16.18% |
Clopper–Pearson holds each side under 2.5% at every proportion. Blaker holds only the sum under 5%, and at some proportions it puts the whole 5% on one side — at a proportion of 0.15 with ten trials, its interval lies above the truth 5.00% of the time and below it essentially never.
That is not a defect in Blaker’s interval. It is what inverting a two-sided test means: the test controls the total error and is indifferent to how it is divided, and at a proportion near the boundary the binomial’s skew makes one side much cheaper to spend on than the other. The interval spends where it is cheap and becomes narrower for it.
It is a defect for a reader who uses either end on its own. The side a bound is read from showed, for a mean, that a safety limit read off one end of a two-sided interval inherits whatever that end’s miss rate is, and that a repair judged on the total can leave that end twice as unsafe as promised. The same thing happens here in exact arithmetic. Blaker’s upper limit, used as a 97.5% upper bound, is exceeded up to 5.00% of the time. Clopper–Pearson’s upper limit, used the same way, is exceeded at most 2.50% of the time at every proportion, which is exactly the promise.
So Clopper–Pearson is not conservative about the thing it guarantees. It guarantees two one-sided statements, and at its worst each of them holds with no slack at all. It is conservative only about the total, because the two worst cases never happen at the same proportion.
The two worst cases never meet
The per-side guarantee explains the conservatism of the total in one observation: Clopper–Pearson’s two sides reach their worst at different proportions.
At thirty trials the chance that the interval lies below the truth comes within a few thousandths of a point of 2.5% at a proportion of 0.594. At that proportion the chance that it lies above the truth is 1.46%, so the total miss there is 3.96% and the coverage 96.04%. The mirror image happens at 0.406. Everywhere else both sides are below their limits at once.
A guarantee on each side therefore implies a total that is almost never 5%. Each side is allowed 2.5%, each side uses its allowance only at a handful of proportions, and the handfuls do not coincide. The average of 97.34% is the shadow of two tight one-sided guarantees on a discrete sample space rather than the sign of a loose two-sided one.
The same arithmetic, read as a test
Every interval here is the set of proportions a test does not reject, so the interval’s coverage at a proportion is one minus that test’s size when the null hypothesis is that proportion. The comparison above is therefore also a comparison of three tests of a proportion, and a test’s size is what a p-value’s flatness is about.
At thirty trials and a null proportion of 0.15, the test that goes with Clopper–Pearson — reject when either one-sided exact p-value is below 2.5% — rejects a true null 1.73% of the time. The test that goes with Blaker rejects it 3.54% of the time. Both are exact in the sense of never exceeding 5%, and the first spends barely a third of what it is allowed.
That unspent size is power given up. A test that could reject 5% of the time under the null and rejects 1.73% is, near the null, less able to reject when the null is false by the same margin. Blaker’s test recovers more than half the difference, and it does so by the same means as its interval: it lets the observed count’s skewed distribution decide which tail to spend on. At a null of one half, where the binomial is symmetric, the two tests are the same test and both reject 4.28% of the time — the discreteness alone, with nothing for Blaker to exploit.
The same unspent size is why combinations of discrete p-values are conservative before any dependence is involved, and why a randomised p-value restores exactness there too. The fix that section alludes to and this one keeps approaching is the same fix.
Near the boundary the guarantee is nearly free
The average width in the first table hides where the price is paid, and the hole in Wilson’s coverage suggests it might not be paid evenly.
Measured as the expected width in units of at three expected counts, at thirty trials:
| expected count | Wilson | Clopper–Pearson | Blaker |
|---|---|---|---|
| 0.2 | 3.682 | 3.797 | 3.503 |
| 1 | 4.655 | 4.951 | 4.611 |
| 5 | 7.684 | 8.501 | 7.997 |
At an expected count of one, Blaker’s interval is narrower than Wilson’s and covers 98.31% where Wilson covers 92.29%. At an expected count of 0.2 it is narrower still. The region where the approximate interval fails worst is the region where an exact interval with a guarantee costs least, and in part of it costs nothing: the guarantee and the width advantage come together.
Away from the boundary the order reverses. At an expected count of five, Clopper–Pearson is 10.6% wider than Wilson and Blaker 4.1% wider, and Wilson’s coverage there is 97.61% — above its label at that particular point. That is where the exact intervals’ reputation for conservatism was earned, and it is also where it matters least, because in the middle of the range the approximate intervals behave.
A p-value and an interval that disagree
The correspondence between an interval and a test has a practical edge. A common arrangement prints a Clopper–Pearson
interval beside a two-sided p-value computed a different way — by summing the probability of every count no more likely
than the one observed, which is Sterne’s construction and the one R’s binom.test uses for its p-value while reporting a
Clopper–Pearson interval.
The two can contradict each other on the same data. At thirty trials with a null proportion of 0.15, nine successes give a p-value of 0.0354 — significant at 5% — and a Clopper–Pearson interval from 0.147 to 0.494, which contains 0.15. The printout says the null is rejected and the interval says it is plausible. At a null of 0.30, four successes produce the same contradiction from the other side, with a p-value of 0.0471 and an interval from 0.038 to 0.307.
Blaker’s interval does not disagree with that p-value at any count at either null. It inverts a different two-sided test from Sterne’s — it adds the opposite tail rather than ranking every count by its probability — but both let the binomial’s own shape decide how the 5% is split, and at thirty trials the two agree on every count at both nulls. It is the equal split, not the exactness, that sets Clopper–Pearson apart from the p-value beside it. Neither disagreement is an error in the arithmetic. Each is two correct answers to two different questions printed as though they were one, and the interval with the per-side guarantee is the one that stands apart, for the reason the rest of this essay gives.
What lying above the truth means
The direction of Blaker’s one-sided misses near zero is worth stating plainly, because it is the same direction as the hole in Wilson’s interval.
An interval that lies entirely above the true proportion has a lower limit above it. Near the boundary that is the interval of a study that saw more events than the truth would usually produce, reporting a lower bound on the rate that is too high. For a harm it overstates the least the risk could be; for a benefit it overstates the least the effect could be. At ten trials and a true proportion of 0.15, Blaker’s interval does this 5.00% of the time, and Clopper–Pearson’s 0.99%.
Wilson’s hole was a lower limit too high as well, in a worse form: at an expected count below 0.18 every study that saw a single event reported a lower bound above the truth. Blaker’s version is the controlled one — never more than 5% on that side, at any proportion — and Clopper–Pearson’s is controlled at 2.5%. The three intervals are three answers to how often a study that finds something is allowed to overstate it.
The price, and how it falls with the sample
The first table was at thirty trials. At ten and at a hundred:
| trials | Clopper–Pearson average | Blaker average | Clopper–Pearson width over Wilson | Blaker width over Wilson |
|---|---|---|---|---|
| 10 | 98.37% | 97.34% | 16.8% | 9.3% |
| 30 | 97.34% | 96.31% | 10.4% | 4.6% |
| 100 | 96.47% | 95.78% | 6.0% | 2.9% |
Both prices fall, and Blaker’s falls faster in width. A guarantee at a hundred trials costs 6% of width in the textbook form and 2.9% in Blaker’s, and on both the average coverage is within one and a half points of 95%. The standard objection to exact intervals — that their conservatism makes them unattractive — was always an objection to a price, and at the sample sizes most studies of a proportion have, the price for Blaker’s form is a few per cent of width.
The worst coverage does not fall with the sample for either of them, and does not need to: it is at or above 95% at every sample size by construction, which the sums confirm at ten, thirty and a hundred trials.
Which guarantee to buy
The measurements make the choice depend on how the interval is read.
If both ends are read together — “the proportion is between these two values” — the guarantee that matters is the total, and Blaker’s interval gives it for less width than Clopper–Pearson’s everywhere and for less width than Wilson’s near the boundary. There is no measured reason to prefer Clopper–Pearson for a two-sided reading.
If either end is read on its own — a lower bound for a rate that must be at least something, an upper bound for a defect rate that must be at most something — the guarantee that matters is per side, and only Clopper–Pearson gives it. The extra width is the cost of that guarantee and is not conservatism.
If no guarantee is wanted and the average is what matters, Wilson and Jeffreys average within a quarter of a point of 95% at thirty trials, with the hole near the boundary that neither can close.
The point of the comparison is that “exact” names a family, not an interval. Two exact intervals built from the same binomial distribution make different promises, and the conventional complaint about one of them is a complaint about a promise the complainant was not reading it for.
What the sums establish, and what they do not
Neither exact interval covers less than 95% at any proportion drawn, at ten, thirty and a hundred trials, and Blaker’s lies inside Clopper–Pearson’s for every count at ten and thirty trials. The nesting is checked count by count rather than on average, because it is the property that makes Blaker a strict improvement for a two-sided reader.
Clopper–Pearson misses on neither side more than 2.5% of the time at any proportion drawn; Blaker misses no more than 5% in total and more than 3.5% on a single side somewhere. Stated together because the pair is the finding: the second interval’s saving is exactly the first interval’s per-side guarantee given up.
Blaker’s interval is narrower on average and less over-covered than Clopper–Pearson’s at every sample size measured.
Not claimed: that Blaker’s interval is the shortest interval with a 95% total guarantee. Other constructions — Sterne’s, and intervals that are not required to be connected — trade differently, and have not been counted here. The worst-side figures for Blaker at a hundred trials are read on a grid that can step past a narrow spike, which is why the claim is “more than 3.5%” rather than the 4.88% the finest search found.
Still open: an interval that is neither conservative nor holed
Every interval here either has a hole in its worst case or pays for not having one. Clopper–Pearson and Blaker cover above 95% on average because discreteness forces an interval that is never below the line to be above it most of the time. Wilson and Jeffreys average 95% because they are allowed below it.
The obvious question is whether anything can cover exactly 95% at every proportion. For a discrete count the answer is yes, and only one way: the interval has to depend on something besides the count, so that the sample space stops being discrete. That construction, what it costs, and why nobody uses it despite its being the only interval here that does what every interval claims, is the coin that makes it exact.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- An interval that covers and says nothing — both name clopper–pearson, conservative interval, coverage, interval width, wilson interval
- A ratio that changes between blocks — both name conservative interval, coverage, interval width
- The condition that cannot be dropped — both name conservative interval, coverage, interval width
- The shortest interval, and the one that does not move — both name coverage, discreteness, interval width
- What a two-unit study should report — both name conservative interval, coverage, interval width
- What the first stage does not know — both name conservative interval, exact test, interval width
Named objects
A flat tag is an object no other essay names yet.
Blaker intervalClopper–PearsonConservative intervalCoverageDiscretenessExact testInterval widthWilson interval