A proportion's interval near the boundary, and the coin

A hole no sample size fills

Wilson's interval is the recommended repair for a proportion, and away from the boundary it wobbles a point or two around 95%. Near zero it has a hole: at an expected count of 0.1765 its coverage is 83.50% at ten trials, 83.79% at a hundred and 83.81% at a thousand, and it never climbs past e to the minus 0.1765, which is 83.82%. The hole is where the interval built on one success stops containing the truth, and it belongs to the count rather than to the sample size.

Worth reading first: More data is not monotonically better.

More data is not monotonically better held a proportion fixed at 0.15 and moved the sample size, and found the textbook interval’s coverage lurching by twelve points between nineteen observations and twenty. It found something reassuring about the better interval too: across the same sample sizes Wilson’s coverage never fell below 93.0%, and its worst single step was 3.6 points. The conclusion was that the oscillation cannot be removed but can be reduced to a wobble.

That conclusion was drawn at a proportion of 0.15. This essay moves the proportion instead, towards the boundary where most of the proportions anyone cares about actually live — adverse event rates, defect rates, conversion rates — and finds that the wobble has a hole in it.

Coverage of the Wilson interval against the expected count, at 10, 30, 100 and 1,000 trialsRead against the expected number of successes the four sample sizes draw the same curve near the boundary. The worst coverage is 83.50% at n = 10, 83.71% at n = 30, 83.79% at n = 100, 83.81% at n = 1000, each at an expected count near 0.177, and the limiting depth is e^(−0.1765) = 83.82%.0.020.10.181100.8000.9001expected number of successes, npactual coverage of a nominal 95% intervalworst 83.50% at n = 10n = 10n = 30n = 100n = 1,000summed exactly over every countthe dip is a count, not a sample size
Fig. 1 Coverage of Wilson’s 95% interval against the expected number of successes, on a logarithmic axis, at ten, thirty, a hundred and a thousand trials. The four curves coincide near the boundary and share one dip. The slider changes the interval.

The same dip at every sample size

Every coverage here is exact. For a given number of trials nn and true proportion pp, the coverage is the probability of every count whose interval contains pp, summed over the n+1n + 1 possible counts. No simulation, so no noise to mistake a real feature for.

The worst coverage of Wilson’s interval anywhere below a proportion of one half:

trials worst coverage at an expected count of
10 83.50% 0.179
30 83.71% 0.177
100 83.79% 0.177
1,000 83.81% 0.177

A hundredfold increase in the sample moves the worst coverage by three tenths of a point. It moves its location a great deal — from a proportion of 0.018 at ten trials to 0.00018 at a thousand — but the location in terms of the expected count, npnp, does not move at all.

That is the first thing the hero figure shows, and it is why the horizontal axis is the expected count. Drawn against the proportion, the four dips would sit at four different places and look like four unrelated features. Drawn against npnp, they are one feature that the sample size merely slides towards zero.

One success, and the interval built on it

The mechanism is visible one count at a time. Coverage at a proportion is the sum of the probabilities of the counts whose intervals contain it, so the question at any pp is which counts those are.

Which counts' Wilson intervals contain the truth, at an expected count of 0.173 in 100 trials. The coverage is the probability of the counts whose intervals contain the true proportion. Just below an expected count of 0.1765, 0 covers, and the coverage is 84.10%.
Fig. 2 At an expected count just below the dip in a hundred trials: the probability of each count, marked by whether its Wilson interval contains the true proportion. Only the interval for zero successes does.

At a very small expected count almost every experiment sees zero successes, a few see one, and hardly any see two. The interval for zero successes runs from 0 up to about 3.84/n3.84/n, so it covers any small proportion. The interval for one success is the one that matters, and its lower limit, as nn grows with np=λnp = \lambda held fixed, converges to the smaller root of

(1λ)2=z2λλ1=1+z22z1+z24=0.1765(1 - \lambda)^2 = z^2 \lambda \quad\Longrightarrow\quad \lambda_1 = 1 + \frac{z^2}{2} - z\sqrt{1 + \frac{z^2}{4}} = 0.1765

with z=1.96z = 1.96. For any expected count below 0.1765, an experiment that sees exactly one success reports an interval that lies entirely above the truth. So at those proportions only zero successes covers, and the coverage is the probability of zero successes, which in the Poisson limit is eλe^{-\lambda}.

Just below λ1\lambda_1 that is e0.1765e^{-0.1765} = 83.82%. Just above it, the interval for one success starts to cover again and the coverage jumps back by the probability of one success, about fifteen points. That jump is the hole’s far wall. At finite nn the probability of zero successes is (1p)n(1 - p)^n, slightly smaller than enpe^{-np}, which is why ten trials read 83.50% and a thousand read 83.81%: the finite samples approach the limit from below and none of them passes it.

The missing coverage is an interval for one success whose lower limit sits at 0.1788/n0.1788/n at ten trials, 0.1767/n0.1767/n at a hundred and 0.1765/n0.1765/n at a thousand. The interval’s upper end is generous — above 4/n4/n in all three — and it is the lower end that is too high. A study that sees a single event and reports Wilson’s interval is claiming, with 95% confidence, that the rate is at least about 0.18 events per study-sized sample, and when the truth is somewhat below that the claim is false every time the single event occurs.

Why the usual picture does not show it

Coverage plots for a proportion are nearly always drawn with the proportion on the horizontal axis, from a little above zero to a little below one. This collection’s own coverage curve runs from 0.01 to 0.99, which is the conventional range.

Coverage of four nominal 95% intervals, n = 100. Computed exactly by summing over all 101 possible counts, not simulated. The Wald interval drops to 63.3% and is jagged everywhere; Clopper–Pearson never falls below 95% and pays for it in width.
Fig. 3 Coverage of four 95% intervals against the true proportion at a hundred trials, drawn over the conventional range from 0.01 to 0.99. Wilson’s curve looks well behaved from end to end, because its hole at this sample size sits at a proportion of 0.0018, off the left edge of the axis.

At a hundred trials the hole is at a proportion of 0.0018, a fifth of the way from zero to the axis’s first point. At a thousand trials it is at 0.00018. Every conventional coverage plot of Wilson’s interval at a sample size above about twenty is drawn over a range that excludes the hole, and a reader comparing curves concludes — correctly for the range drawn — that Wilson hovers around 95%.

The omission is not careless. A linear axis from zero to one cannot show a feature whose width is a fraction of 1/n1/n, and the proportion is the natural axis for someone thinking about a single study. But it is the same shape of failure the tail essay found in the standard picture of the central limit theorem: the picture is drawn at the scale where the bulk is visible, and the problem lives at a scale where it is not. The remedy is the same, too — change the axis to the one on which the feature has a fixed position. For a tail that was a logarithm of the probability; here it is the expected count.

One study, read through the hole

A trial enrols a thousand patients to look for a side effect whose true rate is 1.5 in ten thousand. The expected count is 0.15, just inside the hole.

Most such trials — 86.07% of them — see no events at all. They report Wilson’s interval for zero successes, which runs from zero to about 0.0038, and it contains the true rate. About 12.91% see exactly one event. They report Wilson’s interval for one success, which runs from 0.000177 to 0.00564, and its lower end is above the true rate of 0.00015. Every one of those intervals misses. The remaining trials see two or more events, and their intervals are further up still and miss as well.

So the procedure covers 86.07% of the time, and the whole of its shortfall is carried by the trials that found something. That is an uncomfortable place for an error to live. The trials that report a signal are the ones that get read, and they are exactly the ones whose intervals claim the rate is at least 1.8 in ten thousand when it is 1.5. Clopper–Pearson’s interval for one event in a thousand starts at 0.000025, and the repaired Wilson interval’s at 0.000051; both contain the truth.

This is also a small version of a large pattern in this collection. The winner’s curse is an estimate that is too large because it was selected for reaching a threshold; here the selection is done by the count itself, and the interval of the trial that saw the rare event is the one that overstates it.

A staircase behind it

The first dip is the deepest of a sequence, one for each count.

The interval for two successes has a lower limit that, in the same limit, converges to the smaller root of (2λ)2=z2λ(2 - \lambda)^2 = z^2\lambda, and just below it only zero and one successes cover:

the count that drops out at an expected count of coverage just below
1 0.1765 83.82%
2 0.5485 89.48%
3 1.0203 91.59%
4 1.5555 92.72%

Each step is shallower, because by the time the count xx drops out the counts below it carry more of the probability, and the steps climb towards 95% without reaching it. Measured at a hundred trials the second dip is 89.47% at an expected count of 0.550, within a hundredth of a point of the limit in the table. The hero figure shows the first three as successive notches in the curve. Past an expected count of about five the notches have become the wobble the essay on oscillation described: at a hundred trials the worst coverage between expected counts of five and twelve is 92.68%, at 5.52 — which is roughly where the proportion of 0.15 in that essay sits once the sample passes thirty.

So the essay on oscillation was right about where it looked, and the conclusion does not transfer to the boundary. Within a few expected successes of zero the interval’s coverage is set by the first handful of counts, and it is not a wobble.

The same measurement for the other intervals

Wilson is not singled out because it is bad. It is the interval most often recommended in place of the textbook one, and it beats that one almost everywhere. The question is whether its alternatives share the hole.

The worst coverage of four nominal 95% intervals, from 10 to 1,000 trials. More trials do not bring the worst coverage of the first three to 95%. At 1,000 trials: Wilson 83.81%, Jeffreys 89.77%, Wilson, repaired near zero 92.72%, Agresti–Coull 94.36%. Wilson's approaches e^(−0.1765) = 83.82% and stops there.
Fig. 4 The worst coverage over every proportion for four 95% intervals, from ten trials to a thousand. The dashed rule is the limiting depth of Wilson’s dip.
trials Wilson Jeffreys Wilson, repaired near zero Agresti–Coull
10 83.50% 86.81% 91.58% 92.39%
100 83.79% 88.01% 92.68% 93.90%
1,000 83.81% 89.77% 92.72% 94.36%

Jeffreys’ interval has holes of its own, shallower than Wilson’s. At a thousand trials its worst coverage is 89.77%, at an expected count of 0.108 — a dip near the boundary of the same kind — and at ten, thirty and a hundred trials its worst is a second dip further in, at an expected count between two and two and a half, below 89% at thirty and a hundred. Under 90% at every size measured, and never the same place twice.

Agresti–Coull has no hole at the boundary. It adds two successes and two failures before building a Wald interval, which drags every lower limit down; at an expected count of 0.17 in a hundred trials it covers 99.93%, and at a hundred and a thousand trials its worst coverage is found at an expected count near forty rather than near zero. It pays with width — the whole point of the essay on the shortest interval is that width and coverage trade — and it is conservative near the boundary where Wilson is not.

The repaired Wilson interval is Brown, Cai and DasGupta’s modification, and it exists because of the hole. For one, two or three successes it replaces the lower limit with a value taken from the Poisson distribution — 0.0513/n0.0513/n for one success rather than about 0.1765/n0.1765/n — and leaves everything else alone. At an expected count of 0.17 in a hundred trials the unrepaired interval covers 84.35% and the repaired one covers 98.72%. The repair lifts the worst coverage by more than eight points at every sample size, and what is left is instructive: at a hundred trials the repaired interval’s worst is 92.68% at an expected count of 5.52, which is exactly the unrepaired interval’s own worst in that region. The repair removed the hole and left the wobble that was already there.

Coverage of the Jeffreys interval against the expected count, at 10, 30, 100 and 1,000 trials. Read against the expected number of successes the four sample sizes draw the same curve near the boundary. The worst coverage is 86.81% at n = 10, 88.84% at n = 30, 88.01% at n = 100, 89.77% at n = 1000.
Fig. 5 The same four sample sizes for Jeffreys’ interval. Its dips are shallower than Wilson’s, and there are two that matter — one near an expected count of a tenth, one near two — rather than one.

Where the hole is read

A hole at an expected count of 0.18 sounds like a corner case, and in terms of expected counts it is. In terms of the studies that meet it, it is not.

An expected count of 0.18 is a rate of 1.8 per ten thousand in a study of a thousand, or 1.8 per million in a study of a million. A safety study of a new drug, a quality audit of a production line, a screening programme’s yield in a low-prevalence region: each routinely reports an interval for a proportion whose true value is one or two expected events in the whole study. That is the region where Wilson’s interval covers 84%, 89% or 92% depending on which step the truth sits under, and it is the region where the difference between “the rate is at least” and “the rate could be as low as” decides whether the finding is reported.

The error also has a direction. Everything in the hole comes from lower limits that are too high: the interval for one success claims the rate is higher than it may be. For a harm, that overstates the lower bound on the risk; for a benefit, it overstates the lower bound on the effect. The upper limits are fine. A reader who uses only the upper end of a Wilson interval — “the rate is at most” — is not exposed to the hole at all, which is the same lesson as the side a bound is read from on a different interval: an interval’s two ends can be wrong separately, and the total hides which.

Why the sample size cannot help

The heading of the essay is literal, and it is worth saying why.

A larger sample moves every proportion’s expected count up. For a fixed true proportion, enough trials always carry the expected count past the staircase, and then the interval behaves. That is the sense in which the asymptotic argument that recommends Wilson is correct, and it is the sense in which more data is still better on average.

But the hole is not at a proportion. It is at an expected count, and for every sample size there is a proportion at that expected count — a rarer event for a larger study. No sample size is free of the hole; each sample size has its own proportion that falls in it. An interval procedure is used across studies of many sizes and many rates, and the worst coverage over all of them is the depth of the hole, which no single study’s size can change.

This is also why the right axis for thinking about an interval near the boundary is the number of events rather than the number of trials. The information about a rare proportion is carried by the count of events, and a study of a million that sees one event knows about as much about the rate as a study of a thousand that sees one event — which is to say, very little — and both of their Wilson intervals are wrong in the same way.

What the sums establish, and what they do not

Wilson’s worst coverage sits at the expected count where the interval for one success stops covering, at every sample size measured, within two hundredths of λ1\lambda_1 at ten, a hundred and a thousand trials, and its depth is within a point of eλ1e^{-\lambda_1} and below 85% at all three. The two statements are made together because the first alone would be satisfied by any dip that happened to sit near 0.18, and the second is what ties it to the one-success mechanism.

The repair near zero lifts the worst coverage by more than five points at every sample size from ten trials to a thousand.

What does not survive is Wilson’s interval read as a 95% interval for a rare event. At an expected count of 0.17 in two hundred trials it covers 84.36%.

Every worst-coverage figure is the lowest value found by a search over a fine logarithmic grid of expected counts, refined around the grid’s minimum. Coverage jumps at every interval endpoint and the true infimum is approached from one side of a jump without being attained, so each reported worst is a value the coverage really takes, a hair above the infimum; the Poisson limit in the staircase table is the exact infimum as the sample grows.

Not claimed: that Wilson’s interval is a poor choice in general. Its average coverage over a uniform proportion is 95.24% at thirty trials and 95.09% at a hundred, closer to nominal than any interval with a guarantee, and away from the boundary it is what the essay on oscillation said it was. Nor is it claimed that the staircase’s later steps are measured to the same precision at every sample size: the table gives their Poisson limits and one finite-sample check.

Still open: what a guarantee near the boundary costs

The repaired interval and Agresti–Coull both close the hole by lowering the lower limits for small counts, and both pay in width. The interval that closes it completely — never below 95% at any proportion and any sample size — is Clopper–Pearson’s, which inverts two exact one-sided tests. It is usually dismissed as too conservative, and the dismissal is a statement about its average coverage, which is well above 95%.

What a guarantee on the worst case costs in average coverage and in width, whether a narrower interval with the same guarantee exists, and how the price changes near the boundary where the hole was, are the measurements what a guaranteed minimum costs makes. Behind that sits a sharper question: whether any interval can have neither the hole nor the conservatism, and what it would have to give up to manage it.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Agresti coullBinomial proportionCoverageDiscretenessJeffreys' priorPoisson limitRare eventsWilson interval