A hole no sample size fills
Worth reading first: More data is not monotonically better.
More data is not monotonically better held a proportion fixed at 0.15 and moved the sample size, and found the textbook interval’s coverage lurching by twelve points between nineteen observations and twenty. It found something reassuring about the better interval too: across the same sample sizes Wilson’s coverage never fell below 93.0%, and its worst single step was 3.6 points. The conclusion was that the oscillation cannot be removed but can be reduced to a wobble.
That conclusion was drawn at a proportion of 0.15. This essay moves the proportion instead, towards the boundary where most of the proportions anyone cares about actually live — adverse event rates, defect rates, conversion rates — and finds that the wobble has a hole in it.
The same dip at every sample size
Every coverage here is exact. For a given number of trials and true proportion , the coverage is the probability of every count whose interval contains , summed over the possible counts. No simulation, so no noise to mistake a real feature for.
The worst coverage of Wilson’s interval anywhere below a proportion of one half:
| trials | worst coverage | at an expected count of |
|---|---|---|
| 10 | 83.50% | 0.179 |
| 30 | 83.71% | 0.177 |
| 100 | 83.79% | 0.177 |
| 1,000 | 83.81% | 0.177 |
A hundredfold increase in the sample moves the worst coverage by three tenths of a point. It moves its location a great deal — from a proportion of 0.018 at ten trials to 0.00018 at a thousand — but the location in terms of the expected count, , does not move at all.
That is the first thing the hero figure shows, and it is why the horizontal axis is the expected count. Drawn against the proportion, the four dips would sit at four different places and look like four unrelated features. Drawn against , they are one feature that the sample size merely slides towards zero.
One success, and the interval built on it
The mechanism is visible one count at a time. Coverage at a proportion is the sum of the probabilities of the counts whose intervals contain it, so the question at any is which counts those are.
At a very small expected count almost every experiment sees zero successes, a few see one, and hardly any see two. The interval for zero successes runs from 0 up to about , so it covers any small proportion. The interval for one success is the one that matters, and its lower limit, as grows with held fixed, converges to the smaller root of
with . For any expected count below 0.1765, an experiment that sees exactly one success reports an interval that lies entirely above the truth. So at those proportions only zero successes covers, and the coverage is the probability of zero successes, which in the Poisson limit is .
Just below that is = 83.82%. Just above it, the interval for one success starts to cover again and the coverage jumps back by the probability of one success, about fifteen points. That jump is the hole’s far wall. At finite the probability of zero successes is , slightly smaller than , which is why ten trials read 83.50% and a thousand read 83.81%: the finite samples approach the limit from below and none of them passes it.
The missing coverage is an interval for one success whose lower limit sits at at ten trials, at a hundred and at a thousand. The interval’s upper end is generous — above in all three — and it is the lower end that is too high. A study that sees a single event and reports Wilson’s interval is claiming, with 95% confidence, that the rate is at least about 0.18 events per study-sized sample, and when the truth is somewhat below that the claim is false every time the single event occurs.
Why the usual picture does not show it
Coverage plots for a proportion are nearly always drawn with the proportion on the horizontal axis, from a little above zero to a little below one. This collection’s own coverage curve runs from 0.01 to 0.99, which is the conventional range.
At a hundred trials the hole is at a proportion of 0.0018, a fifth of the way from zero to the axis’s first point. At a thousand trials it is at 0.00018. Every conventional coverage plot of Wilson’s interval at a sample size above about twenty is drawn over a range that excludes the hole, and a reader comparing curves concludes — correctly for the range drawn — that Wilson hovers around 95%.
The omission is not careless. A linear axis from zero to one cannot show a feature whose width is a fraction of , and the proportion is the natural axis for someone thinking about a single study. But it is the same shape of failure the tail essay found in the standard picture of the central limit theorem: the picture is drawn at the scale where the bulk is visible, and the problem lives at a scale where it is not. The remedy is the same, too — change the axis to the one on which the feature has a fixed position. For a tail that was a logarithm of the probability; here it is the expected count.
One study, read through the hole
A trial enrols a thousand patients to look for a side effect whose true rate is 1.5 in ten thousand. The expected count is 0.15, just inside the hole.
Most such trials — 86.07% of them — see no events at all. They report Wilson’s interval for zero successes, which runs from zero to about 0.0038, and it contains the true rate. About 12.91% see exactly one event. They report Wilson’s interval for one success, which runs from 0.000177 to 0.00564, and its lower end is above the true rate of 0.00015. Every one of those intervals misses. The remaining trials see two or more events, and their intervals are further up still and miss as well.
So the procedure covers 86.07% of the time, and the whole of its shortfall is carried by the trials that found something. That is an uncomfortable place for an error to live. The trials that report a signal are the ones that get read, and they are exactly the ones whose intervals claim the rate is at least 1.8 in ten thousand when it is 1.5. Clopper–Pearson’s interval for one event in a thousand starts at 0.000025, and the repaired Wilson interval’s at 0.000051; both contain the truth.
This is also a small version of a large pattern in this collection. The winner’s curse is an estimate that is too large because it was selected for reaching a threshold; here the selection is done by the count itself, and the interval of the trial that saw the rare event is the one that overstates it.
A staircase behind it
The first dip is the deepest of a sequence, one for each count.
The interval for two successes has a lower limit that, in the same limit, converges to the smaller root of , and just below it only zero and one successes cover:
| the count that drops out | at an expected count of | coverage just below |
|---|---|---|
| 1 | 0.1765 | 83.82% |
| 2 | 0.5485 | 89.48% |
| 3 | 1.0203 | 91.59% |
| 4 | 1.5555 | 92.72% |
Each step is shallower, because by the time the count drops out the counts below it carry more of the probability, and the steps climb towards 95% without reaching it. Measured at a hundred trials the second dip is 89.47% at an expected count of 0.550, within a hundredth of a point of the limit in the table. The hero figure shows the first three as successive notches in the curve. Past an expected count of about five the notches have become the wobble the essay on oscillation described: at a hundred trials the worst coverage between expected counts of five and twelve is 92.68%, at 5.52 — which is roughly where the proportion of 0.15 in that essay sits once the sample passes thirty.
So the essay on oscillation was right about where it looked, and the conclusion does not transfer to the boundary. Within a few expected successes of zero the interval’s coverage is set by the first handful of counts, and it is not a wobble.
The same measurement for the other intervals
Wilson is not singled out because it is bad. It is the interval most often recommended in place of the textbook one, and it beats that one almost everywhere. The question is whether its alternatives share the hole.
| trials | Wilson | Jeffreys | Wilson, repaired near zero | Agresti–Coull |
|---|---|---|---|---|
| 10 | 83.50% | 86.81% | 91.58% | 92.39% |
| 100 | 83.79% | 88.01% | 92.68% | 93.90% |
| 1,000 | 83.81% | 89.77% | 92.72% | 94.36% |
Jeffreys’ interval has holes of its own, shallower than Wilson’s. At a thousand trials its worst coverage is 89.77%, at an expected count of 0.108 — a dip near the boundary of the same kind — and at ten, thirty and a hundred trials its worst is a second dip further in, at an expected count between two and two and a half, below 89% at thirty and a hundred. Under 90% at every size measured, and never the same place twice.
Agresti–Coull has no hole at the boundary. It adds two successes and two failures before building a Wald interval, which drags every lower limit down; at an expected count of 0.17 in a hundred trials it covers 99.93%, and at a hundred and a thousand trials its worst coverage is found at an expected count near forty rather than near zero. It pays with width — the whole point of the essay on the shortest interval is that width and coverage trade — and it is conservative near the boundary where Wilson is not.
The repaired Wilson interval is Brown, Cai and DasGupta’s modification, and it exists because of the hole. For one, two or three successes it replaces the lower limit with a value taken from the Poisson distribution — for one success rather than about — and leaves everything else alone. At an expected count of 0.17 in a hundred trials the unrepaired interval covers 84.35% and the repaired one covers 98.72%. The repair lifts the worst coverage by more than eight points at every sample size, and what is left is instructive: at a hundred trials the repaired interval’s worst is 92.68% at an expected count of 5.52, which is exactly the unrepaired interval’s own worst in that region. The repair removed the hole and left the wobble that was already there.
Where the hole is read
A hole at an expected count of 0.18 sounds like a corner case, and in terms of expected counts it is. In terms of the studies that meet it, it is not.
An expected count of 0.18 is a rate of 1.8 per ten thousand in a study of a thousand, or 1.8 per million in a study of a million. A safety study of a new drug, a quality audit of a production line, a screening programme’s yield in a low-prevalence region: each routinely reports an interval for a proportion whose true value is one or two expected events in the whole study. That is the region where Wilson’s interval covers 84%, 89% or 92% depending on which step the truth sits under, and it is the region where the difference between “the rate is at least” and “the rate could be as low as” decides whether the finding is reported.
The error also has a direction. Everything in the hole comes from lower limits that are too high: the interval for one success claims the rate is higher than it may be. For a harm, that overstates the lower bound on the risk; for a benefit, it overstates the lower bound on the effect. The upper limits are fine. A reader who uses only the upper end of a Wilson interval — “the rate is at most” — is not exposed to the hole at all, which is the same lesson as the side a bound is read from on a different interval: an interval’s two ends can be wrong separately, and the total hides which.
Why the sample size cannot help
The heading of the essay is literal, and it is worth saying why.
A larger sample moves every proportion’s expected count up. For a fixed true proportion, enough trials always carry the expected count past the staircase, and then the interval behaves. That is the sense in which the asymptotic argument that recommends Wilson is correct, and it is the sense in which more data is still better on average.
But the hole is not at a proportion. It is at an expected count, and for every sample size there is a proportion at that expected count — a rarer event for a larger study. No sample size is free of the hole; each sample size has its own proportion that falls in it. An interval procedure is used across studies of many sizes and many rates, and the worst coverage over all of them is the depth of the hole, which no single study’s size can change.
This is also why the right axis for thinking about an interval near the boundary is the number of events rather than the number of trials. The information about a rare proportion is carried by the count of events, and a study of a million that sees one event knows about as much about the rate as a study of a thousand that sees one event — which is to say, very little — and both of their Wilson intervals are wrong in the same way.
What the sums establish, and what they do not
Wilson’s worst coverage sits at the expected count where the interval for one success stops covering, at every sample size measured, within two hundredths of at ten, a hundred and a thousand trials, and its depth is within a point of and below 85% at all three. The two statements are made together because the first alone would be satisfied by any dip that happened to sit near 0.18, and the second is what ties it to the one-success mechanism.
The repair near zero lifts the worst coverage by more than five points at every sample size from ten trials to a thousand.
What does not survive is Wilson’s interval read as a 95% interval for a rare event. At an expected count of 0.17 in two hundred trials it covers 84.36%.
Every worst-coverage figure is the lowest value found by a search over a fine logarithmic grid of expected counts, refined around the grid’s minimum. Coverage jumps at every interval endpoint and the true infimum is approached from one side of a jump without being attained, so each reported worst is a value the coverage really takes, a hair above the infimum; the Poisson limit in the staircase table is the exact infimum as the sample grows.
Not claimed: that Wilson’s interval is a poor choice in general. Its average coverage over a uniform proportion is 95.24% at thirty trials and 95.09% at a hundred, closer to nominal than any interval with a guarantee, and away from the boundary it is what the essay on oscillation said it was. Nor is it claimed that the staircase’s later steps are measured to the same precision at every sample size: the table gives their Poisson limits and one finite-sample check.
Still open: what a guarantee near the boundary costs
The repaired interval and Agresti–Coull both close the hole by lowering the lower limits for small counts, and both pay in width. The interval that closes it completely — never below 95% at any proportion and any sample size — is Clopper–Pearson’s, which inverts two exact one-sided tests. It is usually dismissed as too conservative, and the dismissal is a statement about its average coverage, which is well above 95%.
What a guarantee on the worst case costs in average coverage and in width, whether a narrower interval with the same guarantee exists, and how the price changes near the boundary where the hole was, are the measurements what a guaranteed minimum costs makes. Behind that sits a sharper question: whether any interval can have neither the hole nor the conservatism, and what it would have to give up to manage it.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A coverage table with its own error — both name binomial proportion, coverage, discreteness
- An interval that covers and says nothing — both name binomial proportion, coverage, wilson interval
- The same draws for both methods — both name binomial proportion, coverage, discreteness
- The shortest interval, and the one that does not move — both name coverage, discreteness, jeffreys' prior
- A p-value that is not flat is not a p-value — both name coverage, discreteness
- A simulation that stops when it looks settled — both name binomial proportion, coverage
Named objects
A flat tag is an object no other essay names yet.
Agresti coullBinomial proportionCoverageDiscretenessJeffreys' priorPoisson limitRare eventsWilson interval