A count that bets against its interval
Worth reading first: What the 95% refers to.
An interval that covers and says nothing ended on the question every reader of a study eventually asks. All of its numbers — average coverage, worst-case coverage, expected width — describe a procedure, averaged over samples that did not happen. A reader has one sample and one interval. Is there anything that can be said about that interval beyond “the procedure that made it covers 95% of the time”?
The first answer is no, and it is worth seeing why before looking for a better one. Fix the count a study observed. The interval built from it is fixed too, and the true proportion either lies inside it or does not. Conditional on the exact count, coverage is one or zero, and which of the two is the very thing the study does not know. Conditioning all the way down to the observation leaves no frequency to speak of.
The second answer is more useful. Conditioning on a set of counts does leave a frequency: among all studies whose count lands in the set, the interval covers some share of the time, and that share can be computed exactly at every true proportion. If a set exists on which that share is below 95% at every proportion, then a reader whose count lands in it knows the label is wrong for studies like theirs, by an amount they can state, without knowing the proportion at all. If no such set exists, the label survives everything the count can say against it.
A bet that wins whatever the proportion is
Buehler put the question as a bet, which is the cleanest way to hold it. An opponent sees the count but not the proportion and offers to bet that the interval misses, at the odds a 95% label implies — nineteen to one on covering. The opponent may decline to bet on some counts and bet on others. If the opponent has a rule for choosing which counts to bet on that wins in expectation at every true proportion, the label is wrong for the counts in that rule, and wrong in a way the data reveal. A set with that property is a relevant subset.
The textbook interval at forty trials has one, and it is the first place anybody would look.
Take the counts of at most two or at least thirty-eight. Among studies whose count lands there, the textbook interval covers at most 88.96% at any true proportion, and at most proportions far less. Widen the set to counts of at most five or at least thirty-five and the ceiling rises to 94.90% — still below the label, at every proportion. Widen it again to eight and the curve pokes above, to 95.57% near a proportion of 0.12, and the bet no longer wins everywhere.
The set is not rare. At a true proportion of 0.1 a count of five or fewer turns up in 79.4% of studies, and at 0.05 in 98.6% of them. A study of an uncommon event out of forty, which is a common kind of study, nearly always produces a count in the set on which the label is demonstrably wrong.
The middle stretch of the figure, where the coverage is zero, is not a defect. When the proportion is near one half, a count of five is improbable and its interval does not reach anywhere near one half, so the few studies that land in the tails there all miss. The bet is conditional, and the conditioning is what makes it win: wherever the tails are likely, they cover less than 95%, and wherever they cover well they are unlikely.
Why the textbook interval loses on its own tails
The textbook interval is , and its fault at the ends of the range is well known: the plug-in standard error goes to zero as the count does. At a count of zero the interval is the single point zero, which covers no proportion except zero itself. Every study that sees no events reports a certainty it does not have.
That one count is enough to anchor the bet, and the mechanism is clearest if the zero count is removed. Take the counts from one to five and from thirty-five to thirty-nine and leave zero and forty out. At very small proportions nearly every study in that set observes exactly one event, and the textbook interval for a count of one — from zero to about 0.073 — contains every proportion that small. Conditional on the set, coverage climbs to one as the proportion goes to zero, and the bet fails there. Put the zero count back and the picture reverses: at very small proportions the conditional distribution sits almost entirely on zero, whose interval misses, so coverage is held near zero exactly where the other counts would have lifted it.
The counts from one to five then do the rest. Their intervals are too short, because a standard error computed from is computed from an estimate that is itself small and noisy, and an interval centred on a small count with a width set by that count misses the proportions just above it. Where the zero count stops dominating, these intervals take over the job of holding the coverage down. The ceiling of 94.90% is set near a proportion of 0.073, the edge of the count-of-one interval, where the zero count’s grip has loosened and the short intervals have not yet been joined by enough longer ones.
The width of the widest losing set grows with the sample: four counts at each end out of twenty, five out of forty, seven out of a hundred. That is the opposite of what “the problem goes away with more data” would predict, and it fits what the essay on sample sizes found about coverage in general — the textbook interval’s failures at the ends do not shrink so much as move. At a hundred trials the losing set holds more counts, and the narrowest of them — a count of at most one — covers at most 75.25%, against 76.01% at forty and 77.30% at twenty. The textbook interval’s label is refutable from the data at every sample size, and the refuting set does not narrow.
The score interval has no such set
Wilson’s interval inverts the score test rather than plugging in , and at a count of zero it runs from zero to about 0.088 instead of stopping at a point. The same tail bets fail against it at once. On every tail set from one count at each end to ten, the conditional coverage reaches at least 99.31% somewhere, and on most of them it reaches one: the zero count’s interval covers every small proportion, so at small proportions the tails cover nearly always.
The shape is the same as the textbook interval’s — near zero in the middle of the range, where the tails are improbable and their intervals stop short of one half — and the difference is entirely at the ends. A bet needs the curve under 95% at every proportion, and it is the extreme proportions, where the tails are almost certain to occur, that decide whether it is. There the score interval is generous and the textbook interval is a point.
But the tails are only the first shape an opponent would try. There are about sets of counts out of forty, and an opponent may also bet at a stake that varies from count to count — a randomised bet — which is a continuum of strategies. Searching them is hopeless, and a search that found nothing would prove nothing. What settles the question is a duality.
Suppose some prior on the proportion makes every interval in the family at least 95% credible: for every one of the forty-one counts, the posterior probability that the interval built from that count contains the proportion is at least 95%. Then no bet can win at every proportion. A bet that won everywhere would win on average under that prior too, and averaged over the prior a bet on any count is a bet at the posterior’s odds, which favour covering by more than nineteen to one. The converse holds as well, and it is the part that makes this a method rather than a sufficient condition: if no such prior exists, some bet — perhaps randomised — wins everywhere. This is Gordan’s alternative for a finite system of linear inequalities, applied to the forty-one counts, and it turns the search over bets into a search over priors, which is a linear programme.
The programme finds a prior under which every score interval at forty trials is 96.46% credible — every one of them, the same number for each, because the best prior for this purpose is the one that leaves no interval weaker than the rest. It is a lumpy thing: weight on forty-one separate proportions, more in the middle than at the ends, nothing a statistician would write down as a belief. That does not matter. The prior is not offered as anybody’s opinion. It is a certificate, and its only job is to exist.
The obvious prior does not do the job. Under a flat prior, the score interval for a count of twenty is 94.89% credible, slightly under the label, which on its own says nothing about bets: the flat prior is one prior among many, and the certificate needs only one. The textbook interval has no certificate at all, and the figure shows why directly. Under the prior that vouches for every score interval, the textbook interval for a count of zero is 0.00% credible, and no prior on the open interval from zero to one can do better for an interval that is a single point.
A worst case that no count can find
There is something here that looks like a contradiction. The score interval does not cover 95% at every proportion. Its coverage dips below the label at many of them, by a few points across the range most studies live in — to 92.21% at forty trials over proportions from 0.02 to 0.98, as the earlier essay counted — and further at proportions extremely close to zero or one. Somewhere the label is wrong. Why can no reader bet on it?
Because the dips live at proportions, and a bet has to be placed on counts. A dip at a particular proportion is caused by a particular count’s interval stopping just short of it, and the counts that do that at one proportion are different from the counts that do it at the next. Any set of counts chosen to catch the dips near one proportion carries counts whose intervals are generous somewhere else, and the conditional coverage goes above 95% there. For the textbook interval the failure has the opposite structure: the same counts — the small ones — fail across a whole stretch of proportions at once, because their intervals are short for a reason that does not depend on the proportion. That is what a bet can exploit.
So worst-case coverage and refutability by the count are different properties, and neither implies the other. An interval can dip below its label somewhere and still be unrefutable by any reader, as the score interval does. And an interval could in principle cover at least 95% everywhere and still be refutable from the other side — a reader could know their interval covers more than the label claims. That is what happens to the exact interval.
Two numbers that bracket every conditional claim
The duality gives two numbers for each interval rather than one. The upper bound is the most a prior can vouch for every interval in the family at once; a bet that the interval misses more than the label says exists exactly when 95% is above it. The lower bound is the least a prior can hold every interval to; a bet that the interval covers more than the label says exists exactly when 95% is below it. An interval whose two bounds straddle 95% has a label no count can contradict from either side. Each bound is the exact value of a matrix game, found by bisection.
At forty trials the score interval’s bounds are 94.09% and 96.46%. Read as bets, they say that a reader can find a class of counts on which the score interval covers at most about 96.5% at every proportion, and a class on which it covers at least about 94% at every proportion, and nothing sharper than that in either direction. Agresti and Coull’s interval — Wald’s, after adding two successes and two failures — sits at 94.55% and 96.59%, and it has a certificate too. Both straddle the label at every size measured, and both bands narrow towards 95% as the trials increase, to 94.76% and 95.56% for the score interval at a hundred.
The exact interval of Clopper and Pearson is the other kind. Its bounds at forty trials are 96.23% and 97.74%, both above the label. A reader holding a Clopper–Pearson interval can bet that it covers more than 95% and win at every proportion, and the simplest such bet is to bet on every count: the interval’s coverage never falls below 95.03% at forty trials, at any proportion. That is the familiar fact that the exact interval is conservative, restated as something a reader can act on. The lower bound adds information the familiar fact does not: no class of counts can be shown to cover more than about 96.2% everywhere, so the conservatism a reader can actually claim is bounded too.
The textbook interval’s bounds are the most lopsided of the four: an upper bound of 0.00%, which is the count of zero again, and a lower bound of 94.30%. A reader cannot bet that the textbook interval covers more than 95% on any class of counts; a reader can bet that it covers less, and on the counts at the ends the bet is very good.
What a reader holding one interval can say
This is the answer the earlier essay was asking for, and it is narrower and more useful than “the procedure covers 95%”.
A reader with a textbook interval and a count of five or fewer out of forty can say that their interval belongs to a class on which it covers at most 94.90% whatever the proportion is — and at two or fewer, at most 88.96%. That is a statement about this study and studies like it, it has a frequency interpretation, and it does not depend on any prior. It is also a statement that the interval should not have been reported.
A reader with a score interval can say that nothing in their count contradicts the 95%, from either side, and that the interval is the kind of object a coherent reader with some prior could also call 96.46% credible. That second half is the frequency reading and the Bayesian one meeting at a single interval. What a credible interval covers found that a Bayesian interval can be scored as a frequentist procedure; this runs the comparison the other way, and finds that a frequentist interval with no relevant subsets is one some Bayesian could have issued.
A reader with a Clopper–Pearson interval can say more than the label does: at least 95.03% at forty trials, whatever the proportion, and they can point to the reason — the interval is built to hold each tail at 2.5% separately, and the discreteness of the count means both tails are usually held well below that. Whether that is a virtue depends on what the extra width costs, which the coin that makes it exact prices directly by removing it.
None of this makes the conditional answer the only answer. Marginal coverage and conditional coverage are different promises, and a procedure can keep one and break the other in ways a regulator, a reviewer and a reader would weigh differently. What changes here is that the conditional promise has a precise meaning for a single count — no recognisable class of counts on which the label is wrong — and that meaning can be checked by solving one linear programme rather than argued.
How far the certificates reach
The certificates are exact for what they certify. A prior on 240 support points is a real prior, every posterior under it is a finite sum, and the forty-one posteriors in the figure are computed, not estimated. When the programme returns a prior under which every score interval is 96.46% credible, no bet against the score interval at 95% exists, between the support points or anywhere else, because the argument from the prior does not care where the proportion is.
The bets are less complete. The randomised bet the programme returns is guaranteed to win at the 240 support points. Between them, and especially within a few thousandths of zero and one, a bet chosen this way can lose, and the computed lower bounds are therefore the best bet on that support, which is close to but not the same as the best bet on the whole interval. The verdicts that matter do not lean on this. The textbook interval’s refutation rests on a tail set whose conditional coverage is computed exactly at every proportion and at every interval endpoint, and the exact interval’s over-coverage rests on its own unconditional minimum of 95.03%, computed the same way.
Two other limits are worth stating. The sample sizes run from ten to a hundred; the score interval’s upper bound falls from 98.39% at ten to 95.56% at a hundred and its lower bound rises towards it, so the band is closing on the label, but nothing here says whether it ever crosses for some larger n. And every bound is for a two-sided interval at 95%; a one-sided bound, or a different level, is a different family of intervals with its own programme.
Still open: a bet against an interval that uses a coin
Every interval here is a function of the count alone, and that is what makes the bet a bet on counts. Some intervals for a proportion are not. The coin that makes it exact adds a uniform random draw to the count and reaches exactly 95% at every proportion, and the coin that is already there reads that draw from the order of the successes instead of a random number generator. A reader of such an interval sees the count and the coin, and can bet on both.
The question is whether exact coverage at every proportion survives that. Marginally it does by construction. But the coin is observed, and a set defined by the count and the coin — counts of zero with a coin near one, say — is a set an opponent can recognise. Whether a randomised interval that is exact marginally has a relevant subset once its coin is visible, how large the bet against it can be, and whether a reader of the clock’s coin faces the same problem, is the measurement this leaves. The computation is the same programme with the coin’s values added to the counts as things a bet can depend on, and what a prior is worth suggests the certificate, if one exists, will again be a prior with a stated size.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A prior wrong about its own spread — both name binomial proportion, coverage, credible interval, exact enumeration, prior
- The interval at the end of the curve — both name binomial proportion, conditional coverage, confidence interval, coverage, wald interval
- The interval that integrates — both name coverage, credible interval, frequentist interpretation, prior
- The same draws for both methods — both name binomial proportion, confidence interval, coverage, wald interval
- Where the derivative is zero — both name binomial proportion, confidence interval, coverage, exact enumeration
- Where the two schools agree — both name confidence interval, coverage, credible interval, prior
Named objects
A flat tag is an object no other essay names yet.
Binomial proportionClopper–PearsonConditional coverageConditional inferenceConfidence intervalCoverageCredible intervalExact enumerationFrequentist interpretationLeast favourable priorPriorRelevant subsetWald intervalWilson intervalWorst case