A proportion's interval near the boundary, and the coin

The coin that makes it exact

Every interval for a proportion either covers less than 95% somewhere or more than 95% on average, because a count is discrete. One construction covers exactly 95% at every proportion: it adds a uniform random draw to the count. At thirty trials it is 0.9% wider than Wilson's interval and narrower than both exact ones — and two analysts with the same data report different intervals, and one study in forty that sees nothing reports an empty one.

Worth reading first: More data is not monotonically better.

More data is not monotonically better concluded that the oscillation in a proportion’s coverage “does not disappear. It cannot: the underlying cause is that the sample space is discrete”. A hole no sample size fills found the worst of it, and what a guaranteed minimum costs found that the intervals which never fall below 95% average well above it. Between them they describe a trade that looks forced: under the line somewhere, or over it on average.

The trade is forced for any interval that is a function of the count alone. It is not forced for an interval that is allowed to use something else, and there is exactly one something else that makes the sample space continuous.

Clopper–Pearson, mid-p and the randomised interval: coverage across the proportion, 20 trialsClopper–Pearson never falls below 95% and runs up to 99.80%. The mid-p interval, which is the randomised interval with its coin fixed at one half, runs from 92.94% to 99.80%. The randomised interval covers 95% at every proportion, to within the 0.043% of the numerical integration over the coin.0.9000.9250.9500.97510.2000.4000.6000.800the true proportionactual coverage of a nominal 95% intervalClopper–Pearsonmid-prandomisedsummed exactly; the coin integrated at 1,000 pointsonly the coin is flat
Fig. 1 Coverage against the true proportion at twenty trials for three intervals built from the same exact tail probabilities: Clopper–Pearson, the mid-p interval, and the randomised interval, whose coverage is integrated over its coin. The first two oscillate; the third is the flat line at 95%. The slider is the number of trials.

Why no function of the count can be flat

An interval procedure for a proportion assigns an interval to each of the n+1n + 1 possible counts. Its coverage at a proportion pp is the total probability of the counts whose intervals contain pp. As pp moves, it crosses interval endpoints one at a time, and at each crossing a whole count’s probability enters or leaves the sum at once.

So the coverage is a sum of a changing set of binomial probabilities, and it jumps every time the set changes. To be exactly 95% at every pp the jumps would have to be zero, which would need counts with zero probability. No choice of endpoints can do it. Every interval built from the count alone has a coverage function with steps in it, and the only choice is where to put the steps relative to the 95% line — below it sometimes, like Wilson’s, or never below it and therefore mostly above, like Clopper–Pearson’s.

That argument has one loophole. It assumes the interval is a function of the count. If the interval also depends on a continuous quantity drawn independently of the data, the coverage at pp is an average over that quantity, and an average of step functions can be smooth.

The interval with a coin in it

Clopper–Pearson’s lower limit is the proportion at which the probability of seeing the observed count or more is 2.5%. Its upper limit is the proportion at which seeing the count or fewer is 2.5%. Both use the observed count’s own probability in full on each side, which is the reason each side is conservative.

The randomised interval splits that probability with a uniform random number UU. The lower limit is the pp at which

P(X>k)  +  UP(X=k)  =  0.025P(X > k) \;+\; U \cdot P(X = k) \;=\; 0.025

and the upper limit is the pp at which P(X<k)+(1U)P(X=k)=0.025P(X < k) + (1 - U)\cdot P(X = k) = 0.025. Under any true proportion, the quantity P(X>k)+UP(X=k)P(X > k) + U \cdot P(X = k) — evaluated at the observed count and the drawn UU — is exactly uniform on (0,1)(0, 1). That is the continuous analogue of the probability integral transform, and it is the reason a p-value for a continuous statistic is flat. Inverting an exactly uniform pivot gives an interval that covers exactly 95%, at every proportion and every sample size.

The hero figure checks it by summing over the counts and integrating over the coin at a thousand values: the randomised interval’s coverage at twenty trials stays within five hundredths of a point of 95% across the whole range, and the remaining wobble is the integration rule’s, not the method’s. Computed in closed form instead — the length of the stretch of UU that puts each count inside — it is 95% to fifteen digits at every proportion checked.

What the coin costs in width

The first thing to ask of a method that fixes coverage is what it spends to do it.

Average width of five 95% intervals for a proportion, 30 trials, with each one's worst coverage. Averaged over a uniform proportion. Wilson 0.2708, worst coverage 83.71%; randomised 0.2732, worst coverage 95.00%; mid-p 0.2759, worst coverage 92.45%; Blaker 0.2833, worst coverage 95.00%; Clopper–Pearson 0.2990, worst coverage 95.05%. The randomised interval covers exactly 95% at every proportion and is 0.9% wider than Wilson's.
Fig. 2 The average expected width of five 95% intervals at thirty trials, averaged over a uniform proportion, with each interval’s worst coverage beside it. The randomised interval’s width is averaged over its coin as well.
trials Wilson randomised mid-p Blaker Clopper–Pearson
10 0.4354 0.4488 0.4607 0.4760 0.5085
30 0.2708 0.2732 0.2759 0.2833 0.2990
100 0.1523 0.1526 0.1531 0.1567 0.1614

At thirty trials the randomised interval is 0.9% wider than Wilson’s and narrower than both intervals that guarantee their minimum. At a hundred trials the two widths differ by 0.0003, a fifth of a per cent. Wilson’s worst coverage at those sizes is 83.7% and 83.8%; the randomised interval’s is 95.00% at every proportion.

That is the whole of the discreteness problem priced in width. The conservatism of Clopper–Pearson — 10.4% wider than Wilson at thirty trials — is not the price of exactness. It is the price of exactness without a coin. With one, exact coverage costs about a hundredth of the width.

What the coin costs in everything else

If the randomised interval were free it would be in every textbook. It is in almost none, and the reasons are visible in one picture.

The randomised 95% interval for 3 successes in 20 trials, as a function of the coin. The same data give a lower end anywhere from 0.0564 to 0.0323 and an upper end from 0.3783 to 0.3187, depending on a uniform draw nobody observed. The coin at one half gives the mid-p interval, 0.0396 to 0.3561. Clopper–Pearson's 0.0321 to 0.3789 contains every one of them.
Fig. 3 The randomised 95% interval for three successes in twenty trials, as a function of the value the coin gave. The shaded band runs between the two ends. The dashed lines are Clopper–Pearson’s interval; the marked value of one half gives the mid-p interval.

Two analysts with the same data report different intervals. Three successes in twenty trials give a lower limit anywhere from 0.0321 to 0.0573 and an upper limit anywhere from 0.3170 to 0.3789, depending on a draw that has nothing to do with the data. The interval is exact as a procedure and arbitrary as a report. A reader shown one realisation of it cannot tell which part of the width is information and which part is the coin.

An interval can be empty. When nothing is observed, the upper limit needs (1U)P(X=0)(1 - U)\cdot P(X = 0) to reach 2.5% at some proportion, and since P(X=0)P(X = 0) is at most one, any coin above 0.975 leaves no proportion inside. Counted over four thousand values of the coin, the interval for zero successes in twenty trials is empty for 2.5% of them, and counted over two hundred values for each count from one to nineteen it is never empty. So one study in forty that sees no events reports that the proportion is nowhere. The coverage is still exactly 95%, because an empty interval that misses is one of the 5%.

The randomised 95% interval for 0 successes in 20 trials, as a function of the coin. The same data give a lower end anywhere from 0.0004 to 0.0000 and an upper end from 0.1677 to 0.0000, depending on a uniform draw nobody observed. For 0 successes the highest draws of the coin leave the interval empty — 2.5% of them. The coin at one half gives the mid-p interval, 0.0000 to 0.1391. Clopper–Pearson's 0.0000 to 0.1684 contains every one of them.
Fig. 4 The randomised interval for zero successes in twenty trials, as a function of the coin. Its lower end is at or near zero throughout; its upper end falls as the coin rises and reaches zero before the coin reaches one, where the interval is empty.

It violates a principle most readers hold without naming it: that two datasets with the same evidence should lead to the same conclusion. The randomised interval makes the conclusion depend on an event outside the data, which is exactly the property that makes its coverage exact and exactly the property a scientific report is not supposed to have.

The test that goes with it is the most powerful one

The interval’s partner is a test, and for the test there is a classical result that makes the coin harder to dismiss.

The Neyman–Pearson lemma says the most powerful test of one proportion against a larger one rejects for large counts — and, on a discrete sample space, rejects at the boundary count with a probability chosen to spend the level exactly. The randomised test is not a curiosity added to the theory; it is what the theory produces, and the non-randomised exact test is its rounding.

At twenty trials, testing a proportion of 0.1 against larger ones at 5%:

size power at 0.15 power at 0.20 power at 0.30
reject at five or more 4.32% 17.02% 37.04% 76.25%
also reject at four, with probability 0.076 5.00% 18.40% 38.69% 77.24%

The unrandomised test spends 4.32% of its 5%, and loses between one and one and a half points of power at each alternative for it. The coin is flipped only when exactly four successes are seen, and then it rejects about one time in thirteen. That is the entire difference between a test that is exactly at its level and one that is not, and it is the same difference, at the same count, as the one between Clopper–Pearson’s interval and the randomised interval’s.

This reframes the complaint. Nobody objects that the most powerful test is unfair to the data. They object that its result depends on a coin — which is a real objection, and a different one from the claim that the conservative test is the more careful choice. The conservative test is the less powerful one, and its extra caution is the unspent 0.68% of size, not a principled margin.

A coin before the data, and a coin after

The objection to the randomised interval is sharpest when it is put beside a coin the subject already accepts.

A randomised experiment assigns units to arms with a coin, and randomisation is not balance records how much ordinary variation that coin introduces into every estimate: two runs of the same trial on the same people give different answers because the coin fell differently. Nobody treats that as a defect of the analysis. The coin is accepted because it is flipped before the data exist, as part of producing them, and every estimate’s uncertainty already includes it.

The randomised interval’s coin is flipped after the data exist, as part of reading them. The data are fixed, the coin changes the conclusion, and the reader can see that it did. The mathematics of the two coins is the same — each turns a discrete or unbalanced structure into something with an exact reference distribution — and the objection is entirely about when in the process the randomness enters and whether it can be seen.

That distinction is defensible and worth holding. It is also worth seeing that it is a distinction about the legitimacy of a report, not about accuracy, and that the price of honouring it for a proportion is precisely the oscillation measured in the three essays before this one.

Reporting the coin instead of drawing it

The objection is to drawing the coin, not to the construction. The figure above already shows the way round it: instead of drawing UU and reporting one interval, report the whole band — for every proportion, the share of coin values for which it would be inside.

That object has a name, a fuzzy confidence interval, and it is a function from proportions to [0,1][0, 1] rather than a pair of endpoints. A proportion deep inside every realisation has membership one; a proportion outside every realisation has membership zero; a proportion in the shaded band’s edges has membership in between. It carries exactly the information of the randomised interval, it is reproducible, and it is never empty. It is also not a thing any reader knows how to use, and a report that says “the proportion is in this set to degree 0.63” has lost the plainness that made an interval worth printing.

The two objects are tied more closely than they look. Both of the randomised interval’s ends fall steadily as the coin’s value rises, so a proportion sitting exactly at the mid-p interval’s lower end is inside the randomised interval for every coin at or above one half and outside it for every coin below: its membership is exactly one half. The same holds at the upper end. The mid-p interval is the set of proportions whose membership in the fuzzy interval is at least one half — the fuzzy interval cut at its middle, which checking every count at ten and twenty trials against four hundred proportions confirms.

The practical compromise is to fix the coin at one half.

The coin fixed at one half

Setting U=12U = \tfrac{1}{2} gives the mid-p interval: each side counts half of the observed count’s probability. It is a function of the count again, so it is reproducible and never empty — and by the argument at the top of this essay its coverage has steps again.

trials mid-p worst coverage mid-p average coverage worst on one side
10 92.66% 96.83% 4.99%
30 92.45% 95.82% 4.92%
100 92.07% 95.29% 4.88%

The mid-p interval inherits the randomised interval’s centre and not its flatness. Its average is closer to 95% than any interval with a guarantee — within three tenths of a point at a hundred trials — and its worst is about three points below, far shallower than Wilson’s hole. On one side it can miss nearly 5% of the time, twice what a one-sided reader expects.

So the mid-p interval sits exactly where the comparison of guarantees would put it: no guarantee on the minimum, a good average, a moderate width. It is the coin’s expectation made into a function of the count, and it keeps the expectation’s virtues and gives up the coin’s one property that nothing else has.

Where the mid-p interval misses, split by side, 30 trials. The chance the whole interval lies below the truth and the chance it lies above, at every proportion. The worst below is 4.61% and the worst above is 4.61%.
Fig. 5 The mid-p interval at thirty trials, its misses split by side across the proportion. Neither side is held to 2.5%, and each approaches 5% at some proportion.

What the construction settles

The three essays before this one treated discreteness as a fixed cost to be allocated — below the line here, above it there. This one shows that the cost is real only for a report that must be a function of the count.

That changes what the comparisons mean. An interval that covers 97.34% on average is not being careful; it is being forced, by the refusal to depend on anything but the count, to overshoot where the count’s steps fall. An interval that covers 83.8% at its worst is not merely approximate; it is placing its steps on the wrong side of the line where the counts are fewest. And the gap between the two is not a gap between cautious and careless statistics. It is the width of a single step of the binomial distribution, and it can be closed at almost no cost in width by an operation — adding a coin — that almost every reader would regard as illegitimate.

The same resolution applies wherever the discreteness came from. A p-value from a discrete test is conservative for the same reason and becomes exactly flat with the same coin; a permutation test on a small sample has the same steps and the same fix; and a coverage study of any count-based interval, including the ones this collection counts, is measuring where the steps sit rather than whether they exist.

What the sums establish, and what they do not

The randomised interval’s coverage is 95% exactly at every proportion checked, computed in closed form, and integrating over the coin numerically agrees to within two tenths of a point at a coarse rule and five hundredths at a fine one. The two routes are kept together because the closed form alone could be a formula written to give 0.95, and the numerical route counts intervals.

The randomised interval is narrower on average than both intervals that guarantee their minimum, and within 4% of Wilson’s width, at ten, thirty and a hundred trials.

For zero successes the interval is empty for 2.5% of coin values, and for every interior count it is never empty, counted on a grid of four thousand values for the first and two hundred for each of the others.

The randomised test spends its level exactly and the unrandomised one does not. Both are sums over the binomial at a null of 0.1, and the table’s two rows differ only by the probability of rejecting at four successes; the powers are sums at each alternative, not simulations, so the one-and-a-half-point differences are the differences and carry no sampling error.

A caution on the width table. The randomised interval’s average width is averaged over forty values of the coin as well as over a thousand proportions, so its third decimal is good and its fourth is approximate; the comparisons drawn from it — narrower than Blaker and Clopper–Pearson, within a few per cent of Wilson — are several times larger than that error at every sample size.

What does not survive is the mid-p interval read as guaranteeing its level. It is the coin’s average, not the coin, and its worst coverage is 92.45% at thirty trials.

Not claimed: that the randomised interval should be reported. The measurements say what it would buy; the objection to reporting a result that depends on a random draw is a principle about evidence, and a coverage calculation cannot overrule it. Nor that the fuzzy interval is usable as it stands — whether a reader shown a membership function draws better conclusions than one shown an interval is a question about readers, not about sums.

Still open: the coin that is already there

A randomised report is objectionable because the randomness is added. Some data already contain a continuous quantity that is independent of the count and could play the coin’s part — the order in which events arrived, a continuous measurement taken alongside each binary one, the time to each event in a trial that also reports a proportion.

Whether breaking ties in the count with such a quantity delivers the randomised interval’s exact coverage, and whether it is defensible as evidence where the drawn coin is not, depends on the quantity being independent of the proportion given the count. That independence is checkable for a given design, the resulting coverage is an exact sum over the outcomes of that design, and neither has been done.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Clopper–PearsonCoverageDiscretenessFuzzy intervalMid-pRandomised intervalReproducibilityWilson interval