The coin that makes it exact
Worth reading first: More data is not monotonically better.
More data is not monotonically better concluded that the oscillation in a proportion’s coverage “does not disappear. It cannot: the underlying cause is that the sample space is discrete”. A hole no sample size fills found the worst of it, and what a guaranteed minimum costs found that the intervals which never fall below 95% average well above it. Between them they describe a trade that looks forced: under the line somewhere, or over it on average.
The trade is forced for any interval that is a function of the count alone. It is not forced for an interval that is allowed to use something else, and there is exactly one something else that makes the sample space continuous.
Why no function of the count can be flat
An interval procedure for a proportion assigns an interval to each of the possible counts. Its coverage at a proportion is the total probability of the counts whose intervals contain . As moves, it crosses interval endpoints one at a time, and at each crossing a whole count’s probability enters or leaves the sum at once.
So the coverage is a sum of a changing set of binomial probabilities, and it jumps every time the set changes. To be exactly 95% at every the jumps would have to be zero, which would need counts with zero probability. No choice of endpoints can do it. Every interval built from the count alone has a coverage function with steps in it, and the only choice is where to put the steps relative to the 95% line — below it sometimes, like Wilson’s, or never below it and therefore mostly above, like Clopper–Pearson’s.
That argument has one loophole. It assumes the interval is a function of the count. If the interval also depends on a continuous quantity drawn independently of the data, the coverage at is an average over that quantity, and an average of step functions can be smooth.
The interval with a coin in it
Clopper–Pearson’s lower limit is the proportion at which the probability of seeing the observed count or more is 2.5%. Its upper limit is the proportion at which seeing the count or fewer is 2.5%. Both use the observed count’s own probability in full on each side, which is the reason each side is conservative.
The randomised interval splits that probability with a uniform random number . The lower limit is the at which
and the upper limit is the at which . Under any true proportion, the quantity — evaluated at the observed count and the drawn — is exactly uniform on . That is the continuous analogue of the probability integral transform, and it is the reason a p-value for a continuous statistic is flat. Inverting an exactly uniform pivot gives an interval that covers exactly 95%, at every proportion and every sample size.
The hero figure checks it by summing over the counts and integrating over the coin at a thousand values: the randomised interval’s coverage at twenty trials stays within five hundredths of a point of 95% across the whole range, and the remaining wobble is the integration rule’s, not the method’s. Computed in closed form instead — the length of the stretch of that puts each count inside — it is 95% to fifteen digits at every proportion checked.
What the coin costs in width
The first thing to ask of a method that fixes coverage is what it spends to do it.
| trials | Wilson | randomised | mid-p | Blaker | Clopper–Pearson |
|---|---|---|---|---|---|
| 10 | 0.4354 | 0.4488 | 0.4607 | 0.4760 | 0.5085 |
| 30 | 0.2708 | 0.2732 | 0.2759 | 0.2833 | 0.2990 |
| 100 | 0.1523 | 0.1526 | 0.1531 | 0.1567 | 0.1614 |
At thirty trials the randomised interval is 0.9% wider than Wilson’s and narrower than both intervals that guarantee their minimum. At a hundred trials the two widths differ by 0.0003, a fifth of a per cent. Wilson’s worst coverage at those sizes is 83.7% and 83.8%; the randomised interval’s is 95.00% at every proportion.
That is the whole of the discreteness problem priced in width. The conservatism of Clopper–Pearson — 10.4% wider than Wilson at thirty trials — is not the price of exactness. It is the price of exactness without a coin. With one, exact coverage costs about a hundredth of the width.
What the coin costs in everything else
If the randomised interval were free it would be in every textbook. It is in almost none, and the reasons are visible in one picture.
Two analysts with the same data report different intervals. Three successes in twenty trials give a lower limit anywhere from 0.0321 to 0.0573 and an upper limit anywhere from 0.3170 to 0.3789, depending on a draw that has nothing to do with the data. The interval is exact as a procedure and arbitrary as a report. A reader shown one realisation of it cannot tell which part of the width is information and which part is the coin.
An interval can be empty. When nothing is observed, the upper limit needs to reach 2.5% at some proportion, and since is at most one, any coin above 0.975 leaves no proportion inside. Counted over four thousand values of the coin, the interval for zero successes in twenty trials is empty for 2.5% of them, and counted over two hundred values for each count from one to nineteen it is never empty. So one study in forty that sees no events reports that the proportion is nowhere. The coverage is still exactly 95%, because an empty interval that misses is one of the 5%.
It violates a principle most readers hold without naming it: that two datasets with the same evidence should lead to the same conclusion. The randomised interval makes the conclusion depend on an event outside the data, which is exactly the property that makes its coverage exact and exactly the property a scientific report is not supposed to have.
The test that goes with it is the most powerful one
The interval’s partner is a test, and for the test there is a classical result that makes the coin harder to dismiss.
The Neyman–Pearson lemma says the most powerful test of one proportion against a larger one rejects for large counts — and, on a discrete sample space, rejects at the boundary count with a probability chosen to spend the level exactly. The randomised test is not a curiosity added to the theory; it is what the theory produces, and the non-randomised exact test is its rounding.
At twenty trials, testing a proportion of 0.1 against larger ones at 5%:
| size | power at 0.15 | power at 0.20 | power at 0.30 | |
|---|---|---|---|---|
| reject at five or more | 4.32% | 17.02% | 37.04% | 76.25% |
| also reject at four, with probability 0.076 | 5.00% | 18.40% | 38.69% | 77.24% |
The unrandomised test spends 4.32% of its 5%, and loses between one and one and a half points of power at each alternative for it. The coin is flipped only when exactly four successes are seen, and then it rejects about one time in thirteen. That is the entire difference between a test that is exactly at its level and one that is not, and it is the same difference, at the same count, as the one between Clopper–Pearson’s interval and the randomised interval’s.
This reframes the complaint. Nobody objects that the most powerful test is unfair to the data. They object that its result depends on a coin — which is a real objection, and a different one from the claim that the conservative test is the more careful choice. The conservative test is the less powerful one, and its extra caution is the unspent 0.68% of size, not a principled margin.
A coin before the data, and a coin after
The objection to the randomised interval is sharpest when it is put beside a coin the subject already accepts.
A randomised experiment assigns units to arms with a coin, and randomisation is not balance records how much ordinary variation that coin introduces into every estimate: two runs of the same trial on the same people give different answers because the coin fell differently. Nobody treats that as a defect of the analysis. The coin is accepted because it is flipped before the data exist, as part of producing them, and every estimate’s uncertainty already includes it.
The randomised interval’s coin is flipped after the data exist, as part of reading them. The data are fixed, the coin changes the conclusion, and the reader can see that it did. The mathematics of the two coins is the same — each turns a discrete or unbalanced structure into something with an exact reference distribution — and the objection is entirely about when in the process the randomness enters and whether it can be seen.
That distinction is defensible and worth holding. It is also worth seeing that it is a distinction about the legitimacy of a report, not about accuracy, and that the price of honouring it for a proportion is precisely the oscillation measured in the three essays before this one.
Reporting the coin instead of drawing it
The objection is to drawing the coin, not to the construction. The figure above already shows the way round it: instead of drawing and reporting one interval, report the whole band — for every proportion, the share of coin values for which it would be inside.
That object has a name, a fuzzy confidence interval, and it is a function from proportions to rather than a pair of endpoints. A proportion deep inside every realisation has membership one; a proportion outside every realisation has membership zero; a proportion in the shaded band’s edges has membership in between. It carries exactly the information of the randomised interval, it is reproducible, and it is never empty. It is also not a thing any reader knows how to use, and a report that says “the proportion is in this set to degree 0.63” has lost the plainness that made an interval worth printing.
The two objects are tied more closely than they look. Both of the randomised interval’s ends fall steadily as the coin’s value rises, so a proportion sitting exactly at the mid-p interval’s lower end is inside the randomised interval for every coin at or above one half and outside it for every coin below: its membership is exactly one half. The same holds at the upper end. The mid-p interval is the set of proportions whose membership in the fuzzy interval is at least one half — the fuzzy interval cut at its middle, which checking every count at ten and twenty trials against four hundred proportions confirms.
The practical compromise is to fix the coin at one half.
The coin fixed at one half
Setting gives the mid-p interval: each side counts half of the observed count’s probability. It is a function of the count again, so it is reproducible and never empty — and by the argument at the top of this essay its coverage has steps again.
| trials | mid-p worst coverage | mid-p average coverage | worst on one side |
|---|---|---|---|
| 10 | 92.66% | 96.83% | 4.99% |
| 30 | 92.45% | 95.82% | 4.92% |
| 100 | 92.07% | 95.29% | 4.88% |
The mid-p interval inherits the randomised interval’s centre and not its flatness. Its average is closer to 95% than any interval with a guarantee — within three tenths of a point at a hundred trials — and its worst is about three points below, far shallower than Wilson’s hole. On one side it can miss nearly 5% of the time, twice what a one-sided reader expects.
So the mid-p interval sits exactly where the comparison of guarantees would put it: no guarantee on the minimum, a good average, a moderate width. It is the coin’s expectation made into a function of the count, and it keeps the expectation’s virtues and gives up the coin’s one property that nothing else has.
What the construction settles
The three essays before this one treated discreteness as a fixed cost to be allocated — below the line here, above it there. This one shows that the cost is real only for a report that must be a function of the count.
That changes what the comparisons mean. An interval that covers 97.34% on average is not being careful; it is being forced, by the refusal to depend on anything but the count, to overshoot where the count’s steps fall. An interval that covers 83.8% at its worst is not merely approximate; it is placing its steps on the wrong side of the line where the counts are fewest. And the gap between the two is not a gap between cautious and careless statistics. It is the width of a single step of the binomial distribution, and it can be closed at almost no cost in width by an operation — adding a coin — that almost every reader would regard as illegitimate.
The same resolution applies wherever the discreteness came from. A p-value from a discrete test is conservative for the same reason and becomes exactly flat with the same coin; a permutation test on a small sample has the same steps and the same fix; and a coverage study of any count-based interval, including the ones this collection counts, is measuring where the steps sit rather than whether they exist.
What the sums establish, and what they do not
The randomised interval’s coverage is 95% exactly at every proportion checked, computed in closed form, and integrating over the coin numerically agrees to within two tenths of a point at a coarse rule and five hundredths at a fine one. The two routes are kept together because the closed form alone could be a formula written to give 0.95, and the numerical route counts intervals.
The randomised interval is narrower on average than both intervals that guarantee their minimum, and within 4% of Wilson’s width, at ten, thirty and a hundred trials.
For zero successes the interval is empty for 2.5% of coin values, and for every interior count it is never empty, counted on a grid of four thousand values for the first and two hundred for each of the others.
The randomised test spends its level exactly and the unrandomised one does not. Both are sums over the binomial at a null of 0.1, and the table’s two rows differ only by the probability of rejecting at four successes; the powers are sums at each alternative, not simulations, so the one-and-a-half-point differences are the differences and carry no sampling error.
A caution on the width table. The randomised interval’s average width is averaged over forty values of the coin as well as over a thousand proportions, so its third decimal is good and its fourth is approximate; the comparisons drawn from it — narrower than Blaker and Clopper–Pearson, within a few per cent of Wilson — are several times larger than that error at every sample size.
What does not survive is the mid-p interval read as guaranteeing its level. It is the coin’s average, not the coin, and its worst coverage is 92.45% at thirty trials.
Not claimed: that the randomised interval should be reported. The measurements say what it would buy; the objection to reporting a result that depends on a random draw is a principle about evidence, and a coverage calculation cannot overrule it. Nor that the fuzzy interval is usable as it stands — whether a reader shown a membership function draws better conclusions than one shown an interval is a question about readers, not about sums.
Still open: the coin that is already there
A randomised report is objectionable because the randomness is added. Some data already contain a continuous quantity that is independent of the count and could play the coin’s part — the order in which events arrived, a continuous measurement taken alongside each binary one, the time to each event in a trial that also reports a proportion.
Whether breaking ties in the count with such a quantity delivers the randomised interval’s exact coverage, and whether it is defensible as evidence where the drawn coin is not, depends on the quantity being independent of the proportion given the count. That independence is checkable for a given design, the resulting coverage is an exact sum over the outcomes of that design, and neither has been done.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- An interval that covers and says nothing — both name clopper–pearson, coverage, wilson interval
- The shortest interval is the one that misses — both name clopper–pearson, coverage, discreteness
- A coverage table with its own error — both name coverage, discreteness
- The same draws for both methods — both name coverage, discreteness
- The seed is part of the figure — both name coverage, reproducibility
- The shortest interval, and the one that does not move — both name coverage, discreteness
Named objects
A flat tag is an object no other essay names yet.
Clopper–PearsonCoverageDiscretenessFuzzy intervalMid-pRandomised intervalReproducibilityWilson interval