A proportion's interval near the boundary, and the coin

The coin that is already there

The randomised interval covers exactly 95% at every proportion because it adds a drawn coin to the count, and a drawn coin is why nobody reports it. A sequence of trials already holds one: given the count, the order the successes arrived in is equally likely to be any of its arrangements whatever the proportion is. Used as the coin, it keeps twenty trials' coverage within 0.12 points of 95% from proportions of 0.2 to 0.8, where mid-p strays by 2.8, and it gives the same interval to every analyst. It fails where the count has few arrangements — a count of none has one — and it holds only while the order is read as recorded: an analyst choosing among sixteen orders lifts a 5% test on sixteen trials to 7.68%.

Worth reading first: More data is not monotonically better.

The coin that makes it exact found the one interval for a proportion that covers exactly 95% at every proportion: add a uniform random draw to the count, so that the sample space is continuous and nothing in it can oscillate. It also found why the interval is not used. Two analysts with the same data and different draws report different intervals, and an interval whose end depends on a coin that nobody observed is hard to defend as evidence.

The essay ended on a way round that objection. The randomised interval needs a uniform quantity that is independent of the proportion given the count; it does not need that quantity to be drawn by the analyst. Some data already contain one — the order in which events arrived, a continuous measurement taken alongside each binary one — and if such a quantity is genuinely independent of the proportion given the count, it can play the coin’s part with nothing added. Whether it delivers the randomised interval’s exact coverage, and whether it is defensible where a drawn coin is not, were the two questions. The order of arrival answers both, each with a condition.

Coverage across the proportion of the randomised interval with the arrangement of the successes as its coin, 20 trialsBetween proportions of 0.2 and 0.8 the arrangement coin's coverage stays within 0.121% of 95%, where mid-p's strays by 2.76%. Its worst coverage is 92.60%, at a proportion of 0.140, near the ends of the range where the count leaves few arrangements; mid-p's is 92.94%. Averaged over the proportion, 95.63% against mid-p's 96.11% and Clopper–Pearson's 97.71%.0.9000.9250.9500.975100.2000.4000.6000.8001the true proportionactual coverage of a nominal 95% intervalcoin: the order the successes arrived incoin fixed at one half (mid-p)exact sums over counts and arrangementsthe dashed line is a drawn coin's 95%
Fig. 1 Coverage against the true proportion, at twenty trials, of the randomised interval whose coin is the order the successes arrived in, and of the same interval with its coin fixed at one half, which is the mid-p interval. The dashed line is 95%, which a drawn coin reaches exactly. The slider sets the number of trials.

An order that says nothing about the proportion

Suppose a proportion is estimated from nn trials recorded in sequence — patients treated one after another, items off a line, days on which an event did or did not happen — each succeeding independently with the same probability pp. The count kk is what every interval uses. The sequence says which kk of the nn trials succeeded, and given kk, every one of the (nk)\binom{n}{k} possible sets of positions is equally likely, with probability pk(1−p)n−kp^k(1-p)^{n-k} each whatever pp is. The arrangement is ancillary: it carries no information about the proportion at all.

That makes it a coin. Number the arrangements of kk successes in a fixed order — the combinatorial number system’s, in which sets whose successes come later rank higher — and read the observed arrangement’s rank rr as

u=r+12(nk),u = \frac{r + \tfrac12}{\binom{n}{k}},

the midpoint of its share of the unit interval. Given the count, uu is uniform over (nk)\binom{n}{k} equally spaced values and independent of pp, and the randomised interval can be computed with it in place of a drawn uniform: the interval covers pp exactly when P(X>k)+u P(X=k)P(X > k) + u\,P(X = k) lies between 2.5% and 97.5%. Every analyst who reads the same sequence in the same order gets the same uu, and so the same interval.

The only difference from a drawn coin is that this one has finitely many faces, and how many depends on the count.

Exact where the count has faces

The coverage of the arrangement coin is an exact sum: for each count, the share of its arrangements whose uu falls in the covering set, weighted by the count’s probability. The hero figure draws it at twenty trials.

Between proportions of 0.2 and 0.8 it is flat. Its largest departure from 95% there is 0.121 points, against 2.764 points for mid-p, whose coin is fixed at one half. At fifty trials the departure between 0.2 and 0.8 is below a thousandth of a point; at ten, 2.51 points against mid-p’s 4.10, because ten trials offer few arrangements at every count. Where the counts in play have hundreds or thousands of arrangements, the coin is effectively continuous and the interval does exactly what the drawn coin did.

The trouble is at the ends. Near a proportion of zero the likely counts are zero, one and two, and they have one, twenty and 190 arrangements at twenty trials. A count of zero has a single arrangement — no successes anywhere — so its coin always reads one half, and the interval at a count of zero is the mid-p interval, whose upper limit is 1−0.051/n1 - 0.05^{1/n}, 0.139 at twenty trials. Below that proportion every trial with no successes covers, and the arrangement coin’s coverage sits near 97.5%; just above it, none of them does, and the coverage falls to its worst, 92.60%, at a proportion of 0.140. Mid-p’s own worst at twenty trials is 92.94%. Averaged over the proportion, the arrangement coin covers 95.63%, mid-p 96.11% and Clopper–Pearson 97.71%.

How many faces the arrangement coin has at each count of 20 trials, with the chance of each count at a proportion of 0.05. C(20, k) arrangements at each count: one at no successes, 20 at one, 190 at two, 184,756 at 10. At a proportion of 0.05 the counts of none and one carry 73.6% of the probability.
Fig. 2 How many values the arrangement coin can take at each count of twenty trials — the number of arrangements, on a logarithmic scale — with the chance of each count at a proportion of 0.05 as bars. Where the counts are likely, the coin has few faces.

That is the arrangement coin’s limitation stated as a picture. It is exact exactly where exactness matters least — in the middle, where the ordinary intervals are already close to 95% — and it reverts to mid-p where a hole no sample size fills found every interval for a proportion weakest: the rare event, whose count is zero or one and whose arrangements are one or nn. A proportion of 0.05 on twenty trials puts 73.6% of its probability on counts of zero and one, and a coin with one face and a coin with twenty cannot reproduce a uniform.

A p-value that is flat

The same coin turns the exact test of a proportion into one whose p-value is uniform under the null, which is the property a p-value that is not flat set as the test of whether a p-value is a p-value at all. An exact binomial p-value is a step function of the count: with sixteen trials and a null of one half, the largest gap between its distribution and the uniform — the Kolmogorov distance — is 0.196, and mid-p’s, which fixes the coin at one half, is 0.098. Read with the arrangement coin, the p-value takes one value per arrangement, 65,536 of them, and its distance from the uniform is 0.00001: flat to five decimal places.

At a null of 0.3 on the same sixteen trials the distance is 0.0017, against mid-p’s 0.105; at a null of 0.1 on twenty trials, 0.061 against 0.143. The pattern is the one the coverage showed. Where the null puts its probability on counts with many arrangements, the order of arrival supplies all the continuity the count lacks; where it puts it on counts of zero and one, it supplies a little, and the test is still lumpy at the rare end.

Nothing about the interval’s width is given up for this. At twenty trials, averaged over the proportion, the arrangement coin’s interval is 0.3313 wide, a drawn coin’s 0.3300, mid-p’s 0.3348, Wilson’s 0.3254 and Clopper–Pearson’s 0.3661. The coin already in the data costs what a drawn coin costs, a little more than Wilson and a good deal less than the interval that buys its guarantee with width, and it does so without the oscillation in coverage that no function of the count alone can remove.

An interval that two analysts agree on

The arrangement coin answers the objection that stopped the drawn coin being used. Two analysts with the same sequence compute the same rank and the same interval, and a reader can check it. It also answers the second objection, the one the essay on the drawn coin found in the tails: a drawn coin gives an empty interval to one study in forty that sees no successes, because a coin above 97.5% at a count of zero leaves no proportion inside. The arrangement coin’s single face at a count of zero is one half, and its interval there is never empty.

What it asks in return is an agreement about the order. The rank is a function of the sequence read in a particular order, and it is ancillary only if that order was fixed before the data were seen and the trials in it are exchangeable. Each of those two conditions can fail, and they fail differently.

When the order carries information

The first condition is that the order carries no information about the proportion. That holds when every trial has the same success probability. When the probability drifts across the sequence — a treatment that works better as a team gains experience, a process that degrades over a shift — later trials succeed more often, the arrangements with late successes become more likely than those with early ones, and the rank is no longer uniform given the count.

Which side the arrangement coin's interval misses the average proportion on, when the success probability drifts across sixteen trials. The average proportion is 0.4 throughout. With no drift the arrangement coin misses 2.50% below and 2.50% above. At a drift of 0.3, 1.25% and 2.83%; its total coverage is 95.92%, a drawn coin's 96.54% and mid-p's 97.57%.
Fig. 3 With the success probability rising linearly across sixteen trials around an average of 0.4, the share of sequences whose arrangement-coin interval lies wholly above the average proportion and wholly below it, against the size of the drift. Every one of the 65,536 sequences is enumerated at each drift.

The damage is modest. With the probability rising from 0.1 at the first of sixteen trials to 0.7 at the last — an average of 0.4, a drift of 0.3 either way — the arrangement coin’s interval covers the average proportion 95.92% of the time, a drawn coin’s 96.54% and mid-p’s 97.57%; all three cover more than with no drift, because a count built from unequal probabilities varies less than a binomial one. What changes for the arrangement coin is the balance: the average proportion lies below its interval in 1.25% of sequences and above it in 2.83%, against 2.50% and 2.50% with no drift. Late successes raise the rank, a higher coin value moves the interval down, and the interval errs low.

So a drifting process tilts the arrangement coin’s interval in the direction of the drift’s sign, and the tilt is a few tenths of a point at a drift large enough to be obvious in the data. A process that drifts that much is one where a single proportion is the wrong summary, and an analysis that noticed the drift would not report one; within the drift an analyst would plausibly miss, the arrangement coin behaves almost as the drawn one does.

When the order is chosen

The second condition is harder, because nothing in the data enforces it. A sequence can be read in more than one order — by time of treatment, by time of recording, by site, by patient identifier — and each gives a different rank and a different interval. If the order is chosen after the interval has been seen, the coin has been flipped until it landed the way the analyst wanted.

The size of an exact test whose coin is the arrangement, when the analyst chooses which order to read, sixteen trials. Reading the recorded order only, the test of p = 0.5 has size 4.999%; choosing the most favourable of 2, 4, 8, 16 orderings, 6.659%, 7.275%, 7.678%, 7.681% — up to the 7.681% of rejecting whenever any coin value would. For p = 0.3: 4.969%, 5.179%, 5.179%, 5.179%, 5.179%, against a ceiling of 5.179%.
Fig. 4 The size of the two-sided exact test of a proportion whose coin is the arrangement, on sixteen trials, when the analyst reports whichever of several orderings gives the smallest p-value — the recorded order, its reverse and others fixed in advance — for a null of 0.5 and of 0.3. The dashed lines are the ceiling: rejecting whenever any coin value would. Every sequence is enumerated.

Read in the recorded order only, the test of p=0.5p = 0.5 on sixteen trials has size 4.999%. Allowed to choose between the recorded order and its reverse, 6.659%; among four orders, 7.275%; among eight, 7.678%; among sixteen, 7.681%, which is the ceiling — the size of rejecting whenever any value of the coin would, the most liberal test the randomisation allows. Testing p=0.3p = 0.3 the rise is smaller, from 4.969% to 5.179% with two orders and no further, because at that null the coin decides the outcome on less of the probability.

The ceiling is what makes the failure bounded. The coin can only move a test’s result at the counts where it matters — the counts on the boundary between rejecting and not — and choosing the coin can at worst turn every such count into a rejection. At sixteen trials that ceiling sits 2.7 points above the nominal level at a null of one half and 0.2 points above it at 0.3. A chosen order cannot do worse than that, and a drawn coin re-drawn until it lands well cannot either; the two are the same failure, and the remedy is the same — fix the order, like the seed of a simulation whose seed is part of the figure, before anything is computed.

The objection that remains

Reproducibility answers the practical complaint about a drawn coin. It does not answer the principled one, and it is worth stating plainly because the arrangement coin makes it sharper rather than weaker.

Two sequences with the same count have the same likelihood for the proportion: the probability of either, as a function of pp, is pk(1−p)n−kp^k(1-p)^{n-k}, and nothing in the order changes it. An interval built with the arrangement coin nevertheless gives them different intervals — one sequence with its successes early and another with them late, the same seven successes in twenty trials, can have intervals whose ends differ by a few hundredths. A reader who holds that evidence about pp is what the likelihood says, and nothing else, will find that indefensible, and the arrangement coin is exactly the kind of statistic the likelihood ignores. A drawn coin at least announces that it is noise; the order of arrival looks like data and is, for this purpose, noise.

The defence is the one the randomised interval always had, and it is a defence of the procedure rather than of any one interval. What 95 per cent means is a statement about the procedure’s long-run coverage, and the arrangement coin is the one way found here of making that statement exactly true at every proportion, with a coin anyone can recompute. Whether that is worth an interval that depends on something the likelihood ignores is a choice between two principles, and the measurements here only price it: exact coverage wherever the count has arrangements to spare, against intervals that differ by a few hundredths between sequences the likelihood cannot tell apart.

In practice the choice is easier than the principle, because the two intervals rarely disagree about anything a reader acts on. Where the count has many arrangements the coin moves the interval’s ends by less than the ends’ own rounding; where it has few, it is mid-p. The cases in which an interval with the coin and without it would lead to different decisions are the boundary counts of a test, and those are exactly the cases the chosen-order figure shows are the ones a coin can be abused on.

What the order of arrival can be used for

As the randomised interval’s coin, where the count has many arrangements. Between proportions of about 0.2 and 0.8 at twenty trials, or more widely at larger samples, it delivers the drawn coin’s exact 95% to within a tenth of a point, reproducibly.

Not as a repair for rare events. At a count of zero it has one face and the interval is mid-p’s; near the proportion where that interval’s upper limit sits, its coverage is as bad as mid-p’s, 92.6% at twenty trials. The problem what a guaranteed minimum costs priced for rare events is the one problem this coin cannot touch.

With the order stated before the data, and the trials’ exchangeability argued rather than assumed. A drift tilts the interval by a few tenths of a point; a chosen order can lift a 5% test to the most liberal test the randomisation allows. The first is a modelling question and the second is a reporting one, and both are answered by writing the order down in advance.

Every coverage here is an exact sum — over counts and, within each count, over the arrangements whose coin value falls inside the covering set — and the drift and chosen-order figures enumerate every sequence of sixteen trials. The arrangement ranks are checked to be distinct and to run from zero to one less than the number of arrangements at every count of ten trials, and the enumerated coverage with no drift is checked against the closed sum. Not measured: a continuous measurement alongside each binary outcome as the coin, which is not ancillary — the measurement’s distribution given a success depends on how far above the threshold the success fell, which depends on the proportion — and times to events in a trial that also reports a proportion, whose independence of the proportion given the count depends on the design.

Still open: a coin with more faces at the rare end

The arrangement coin fails at the ends because a count of zero or one has too few arrangements. A process observed continuously has more: the times at which a Poisson process’s events fell are, given their number, uniform on the observation window and independent of the rate, so a single event in a window carries a continuous coin — its position — and so do two. For a proportion built from a continuously observed process, the arrival times might supply the faces the order of trials cannot.

A count of zero is still a problem, since no events means no times; but a count of one would have a continuum of faces rather than nn. Whether an interval for a rate built this way delivers exact coverage down to an expected count of one, where every interval for a count has its deepest hole, and what the zero-count case forces it to give up, is computable exactly on the same terms as the sums here and has not been computed.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Binomial proportionClopper–PearsonCoverageDiscretenessExchangeabilityMid-pRandomised intervalReproducibility