The coin that is already there
Worth reading first: More data is not monotonically better.
The coin that makes it exact found the one interval for a proportion that covers exactly 95% at every proportion: add a uniform random draw to the count, so that the sample space is continuous and nothing in it can oscillate. It also found why the interval is not used. Two analysts with the same data and different draws report different intervals, and an interval whose end depends on a coin that nobody observed is hard to defend as evidence.
The essay ended on a way round that objection. The randomised interval needs a uniform quantity that is independent of the proportion given the count; it does not need that quantity to be drawn by the analyst. Some data already contain one — the order in which events arrived, a continuous measurement taken alongside each binary one — and if such a quantity is genuinely independent of the proportion given the count, it can play the coin’s part with nothing added. Whether it delivers the randomised interval’s exact coverage, and whether it is defensible where a drawn coin is not, were the two questions. The order of arrival answers both, each with a condition.
An order that says nothing about the proportion
Suppose a proportion is estimated from trials recorded in sequence — patients treated one after another, items off a line, days on which an event did or did not happen — each succeeding independently with the same probability . The count is what every interval uses. The sequence says which of the trials succeeded, and given , every one of the possible sets of positions is equally likely, with probability each whatever is. The arrangement is ancillary: it carries no information about the proportion at all.
That makes it a coin. Number the arrangements of successes in a fixed order — the combinatorial number system’s, in which sets whose successes come later rank higher — and read the observed arrangement’s rank as
the midpoint of its share of the unit interval. Given the count, is uniform over equally spaced values and independent of , and the randomised interval can be computed with it in place of a drawn uniform: the interval covers exactly when lies between 2.5% and 97.5%. Every analyst who reads the same sequence in the same order gets the same , and so the same interval.
The only difference from a drawn coin is that this one has finitely many faces, and how many depends on the count.
Exact where the count has faces
The coverage of the arrangement coin is an exact sum: for each count, the share of its arrangements whose falls in the covering set, weighted by the count’s probability. The hero figure draws it at twenty trials.
Between proportions of 0.2 and 0.8 it is flat. Its largest departure from 95% there is 0.121 points, against 2.764 points for mid-p, whose coin is fixed at one half. At fifty trials the departure between 0.2 and 0.8 is below a thousandth of a point; at ten, 2.51 points against mid-p’s 4.10, because ten trials offer few arrangements at every count. Where the counts in play have hundreds or thousands of arrangements, the coin is effectively continuous and the interval does exactly what the drawn coin did.
The trouble is at the ends. Near a proportion of zero the likely counts are zero, one and two, and they have one, twenty and 190 arrangements at twenty trials. A count of zero has a single arrangement — no successes anywhere — so its coin always reads one half, and the interval at a count of zero is the mid-p interval, whose upper limit is , 0.139 at twenty trials. Below that proportion every trial with no successes covers, and the arrangement coin’s coverage sits near 97.5%; just above it, none of them does, and the coverage falls to its worst, 92.60%, at a proportion of 0.140. Mid-p’s own worst at twenty trials is 92.94%. Averaged over the proportion, the arrangement coin covers 95.63%, mid-p 96.11% and Clopper–Pearson 97.71%.
That is the arrangement coin’s limitation stated as a picture. It is exact exactly where exactness matters least — in the middle, where the ordinary intervals are already close to 95% — and it reverts to mid-p where a hole no sample size fills found every interval for a proportion weakest: the rare event, whose count is zero or one and whose arrangements are one or . A proportion of 0.05 on twenty trials puts 73.6% of its probability on counts of zero and one, and a coin with one face and a coin with twenty cannot reproduce a uniform.
A p-value that is flat
The same coin turns the exact test of a proportion into one whose p-value is uniform under the null, which is the property a p-value that is not flat set as the test of whether a p-value is a p-value at all. An exact binomial p-value is a step function of the count: with sixteen trials and a null of one half, the largest gap between its distribution and the uniform — the Kolmogorov distance — is 0.196, and mid-p’s, which fixes the coin at one half, is 0.098. Read with the arrangement coin, the p-value takes one value per arrangement, 65,536 of them, and its distance from the uniform is 0.00001: flat to five decimal places.
At a null of 0.3 on the same sixteen trials the distance is 0.0017, against mid-p’s 0.105; at a null of 0.1 on twenty trials, 0.061 against 0.143. The pattern is the one the coverage showed. Where the null puts its probability on counts with many arrangements, the order of arrival supplies all the continuity the count lacks; where it puts it on counts of zero and one, it supplies a little, and the test is still lumpy at the rare end.
Nothing about the interval’s width is given up for this. At twenty trials, averaged over the proportion, the arrangement coin’s interval is 0.3313 wide, a drawn coin’s 0.3300, mid-p’s 0.3348, Wilson’s 0.3254 and Clopper–Pearson’s 0.3661. The coin already in the data costs what a drawn coin costs, a little more than Wilson and a good deal less than the interval that buys its guarantee with width, and it does so without the oscillation in coverage that no function of the count alone can remove.
An interval that two analysts agree on
The arrangement coin answers the objection that stopped the drawn coin being used. Two analysts with the same sequence compute the same rank and the same interval, and a reader can check it. It also answers the second objection, the one the essay on the drawn coin found in the tails: a drawn coin gives an empty interval to one study in forty that sees no successes, because a coin above 97.5% at a count of zero leaves no proportion inside. The arrangement coin’s single face at a count of zero is one half, and its interval there is never empty.
What it asks in return is an agreement about the order. The rank is a function of the sequence read in a particular order, and it is ancillary only if that order was fixed before the data were seen and the trials in it are exchangeable. Each of those two conditions can fail, and they fail differently.
When the order carries information
The first condition is that the order carries no information about the proportion. That holds when every trial has the same success probability. When the probability drifts across the sequence — a treatment that works better as a team gains experience, a process that degrades over a shift — later trials succeed more often, the arrangements with late successes become more likely than those with early ones, and the rank is no longer uniform given the count.
The damage is modest. With the probability rising from 0.1 at the first of sixteen trials to 0.7 at the last — an average of 0.4, a drift of 0.3 either way — the arrangement coin’s interval covers the average proportion 95.92% of the time, a drawn coin’s 96.54% and mid-p’s 97.57%; all three cover more than with no drift, because a count built from unequal probabilities varies less than a binomial one. What changes for the arrangement coin is the balance: the average proportion lies below its interval in 1.25% of sequences and above it in 2.83%, against 2.50% and 2.50% with no drift. Late successes raise the rank, a higher coin value moves the interval down, and the interval errs low.
So a drifting process tilts the arrangement coin’s interval in the direction of the drift’s sign, and the tilt is a few tenths of a point at a drift large enough to be obvious in the data. A process that drifts that much is one where a single proportion is the wrong summary, and an analysis that noticed the drift would not report one; within the drift an analyst would plausibly miss, the arrangement coin behaves almost as the drawn one does.
When the order is chosen
The second condition is harder, because nothing in the data enforces it. A sequence can be read in more than one order — by time of treatment, by time of recording, by site, by patient identifier — and each gives a different rank and a different interval. If the order is chosen after the interval has been seen, the coin has been flipped until it landed the way the analyst wanted.
Read in the recorded order only, the test of on sixteen trials has size 4.999%. Allowed to choose between the recorded order and its reverse, 6.659%; among four orders, 7.275%; among eight, 7.678%; among sixteen, 7.681%, which is the ceiling — the size of rejecting whenever any value of the coin would, the most liberal test the randomisation allows. Testing the rise is smaller, from 4.969% to 5.179% with two orders and no further, because at that null the coin decides the outcome on less of the probability.
The ceiling is what makes the failure bounded. The coin can only move a test’s result at the counts where it matters — the counts on the boundary between rejecting and not — and choosing the coin can at worst turn every such count into a rejection. At sixteen trials that ceiling sits 2.7 points above the nominal level at a null of one half and 0.2 points above it at 0.3. A chosen order cannot do worse than that, and a drawn coin re-drawn until it lands well cannot either; the two are the same failure, and the remedy is the same — fix the order, like the seed of a simulation whose seed is part of the figure, before anything is computed.
The objection that remains
Reproducibility answers the practical complaint about a drawn coin. It does not answer the principled one, and it is worth stating plainly because the arrangement coin makes it sharper rather than weaker.
Two sequences with the same count have the same likelihood for the proportion: the probability of either, as a function of , is , and nothing in the order changes it. An interval built with the arrangement coin nevertheless gives them different intervals — one sequence with its successes early and another with them late, the same seven successes in twenty trials, can have intervals whose ends differ by a few hundredths. A reader who holds that evidence about is what the likelihood says, and nothing else, will find that indefensible, and the arrangement coin is exactly the kind of statistic the likelihood ignores. A drawn coin at least announces that it is noise; the order of arrival looks like data and is, for this purpose, noise.
The defence is the one the randomised interval always had, and it is a defence of the procedure rather than of any one interval. What 95 per cent means is a statement about the procedure’s long-run coverage, and the arrangement coin is the one way found here of making that statement exactly true at every proportion, with a coin anyone can recompute. Whether that is worth an interval that depends on something the likelihood ignores is a choice between two principles, and the measurements here only price it: exact coverage wherever the count has arrangements to spare, against intervals that differ by a few hundredths between sequences the likelihood cannot tell apart.
In practice the choice is easier than the principle, because the two intervals rarely disagree about anything a reader acts on. Where the count has many arrangements the coin moves the interval’s ends by less than the ends’ own rounding; where it has few, it is mid-p. The cases in which an interval with the coin and without it would lead to different decisions are the boundary counts of a test, and those are exactly the cases the chosen-order figure shows are the ones a coin can be abused on.
What the order of arrival can be used for
As the randomised interval’s coin, where the count has many arrangements. Between proportions of about 0.2 and 0.8 at twenty trials, or more widely at larger samples, it delivers the drawn coin’s exact 95% to within a tenth of a point, reproducibly.
Not as a repair for rare events. At a count of zero it has one face and the interval is mid-p’s; near the proportion where that interval’s upper limit sits, its coverage is as bad as mid-p’s, 92.6% at twenty trials. The problem what a guaranteed minimum costs priced for rare events is the one problem this coin cannot touch.
With the order stated before the data, and the trials’ exchangeability argued rather than assumed. A drift tilts the interval by a few tenths of a point; a chosen order can lift a 5% test to the most liberal test the randomisation allows. The first is a modelling question and the second is a reporting one, and both are answered by writing the order down in advance.
Every coverage here is an exact sum — over counts and, within each count, over the arrangements whose coin value falls inside the covering set — and the drift and chosen-order figures enumerate every sequence of sixteen trials. The arrangement ranks are checked to be distinct and to run from zero to one less than the number of arrangements at every count of ten trials, and the enumerated coverage with no drift is checked against the closed sum. Not measured: a continuous measurement alongside each binary outcome as the coin, which is not ancillary — the measurement’s distribution given a success depends on how far above the threshold the success fell, which depends on the proportion — and times to events in a trial that also reports a proportion, whose independence of the proportion given the count depends on the design.
Still open: a coin with more faces at the rare end
The arrangement coin fails at the ends because a count of zero or one has too few arrangements. A process observed continuously has more: the times at which a Poisson process’s events fell are, given their number, uniform on the observation window and independent of the rate, so a single event in a window carries a continuous coin — its position — and so do two. For a proportion built from a continuously observed process, the arrival times might supply the faces the order of trials cannot.
A count of zero is still a problem, since no events means no times; but a count of one would have a continuum of faces rather than . Whether an interval for a rate built this way delivers exact coverage down to an expected count of one, where every interval for a count has its deepest hole, and what the zero-count case forces it to give up, is computable exactly on the same terms as the sums here and has not been computed.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A coverage table with its own error — both name binomial proportion, coverage, discreteness
- An interval that covers and says nothing — both name binomial proportion, clopper–pearson, coverage
- The same draws for both methods — both name binomial proportion, coverage, discreteness
- The shortest interval is the one that misses — both name clopper–pearson, coverage, discreteness
- A simulation that stops when it looks settled — both name binomial proportion, coverage
- An upper limit is the finding — both name clopper–pearson, coverage
Named objects
A flat tag is an object no other essay names yet.
Binomial proportionClopper–PearsonCoverageDiscretenessExchangeabilityMid-pRandomised intervalReproducibility