A proportion's interval near the boundary, and the coin

Two windows and a ratio

Two counts from two windows give an interval for the ratio of their rates through the conditional binomial: given the total, the first window's count is binomial with a proportion that depends only on the ratio. The arrival times in both windows supply that binomial with a continuous coin at every total from one up, and the interval then covers exactly 95% plus a twentieth of the chance of no events at all. A total of zero can only report the whole of (0, ∞). And the exact conditional test's familiar conservatism is not the zero total's: the zero is the larger cause only while the expected total is below ln 2, and at an expected total of ten the exact interval still covers 97.8% for reasons that have nothing to do with it.

Worth reading first: More data is not monotonically better.

The coin a clock supplies built an interval for one Poisson mean that covers exactly 95% at every expected count beyond ln⁡40\ln 40. It used the times at which the events arrived: given the count, those times are uniform on the window whatever the rate, so they are a coin nobody drew and every analyst can read. The only count left without a coin was zero, and the whole of the interval’s departure from 95% came from what a count of zero was made to report.

The comparisons people make with counts are almost never about one rate. They are about two — adverse events in two arms of a trial, failures in two fleets, cases in two seasons — and the quantity of interest is the ratio of the rates. The essay ended on that case, with two questions. Do the arrival times in two windows supply a coin for the ratio’s interval, so that it too is exact at every total except zero? And is the well-known conservatism of the usual exact test for a ratio at small counts entirely the zero total’s fault, as the one-window hole had been entirely the zero count’s?

The answer to the first is yes, with a twist about what zero must report. The answer to the second is no, and the reason is worth more than the yes.

A ratio is a proportion given the total

Two windows of lengths t1t_1 and t2t_2, with Poisson counts k1k_1 and k2k_2 at rates λ1\lambda_1 and λ2\lambda_2. The standard move is to condition on the total n=k1+k2n = k_1 + k_2. Given nn, the first window’s count is binomial on nn trials with proportion

p=θt1θt1+t2,θ=λ1/λ2,p = \frac{\theta t_1}{\theta t_1 + t_2}, \qquad \theta = \lambda_1/\lambda_2,

and nothing else about the rates enters. So an interval for pp given nn — any of the intervals this field has measured for a proportion — is an interval for the ratio, through a monotone map. The usual exact interval for a ratio of Poisson rates is the Clopper–Pearson interval for this conditional binomial, and the usual exact test is the binomial test it inverts.

Its coverage is a sum over totals. With the total’s mean M=μ1+μ2M = \mu_1 + \mu_2, the unconditional coverage is e−Me^{-M} times whatever a total of zero reports, plus the conditional coverage at each positive total weighted by its Poisson probability. Every number below is that sum, computed exactly.

A coin from both windows

Given both counts, each event’s position within its own window is uniform, whatever either rate is, and independent of every other event’s. So the arrival times carry a continuous coin for the conditional binomial at every total from one up. The simplest reads all nn events in window-relative time — each window rescaled to unit length — and takes the latest: u=t(n) nu = t_{(n)}^{\,n} is uniform given the counts. It exists whichever window the events fell in, so a total of one, with its single event in either window, has a coin with a continuum of faces.

With that coin the interval for pp is the randomised interval the drawn coin made exact, and it covers exactly 95% at every proportion and every total from one up. A total of zero has no times and no counts to divide, and the only interval for a ratio that can be defended from no events is the whole of (0,∞)(0, \infty): any narrower report would be a claim about θ\theta made from nothing, and it would fail at every θ\theta it left out, with probability e−Me^{-M}, which is as large as one likes when MM is small. So the zero total covers with certainty, and the unconditional coverage is

0.95+0.05 e−M.0.95 + 0.05\,e^{-M}.

Coverage of an interval for a ratio of two rates against the expected total. The unconditional coverage of three intervals for the ratio of two Poisson rates, each built conditionally on the total and each reporting the whole of (0, ∞) when the total is zero, at the conditional proportion 0.5, against the expected total count. With the arrival times in both windows as the coin the interval covers exactly 95% plus a twentieth of e raised to minus the expected total: 96.84% at an expected total of one, 95.25% at three, 95.00% from about ten. The exact conditional interval covers 100.00%, 99.79%, 97.82% and 96.55% at forty; mid-p 99.11% at three and 95.89% at ten.
Fig. 1 Unconditional coverage of three intervals for a ratio of two rates against the expected total count, at a conditional proportion of one half.

The arrival-coin interval covers 96.84% at an expected total of one, 95.25% at three and 95.00% from about ten, following that formula to the last digit. The exact conditional interval covers 100.00% at one, 99.79% at three, 97.82% at ten and still 96.55% at forty. Mid-p, the conditional interval with the coin fixed at one half, covers 99.11% at three and 95.89% at ten.

The sums can be checked by a route that shares nothing with them but the interval itself. Two routes to every number is the habit here, so the two windows were simulated directly: Poisson counts in each, every event’s time drawn, the coin read from the latest time, the interval computed and checked against the truth. Over a hundred thousand simulated pairs of windows the arrival-coin interval covers 96.84% at an expected total of one, with a standard error of 0.06 points, and 95.36% at three, with 0.07; the sums say 96.84% and 95.25%. Both are within the simulation’s own error of the formula — the second at about one and a half of its standard errors.

There is one difference from the single window, and it is the twist. For one mean the zero count had to be given a stated interval, and the choice of that interval moved the worst coverage by two and a half points. For a ratio there is no choice: the zero total’s interval is everything, so the departure from 95% is always on the safe side and always exactly 0.05 e−M0.05\,e^{-M}. The ratio problem is better behaved than the rate problem at zero, because a ratio from no events has no candidate answer to be wrong about.

Whose conservatism it is

Now the second question. The exact conditional test for a ratio is known to be conservative at small counts — its actual size well below 5% — and the zero total contributes to that, since a total of zero never rejects. The question is how much.

Whose conservatism the exact conditional interval's is. The exact conditional interval's coverage above 95% at the conditional proportion 0.5, split into the part a total of zero contributes — a twentieth of e to the minus the total's mean, which the arrival coin also carries — and the part every positive total adds through the discreteness of the binomial. The zero total's part is the larger only while the expected total is below 0.69. At an expected total of 2 it is 13.7% of the excess; at 5, 0.8%; at 10 the excess is 2.82% and the zero total's share of it rounds to nothing.
Fig. 2 The exact conditional interval’s coverage above 95%, split into the part a total of zero contributes and the part every positive total adds through the discreteness of the binomial.

The excess splits exactly. The zero total contributes 0.05 e−M0.05\,e^{-M} — the same amount the arrival coin carries. Every positive total contributes the amount by which its conditional Clopper–Pearson coverage exceeds 95%, weighted by its probability. The zero total’s part is the larger of the two only while the expected total is below 0.693. At an expected total of two it is 13.7% of the excess, at five 0.8%, and at ten the excess is 2.82% with the zero total’s part too small to show.

The crossing is not an accident. At a proportion of one half, the Clopper–Pearson interval for a binomial on up to five trials contains one half at every count, so for small totals it covers with certainty and its excess from each is the full 5%. While totals above five are rare, its discreteness part is therefore almost exactly 0.05(1−e−M)0.05(1 - e^{-M}) against the zero total’s 0.05 e−M0.05\,e^{-M}, and the two are equal at M=ln⁡2=0.6931M = \ln 2 = 0.6931 — the expected total at which no events at all is exactly as likely as some. The zero total is the larger cause of conservatism precisely while it is the more likely outcome, and not beyond.

So the conditional test’s conservatism is not the zero total’s. It is the same discreteness a guaranteed minimum always charges a count, now charged at every total, and it persists long after zero totals have become rare. That is the opposite of the single-window result, where the hole at low counts was entirely the zero count’s and a continuous coin removed everything else. Here the continuous coin removes everything else too — but everything else turns out to be nearly all of it.

When the windows are not the same length

Everything above is at a conditional proportion of one half — equal rates in windows of equal length. Real comparisons are rarely that symmetric. A surveillance system compares a short exposure period with a long baseline; a trial’s control arm is followed for longer than its treatment arm; and the conditional proportion is then far from a half even when the rates are equal. With the second window four times as long as the first and the rates equal, p=1/5p = 1/5.

The arrival coin does not notice. Its coverage is 0.95+0.05 e−M0.95 + 0.05\,e^{-M} at every proportion, because the coin is uniform given the counts and the randomised binomial interval is exact at every pp; the windows’ lengths enter only through pp. The exact conditional interval does notice. At p=1/5p = 1/5 it covers 98.78% at an expected total of ten and 97.59% at twenty, against 97.82% and 97.07% at one half, so an unbalanced design makes it more conservative, not less. Mid-p covers 97.47% at ten and 95.60% at twenty, against 95.89% and 95.43%.

The zero total’s share barely moves: the crossing below which it is the larger cause is at an expected total of 0.698 rather than 0.693. So the lesson of the balanced case carries over to the unbalanced one with more force. The zero total is a small-count curiosity; the discreteness of the conditional binomial is the conservatism, and an unbalanced design, which puts the conditional proportion near the boundary where every binomial interval oscillates most, makes it worse.

What exactness costs when one window is empty

A randomised interval buys its exact coverage in an unfamiliar currency, and the ratio makes the currency visible. Suppose the first window recorded no events and the second recorded five. Every conventional interval for the ratio then has a lower limit of zero: no events in the first window cannot rule out a first rate of nothing.

The randomised interval’s lower limit at a first-window count of zero is zero only when the coin exceeds 2.5%. When the coin falls below 2.5% — when the latest of the five arrivals, read in window-relative time, came before about 0.48 of the way through its window, since 0.0251/5=0.4780.025^{1/5} = 0.478 — the lower limit is positive, and the interval excludes a ratio of zero on a sample that contains no evidence against one. That happens on exactly 2.5% of such samples, and it is exactly the 2.5% the interval needs to miss on the lower side when the first rate really is close to zero. The upper limit moves with the coin too: at the coin’s middle value it is a ratio of 0.82 on five events, the mid-p value.

That is what a coin does: it spreads the binomial’s lumps of probability over a continuum, and occasionally the spreading puts an endpoint where no fixed rule would. The single-window essay faced the same thing at a count of one and argued that an interval read from the data, recomputable by anyone, is defensible where a drawn one is not. The ratio adds one more reason. A total of zero is the only outcome on which the interval says nothing, and it says nothing exactly; on every other outcome its errors are the 5% it promised, distributed evenly over the ratios a reader might be asking about rather than concentrated where the counts are smallest.

Across the ratio, at one total

The excess above was read at one proportion. A discrete interval’s coverage oscillates with the proportion, and the arrival coin’s should not.

Coverage across the ratio at one total. Coverage of each interval against the conditional proportion the ratio implies, at a total of mean 5. The arrival coin covers 95.034% at every proportion — 95% plus a twentieth of the chance of a total of zero. Mid-p ranges from 96.50% at 0.54 upwards and the exact conditional interval from 98.61% upwards, both oscillating with the proportion as every interval on a discrete count does.
Fig. 3 Coverage of the three intervals against the conditional proportion the ratio implies, at an expected total of five.

At an expected total of five the arrival coin covers 95.034% at every proportion, flat. Mid-p’s coverage oscillates and its lowest point on the grid is 96.50%; the exact conditional interval’s lowest is 98.61%. At this total, then, mid-p — the usual pragmatic repair — never falls below 95% at all, and still spends most of its range more than a point above it. Its familiar under-coverage needs larger totals: the single-window essay found mid-p reaching 91.66% for one Poisson mean, and for the ratio the averaging over totals smooths most of that away at small expected counts.

What the conservatism costs in power

An interval that covers 98% is a test whose size is 2%, and a test that rejects less often than it may also rejects less often when it should. The power of the test each interval inverts makes that concrete.

The test each interval for a ratio inverts. The probability that each interval leaves out a ratio of one, against the true ratio of the first window's rate to the second's, for two windows of equal length with the second's mean fixed at 5. At a ratio of one — the size of the test — the exact conditional interval rejects 2.18%, mid-p 4.11% and the arrival coin 5.00%, which is 5% less the zero total's share. At a ratio of two they reject 17.8%, 23.9% and 24.5%; at three 54.2%, 61.6% and 62.2%. The exact conditional test's conservatism is paid for in power at every ratio.
Fig. 4 The probability that each interval excludes a ratio of one, against the true ratio, with two windows of equal length and the second window’s expected count fixed at five.

The setting is a two-arm comparison with five events expected in the reference arm. At a true ratio of one the exact conditional test rejects 2.18% of the time, mid-p 4.11%, and the arrival coin 5.00% — five per cent less the zero total’s share, which at an expected total of ten is about two millionths. At a true ratio of two the three reject 17.8%, 23.9% and 24.5%; at three, 54.2%, 61.6% and 62.2%.

So the exact conditional test gives up about seven points of power at a doubling of the rate and eight at a tripling, in exchange for a size of 2.18% where 5% was allowed. Mid-p recovers most of that and the arrival coin the rest. The difference between mid-p and the arrival coin is small at these counts — under a point of power — and the difference between either and the exact test is not.

For a safety comparison that is the difference that matters. A monitoring board watching for a doubling of an adverse event’s rate, with five events expected in the reference arm, sees it flagged by the exact conditional test in fewer than one look in five and by the arrival-coin test in about one look in four — and pays for the difference only with the false alarms it was entitled to: the arrival test’s size is the 5% the board asked for, and the exact test’s 2.18% is a level nobody chose.

What it costs in width

The same comparison in the interval’s width, read on the scale of the conditional proportion where it is bounded.

What each interval for a ratio costs in width. The expected width of each interval for the conditional proportion 0.5, given at least one event, against the expected total. At a total of mean 2 the exact conditional interval is 0.895 wide, mid-p 0.853 and the arrival coin 0.806; at 10 they are 0.612, 0.563 and 0.553. The arrival coin is the narrowest at every total, and its advantage over mid-p shrinks as the total grows while its advantage over the exact interval does not.
Fig. 5 The expected width of each interval for the conditional proportion, given at least one event, against the expected total.

At an expected total of two the exact conditional interval is 0.895 wide on the proportion scale, mid-p 0.853 and the arrival coin 0.806; at ten, 0.612, 0.563 and 0.553. The arrival coin is the narrowest at every total. Its advantage over the exact interval is about a tenth of the width at small and moderate totals; its advantage over mid-p shrinks with the total, from about five per cent at two to under two per cent at ten, because mid-p’s fixed coin at one half becomes a better stand-in for a uniform one as the binomial’s atoms shrink.

The objection, and the design that answers it

The objection to a randomised interval is that the randomness was added, and the coin that is already there answered it for a proportion by finding a coin in the data. For a ratio the answer is stronger. The arrival times are recorded in every surveillance system, every trial’s adverse-event log and every reliability study; nobody has to draw anything, and two analysts reading the same log compute the same interval.

The condition is the one the single window had. The times have to be uniform within each window given the counts, which holds when each window’s rate is constant across it. A rate that drifts within a window — more failures late in a service life, more cases late in a season — makes the latest arrival come later than a uniform would, and the coin leans. The single-window essay measured that lean for a linear drift and found a folded coin, read from distances to the window’s middle, that is immune to it; the same folded coin applies here window by window, and it would need checking against drifts that differ between the two windows, which is the realistic case for two arms of a trial.

Still open: rates that drift differently in the two windows

The next measurement is that check. In a two-arm comparison the arms’ rates can drift in different directions — a treatment’s adverse events concentrated early, the control’s spread evenly — and a coin read from both windows at once then mixes two leans. The folded coin is uniform given the counts under any drift symmetric about a linear trend within each window, whatever the two trends are, so it should survive; the last-arrival coin should not. Whether the folded coin’s coverage stays at 0.95+0.05 e−M0.95 + 0.05\,e^{-M} when the windows drift in opposite directions, and what the last-arrival coin’s lean costs in coverage at the totals where the exact conditional test’s conservatism is largest, is a finite sum of the same kind as everything here and has not been computed.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Clopper–PearsonConditional testConservative intervalCoverageDiscretenessExchangeabilityInterval widthMid-pRandomised intervalStatistical power