Two windows and a ratio
Worth reading first: More data is not monotonically better.
The coin a clock supplies built an interval for one Poisson mean that covers exactly 95% at every expected count beyond . It used the times at which the events arrived: given the count, those times are uniform on the window whatever the rate, so they are a coin nobody drew and every analyst can read. The only count left without a coin was zero, and the whole of the interval’s departure from 95% came from what a count of zero was made to report.
The comparisons people make with counts are almost never about one rate. They are about two — adverse events in two arms of a trial, failures in two fleets, cases in two seasons — and the quantity of interest is the ratio of the rates. The essay ended on that case, with two questions. Do the arrival times in two windows supply a coin for the ratio’s interval, so that it too is exact at every total except zero? And is the well-known conservatism of the usual exact test for a ratio at small counts entirely the zero total’s fault, as the one-window hole had been entirely the zero count’s?
The answer to the first is yes, with a twist about what zero must report. The answer to the second is no, and the reason is worth more than the yes.
A ratio is a proportion given the total
Two windows of lengths and , with Poisson counts and at rates and . The standard move is to condition on the total . Given , the first window’s count is binomial on trials with proportion
and nothing else about the rates enters. So an interval for given — any of the intervals this field has measured for a proportion — is an interval for the ratio, through a monotone map. The usual exact interval for a ratio of Poisson rates is the Clopper–Pearson interval for this conditional binomial, and the usual exact test is the binomial test it inverts.
Its coverage is a sum over totals. With the total’s mean , the unconditional coverage is times whatever a total of zero reports, plus the conditional coverage at each positive total weighted by its Poisson probability. Every number below is that sum, computed exactly.
A coin from both windows
Given both counts, each event’s position within its own window is uniform, whatever either rate is, and independent of every other event’s. So the arrival times carry a continuous coin for the conditional binomial at every total from one up. The simplest reads all events in window-relative time — each window rescaled to unit length — and takes the latest: is uniform given the counts. It exists whichever window the events fell in, so a total of one, with its single event in either window, has a coin with a continuum of faces.
With that coin the interval for is the randomised interval the drawn coin made exact, and it covers exactly 95% at every proportion and every total from one up. A total of zero has no times and no counts to divide, and the only interval for a ratio that can be defended from no events is the whole of : any narrower report would be a claim about made from nothing, and it would fail at every it left out, with probability , which is as large as one likes when is small. So the zero total covers with certainty, and the unconditional coverage is
The arrival-coin interval covers 96.84% at an expected total of one, 95.25% at three and 95.00% from about ten, following that formula to the last digit. The exact conditional interval covers 100.00% at one, 99.79% at three, 97.82% at ten and still 96.55% at forty. Mid-p, the conditional interval with the coin fixed at one half, covers 99.11% at three and 95.89% at ten.
The sums can be checked by a route that shares nothing with them but the interval itself. Two routes to every number is the habit here, so the two windows were simulated directly: Poisson counts in each, every event’s time drawn, the coin read from the latest time, the interval computed and checked against the truth. Over a hundred thousand simulated pairs of windows the arrival-coin interval covers 96.84% at an expected total of one, with a standard error of 0.06 points, and 95.36% at three, with 0.07; the sums say 96.84% and 95.25%. Both are within the simulation’s own error of the formula — the second at about one and a half of its standard errors.
There is one difference from the single window, and it is the twist. For one mean the zero count had to be given a stated interval, and the choice of that interval moved the worst coverage by two and a half points. For a ratio there is no choice: the zero total’s interval is everything, so the departure from 95% is always on the safe side and always exactly . The ratio problem is better behaved than the rate problem at zero, because a ratio from no events has no candidate answer to be wrong about.
Whose conservatism it is
Now the second question. The exact conditional test for a ratio is known to be conservative at small counts — its actual size well below 5% — and the zero total contributes to that, since a total of zero never rejects. The question is how much.
The excess splits exactly. The zero total contributes — the same amount the arrival coin carries. Every positive total contributes the amount by which its conditional Clopper–Pearson coverage exceeds 95%, weighted by its probability. The zero total’s part is the larger of the two only while the expected total is below 0.693. At an expected total of two it is 13.7% of the excess, at five 0.8%, and at ten the excess is 2.82% with the zero total’s part too small to show.
The crossing is not an accident. At a proportion of one half, the Clopper–Pearson interval for a binomial on up to five trials contains one half at every count, so for small totals it covers with certainty and its excess from each is the full 5%. While totals above five are rare, its discreteness part is therefore almost exactly against the zero total’s , and the two are equal at — the expected total at which no events at all is exactly as likely as some. The zero total is the larger cause of conservatism precisely while it is the more likely outcome, and not beyond.
So the conditional test’s conservatism is not the zero total’s. It is the same discreteness a guaranteed minimum always charges a count, now charged at every total, and it persists long after zero totals have become rare. That is the opposite of the single-window result, where the hole at low counts was entirely the zero count’s and a continuous coin removed everything else. Here the continuous coin removes everything else too — but everything else turns out to be nearly all of it.
When the windows are not the same length
Everything above is at a conditional proportion of one half — equal rates in windows of equal length. Real comparisons are rarely that symmetric. A surveillance system compares a short exposure period with a long baseline; a trial’s control arm is followed for longer than its treatment arm; and the conditional proportion is then far from a half even when the rates are equal. With the second window four times as long as the first and the rates equal, .
The arrival coin does not notice. Its coverage is at every proportion, because the coin is uniform given the counts and the randomised binomial interval is exact at every ; the windows’ lengths enter only through . The exact conditional interval does notice. At it covers 98.78% at an expected total of ten and 97.59% at twenty, against 97.82% and 97.07% at one half, so an unbalanced design makes it more conservative, not less. Mid-p covers 97.47% at ten and 95.60% at twenty, against 95.89% and 95.43%.
The zero total’s share barely moves: the crossing below which it is the larger cause is at an expected total of 0.698 rather than 0.693. So the lesson of the balanced case carries over to the unbalanced one with more force. The zero total is a small-count curiosity; the discreteness of the conditional binomial is the conservatism, and an unbalanced design, which puts the conditional proportion near the boundary where every binomial interval oscillates most, makes it worse.
What exactness costs when one window is empty
A randomised interval buys its exact coverage in an unfamiliar currency, and the ratio makes the currency visible. Suppose the first window recorded no events and the second recorded five. Every conventional interval for the ratio then has a lower limit of zero: no events in the first window cannot rule out a first rate of nothing.
The randomised interval’s lower limit at a first-window count of zero is zero only when the coin exceeds 2.5%. When the coin falls below 2.5% — when the latest of the five arrivals, read in window-relative time, came before about 0.48 of the way through its window, since — the lower limit is positive, and the interval excludes a ratio of zero on a sample that contains no evidence against one. That happens on exactly 2.5% of such samples, and it is exactly the 2.5% the interval needs to miss on the lower side when the first rate really is close to zero. The upper limit moves with the coin too: at the coin’s middle value it is a ratio of 0.82 on five events, the mid-p value.
That is what a coin does: it spreads the binomial’s lumps of probability over a continuum, and occasionally the spreading puts an endpoint where no fixed rule would. The single-window essay faced the same thing at a count of one and argued that an interval read from the data, recomputable by anyone, is defensible where a drawn one is not. The ratio adds one more reason. A total of zero is the only outcome on which the interval says nothing, and it says nothing exactly; on every other outcome its errors are the 5% it promised, distributed evenly over the ratios a reader might be asking about rather than concentrated where the counts are smallest.
Across the ratio, at one total
The excess above was read at one proportion. A discrete interval’s coverage oscillates with the proportion, and the arrival coin’s should not.
At an expected total of five the arrival coin covers 95.034% at every proportion, flat. Mid-p’s coverage oscillates and its lowest point on the grid is 96.50%; the exact conditional interval’s lowest is 98.61%. At this total, then, mid-p — the usual pragmatic repair — never falls below 95% at all, and still spends most of its range more than a point above it. Its familiar under-coverage needs larger totals: the single-window essay found mid-p reaching 91.66% for one Poisson mean, and for the ratio the averaging over totals smooths most of that away at small expected counts.
What the conservatism costs in power
An interval that covers 98% is a test whose size is 2%, and a test that rejects less often than it may also rejects less often when it should. The power of the test each interval inverts makes that concrete.
The setting is a two-arm comparison with five events expected in the reference arm. At a true ratio of one the exact conditional test rejects 2.18% of the time, mid-p 4.11%, and the arrival coin 5.00% — five per cent less the zero total’s share, which at an expected total of ten is about two millionths. At a true ratio of two the three reject 17.8%, 23.9% and 24.5%; at three, 54.2%, 61.6% and 62.2%.
So the exact conditional test gives up about seven points of power at a doubling of the rate and eight at a tripling, in exchange for a size of 2.18% where 5% was allowed. Mid-p recovers most of that and the arrival coin the rest. The difference between mid-p and the arrival coin is small at these counts — under a point of power — and the difference between either and the exact test is not.
For a safety comparison that is the difference that matters. A monitoring board watching for a doubling of an adverse event’s rate, with five events expected in the reference arm, sees it flagged by the exact conditional test in fewer than one look in five and by the arrival-coin test in about one look in four — and pays for the difference only with the false alarms it was entitled to: the arrival test’s size is the 5% the board asked for, and the exact test’s 2.18% is a level nobody chose.
What it costs in width
The same comparison in the interval’s width, read on the scale of the conditional proportion where it is bounded.
At an expected total of two the exact conditional interval is 0.895 wide on the proportion scale, mid-p 0.853 and the arrival coin 0.806; at ten, 0.612, 0.563 and 0.553. The arrival coin is the narrowest at every total. Its advantage over the exact interval is about a tenth of the width at small and moderate totals; its advantage over mid-p shrinks with the total, from about five per cent at two to under two per cent at ten, because mid-p’s fixed coin at one half becomes a better stand-in for a uniform one as the binomial’s atoms shrink.
The objection, and the design that answers it
The objection to a randomised interval is that the randomness was added, and the coin that is already there answered it for a proportion by finding a coin in the data. For a ratio the answer is stronger. The arrival times are recorded in every surveillance system, every trial’s adverse-event log and every reliability study; nobody has to draw anything, and two analysts reading the same log compute the same interval.
The condition is the one the single window had. The times have to be uniform within each window given the counts, which holds when each window’s rate is constant across it. A rate that drifts within a window — more failures late in a service life, more cases late in a season — makes the latest arrival come later than a uniform would, and the coin leans. The single-window essay measured that lean for a linear drift and found a folded coin, read from distances to the window’s middle, that is immune to it; the same folded coin applies here window by window, and it would need checking against drifts that differ between the two windows, which is the realistic case for two arms of a trial.
Still open: rates that drift differently in the two windows
The next measurement is that check. In a two-arm comparison the arms’ rates can drift in different directions — a treatment’s adverse events concentrated early, the control’s spread evenly — and a coin read from both windows at once then mixes two leans. The folded coin is uniform given the counts under any drift symmetric about a linear trend within each window, whatever the two trends are, so it should survive; the last-arrival coin should not. Whether the folded coin’s coverage stays at when the windows drift in opposite directions, and what the last-arrival coin’s lean costs in coverage at the totals where the exact conditional test’s conservatism is largest, is a finite sum of the same kind as everything here and has not been computed.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The shortest interval is the one that misses — both name clopper–pearson, conservative interval, coverage, discreteness, interval width
- An interval that covers and says nothing — both name clopper–pearson, conservative interval, coverage, interval width
- A coverage table with its own error — both name coverage, discreteness, statistical power
- A quantile over the recent scores — both name coverage, exchangeability, statistical power
- A ratio that changes between blocks — both name conservative interval, coverage, interval width
- Marginal is not conditional — both name coverage, exchangeability, interval width
Named objects
A flat tag is an object no other essay names yet.
Clopper–PearsonConditional testConservative intervalCoverageDiscretenessExchangeabilityInterval widthMid-pRandomised intervalStatistical power