Intervals, counted

An interval that covers and says nothing

A procedure returning the whole line 95% of the time and the empty set otherwise has coverage exactly 95% at every parameter value. Two real intervals at forty observations have expected widths of 0.2418 and 0.2417 and worst-case coverages of 55.31% and 92.21%.

Worth reading first: What the 95% refers to.

The essay that opened this question established what a 95% interval is claiming: not something about the interval in front of a reader, but something about the procedure that produced it, countable by building every possible sample. Counting it showed the interval taught first covering 87.6% of the time.

The natural next question is whether an interval that does cover 95% is therefore a good interval. It is not, and the counterexample is short enough to write down.

Take a procedure that ignores the data entirely, flips a coin weighted 0.95, and returns the whole real line when the coin says heads and the empty set when it says tails. The parameter is in the interval exactly when the coin says heads, so the coverage is exactly 95% at every parameter value — not asymptotically, not approximately, and not only for some parameters. By the definition of coverage it is a perfect 95% procedure.

A 95% procedure that never looked at the data. A procedure that returns the whole real line with probability 0.95 and the empty set otherwise has coverage exactly 95% at every parameter value — the parameter is in the interval whenever the coin says so, and the coin does not depend on anything. Its expected width is not a number. Beside it, a real interval for a proportion at n = 40: coverage 95.21% averaged across the range, and finite in every draw. A definition of validity that admits the first has not said what an interval is for.
Fig. 1 A procedure that returns the whole real line with probability 0.95 and the empty set otherwise, beside a real interval for a proportion at forty observations. The first has exactly the advertised coverage and an expected width that is not a number.

Coverage is a necessary property and it is not a sufficient one. What else has to be said is the subject of the rest of this essay, and it turns out to be at least two more things.

Why the coin is not a trick

The usual response to the coin is that it is a degenerate case and real procedures do not behave like it. That is true and it is the wrong lesson, because the coin’s failure is not a failure of degeneracy — it is a failure of ignoring the data, and the interesting question is what requiring an interval to use the data actually buys.

The definition of coverage says nothing about the data at all. It says: over repeated samples, the procedure’s output contains the parameter 95% of the time. A procedure can achieve that by looking at the sample and building a sensible interval around what it saw, or by not looking and flipping a coin, and the definition does not distinguish them.

That is not a quirk of an artificial example. It is the reason a valid procedure can be a poor one in ordinary ways too: an interval can be valid and much too wide, valid and centred on the wrong place with compensating width, or valid on average and badly invalid where the parameter actually is. Each of those is a real procedure that the coin makes visible.

The coin’s other use is that it is exactly 95%, where every real procedure for a proportion is something else — a consequence of discreteness that the earlier essay had to work around. So the one procedure with perfect coverage in this whole essay is the one nobody would use, which is as clean a statement of the essay’s point as is available.

The second number is the width

The coin fails on width, and width is the obvious repair: require the interval to cover 95% of the time and to be short. That rules the coin out immediately, since its expected width is infinite.

It does not rule much else out.

Three promises, and no procedure keeps all three. Average coverage and worst-case coverage for four 95% intervals for a proportion at n = 40, computed exactly. Their expected widths are 0.2418, 0.2417, 0.2472, 0.2641 in the same order. The textbook interval and the score interval have the same expected width to four digits — 0.2418 and 0.2417 — and worst-case coverages of 55.31% and 92.21%. The exact interval never breaks its promise and is 9.3% wider than the score interval to do it. Each of the three columns orders the four procedures differently.
Fig. 2 Average coverage and worst-case coverage for four 95% intervals for a proportion at forty observations, computed exactly over all forty-one counts. Their expected widths are 0.2418, 0.2417, 0.2472 and 0.2641 in the same order.

The textbook interval and the score interval have expected widths of 0.2418 and 0.2417 — the same number to four digits. Two procedures indistinguishable on width, averaging 91.22% and 95.21% coverage, and falling to 55.31% and 92.21% at their worst proportions.

A reader given both intervals for the same data would see two intervals of nearly the same length. A reader given both procedures’ expected widths would see no difference at all. The thing that separates them is a third number.

The third number is the worst case

Coverage is not one number. It is a function of the parameter, and reporting it as a single figure means averaging it — over what, is a choice nobody usually states.

Three procedures, and only one of them keeps its promise. The exact coverage of three 95% intervals for a proportion at n = 40, computed by summing over all 41 possible counts rather than simulated. The textbook interval averages 91.22% and falls to 55.31% at a proportion of 0.020. The score interval averages 95.21% and never falls below 92.21%. The exact interval never falls below 95.19% and averages 97.00%, which is its own kind of failure.
Fig. 3 The exact coverage of three 95% intervals for a proportion at forty observations, computed by summing over all forty-one possible counts. The textbook interval averages 91.22% and falls to 55.31%; the score interval averages 95.21% and never falls below 92.21%.

The textbook interval’s average of 91.22% reads as “a few points short”. Its coverage at a true proportion of 0.02 is 55.31%, which is not a few points short of anything — it is a procedure that misses the parameter nearly half the time, labelled 95%.

The average and the worst case are different statements about different things, and a confidence level is a claim about the second. “This procedure covers at least 95% of the time” is what the phrase means, and a procedure whose coverage averages 95% while dipping to 55% has not made that claim, it has made a weaker one that sounds identical.

Four procedures and three orderings

The four procedures in the sweep are worth reading as a table rather than a ranking, because each column orders them differently.

By average coverage, closest to 95% first: the score interval at 95.21%, then adding two successes and two failures at 95.86%, then the exact interval at 97.00%, then the textbook interval at 91.22%.

By worst-case coverage, highest first: the exact interval at 95.19%, then adding two and two at 93.82%, then the score interval at 92.21%, then the textbook interval at 55.31%.

By expected width, narrowest first: the score interval at 0.2417, the textbook interval at 0.2418, adding two and two at 0.2472, the exact interval at 0.2641.

Three orderings, three different procedures in first place, one procedure last in two of them. There is no procedure that is best on all three and no reason there should be: the three properties are not three views of one thing, they are three separate demands, and a construction that satisfies one exactly is trading against the others to do it.

The averages converge and the promises do not

A reader might reasonably expect the gap to close with data. It does, on one of the two numbers.

The average catches up and the worst case does not. The average and the worst-case coverage of two procedures, against the number of observations, computed exactly at every size. The textbook interval's average climbs from 79.02% at 10 observations to 93.30% at 100, which looks like convergence. Its worst case climbs from 18.29% to 80.22% and is still 13.1 points behind its own average. The score interval's two lines are within 1.7 points of each other throughout.
Fig. 4 The average and the worst-case coverage of two procedures against the number of observations, computed exactly at every size. The textbook interval’s average climbs from 79.02% to 93.70%; its worst case climbs from 18.29% to 84.74% and is still ten points behind its own average at a hundred observations.

At ten observations the textbook interval averages 79.02% and covers 18.29% at its worst proportion. At ninety-four it averages 93.70% — which looks like a procedure settling onto its promise — and covers 84.74% at its worst.

The average converges and the minimum lags, and the lag is where the promise lives. The score interval’s two lines run within a couple of points of each other at every size, which is what a procedure whose average is informative looks like.

This is the same reading a neighbouring essay in this field makes about coverage that does not improve monotonically: the behaviour of an interval for a discrete parameter is not summarised by any single number, and a reader who takes one is taking the one the procedure’s author found flattering.

Where the textbook interval’s failure comes from

The 55.31% is at a proportion of 0.02, and the mechanism is worth stating because it explains why no amount of averaging rescues it.

The textbook interval is p^±zp^(1p^)/n\hat p \pm z\sqrt{\hat p(1-\hat p)/n}, and its width is a function of the observed proportion. At a true proportion of 0.02 with forty observations the most likely count is zero, which gives p^\hat{p} = 0 and a width of exactly zero: the interval is the single point 0, and it misses any true proportion above it. A count of one gives a width of 0.096 around 0.025, which happens to cover. So the coverage at that proportion is essentially the probability of seeing at least one event, and at 0.02 with forty draws that is about 55%.

The failure is therefore not noise or slow convergence. It is a construction whose width is estimated from the same quantity it is an interval for, collapsing exactly where that quantity is small. The score interval avoids it by inverting a test rather than plugging in an estimate, so its width at a count of zero is not zero — which is the whole of the difference between them and is invisible in any comparison of expected widths.

That is the same defect the delta method exhibits at a flat point: a standard error read off an estimate is zero wherever the estimate happens to make it zero, and the interval built on it is a point.

The fourth promise, and what it costs

There is a procedure that keeps the promise as stated — coverage at least 95% at every proportion — and the sweep above says exactly what it costs.

The exact interval’s worst-case coverage is 95.19%, above the level everywhere by construction. Its average is 97.00% and its expected width is 0.2641, which is 9.3% wider than the score interval’s.

That is the trade in its clearest form. A procedure can honour a floor everywhere, and it pays in width, and the payment is not small: nine per cent of width is what the last ten per cent of a trial’s units would have bought. Nothing here says that is a bad bargain; it says it is a bargain, and that a reader shown only a coverage figure has not been told which side of it a procedure sits on.

There is a second reading of that 9.3% worth having. Width falls like 1/n1/\sqrt{n}, so a nine per cent wider interval is what a trial 19% smaller would have produced at the same construction. Paying for a guaranteed floor therefore costs about a fifth of the study — expressed in the currency a study is built in rather than in a ratio of widths, which is the form the decision is actually made in.

And over-covering is a defect too. An interval covering 97% when it says 95% is wider than it needs to be, which means a study concluded less than its data supported. That is a real cost borne by a reader who cannot see it, and it is the reason the exact interval is not simply the right answer.

What “95%” would have to mean to be checkable

Collecting the three properties gives a version of the phrase that a reader could actually verify, and it is longer than the phrase.

A procedure described as 95% should mean: its coverage is at least 95% at every value of the parameter, and its expected width is as small as that permits. The first clause is what the exact interval satisfies and the textbook interval does not; the second is what the exact interval sacrifices.

Those two cannot both be satisfied exactly for a discrete parameter, which is not a defect in any procedure but a fact about the problem: the coverage of any procedure for a proportion is a step function of the parameter, so a floor can only be met by exceeding it nearly everywhere. The available constructions are therefore points on a trade, and the honest description of one is which point it sits at — which is three numbers and not one.

The word the phrase is missing is the quantifier. “Covers 95% of the time” is ambiguous between “for every parameter” and “on average over parameters”, the two readings differ by thirty-six points for the interval taught first, and neither the phrase nor any software output says which is meant.

What a defensible report looks like

Prefer exact arithmetic where the sample space is finite. Every number here is a finite sum over the n + 1 possible counts, and the differences being reported are a few points wide — small enough that a simulation at any affordable size would leave them arguable. That is the same preference an enumerated coverage rests on, and it is available wherever the parameter is discrete.

Name the procedure. “A 95% confidence interval for the proportion” describes at least four different objects whose worst-case coverages run from 55% to 95%, and software defaults differ. A reader told which construction was used can look up its properties; a reader told only the level cannot, and the level is the number that does not vary between them.

State the worst case, not the average. For a discrete parameter it is computable exactly — a finite sum over the possible counts at each of a grid of parameter values — so it is a number rather than an estimate. Every figure in this essay is such a sum.

State the expected width beside it. Two procedures with the same coverage profile and different widths are different procedures, and the one a study should use is the shorter. This is the same pairing a width reported beside a coverage is about: neither number means much without the other.

Prefer a procedure whose average and worst case are close. That is a property a reader can use directly: where the two are close, the average is informative and a single number will do; where they are far apart, the average is a summary of a function that is not well summarised. At forty observations the score interval’s two differ by three points and the textbook interval’s by thirty-six.

Read a discrete parameter’s coverage as a function, not a number. It oscillates, and the oscillation does not settle with the sample, so any single figure quoted for it is a summary of a curve whose shape is the finding. The same is true of the width.

And say what the average was taken over. “Averages 95%” is a statement about a weighting of the parameter space, and the weighting here is uniform on (0.02, 0.98), which is a choice. Weighted towards the middle of the range the textbook interval looks much better; weighted towards the ends it looks much worse. An average with no stated weighting is not a measurement.

What is claimed here and what is not

Everything is exact arithmetic, not simulation. For a proportion the sample space is the n + 1 possible counts with known probabilities, so the coverage and the expected width at any parameter value are finite sums. There is no Monte Carlo error anywhere in this essay, which matters because the differences being shown between the score interval and the textbook one are a few points wide at most proportions and would be arguable if they were simulated.

The coin is an existence proof, not an argument about real intervals. No one proposes it and no software implements it. Its job is to show that the definition of coverage does not by itself pick out useful procedures, so that the question “what else has to be required” has to be answered rather than assumed away.

The four procedures are not the only four. Several more constructions for a proportion exist — mid-p variants, likelihood-ratio intervals, Bayesian intervals read as frequentist ones — and each sits at its own point on the same trade. The four here are chosen because they are the ones software offers by default and because they span the range from “ignores the discreteness” to “honours the floor everywhere”.

The range is (0.02, 0.98) and the size is forty. The textbook interval’s worst case is at the edge of that range, and extending the range towards zero makes it worse without bound — its coverage at a true proportion of 0.001 with forty observations is essentially nothing, since almost every sample gives a count of zero and an interval of zero width. Where the range is cut is therefore a choice that flatters the textbook interval, and it is cut where it is so that the comparison is not about a degenerate case.

And “expected width” treats an infinite interval as infinite. The coin’s expected width is not a large number, it is undefined, and the figure says so rather than substituting a convenient bound. That is the honest reading and it is also the one that makes the point: a summary that averaged width would have to decide what to do with an infinite draw, and every choice it could make would hide something.

Two properties that usually move together, and the case where they do not

The reason this essay needs the coin at all is that on ordinary procedures the three properties move together, so a reader who checks one has effectively checked the others.

A construction that covers well is usually one that uses the data sensibly, and a construction that uses the data sensibly is usually short. The score interval is a case: best or near-best on all three columns at forty observations. Nothing goes wrong, and a reader who checked only its coverage would have been right about everything else.

What the coin and the textbook interval supply is the two ways that correlation can break. The coin has perfect coverage and no other virtue, because it does not use the data. The textbook interval has a competitive width and a catastrophic worst case, because it uses the data in a way that collapses where the data is thin.

Finding such a case is how a conflated account gets taken apart, and it is a move that recurs throughout the subject — a bias that cancels while the coverage does not is the same shape in a different field. Where two properties usually travel together, the informative measurement is the one that separates them.

Still open: what the interval in front of a reader is worth

Every property in this essay is a property of a procedure, averaged over samples that did not happen. That essay made the point and this one has stayed inside it, comparing procedures on their long-run behaviour.

A reader has one interval. The question neither essay has touched is whether anything can be said about that interval — whether a study that observed a particular count can say more than “the procedure that produced this covers 95% of the time”. For a proportion there is a specific reason to think it can: the counts at the ends of the range produce intervals with very different reliability from the ones in the middle, and the count is observed.

What such a statement would look like, whether conditioning on the observed count leaves a statement with a frequency interpretation at all, and what it does to the three numbers above, is the next thing this field has to work out.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Binomial proportionClopper–PearsonClosed formConfidence intervalConservative intervalCoverageExact enumerationFrequentist interpretationInterval widthNominal levelOvercoverageWald intervalWilson intervalWorst case