The shortest interval, and the one that does not move
Worth reading first: What a credible interval covers.
A posterior is a distribution, and a distribution does not come with an interval attached. Asking for a set holding 95% of it is asking for one of infinitely many sets, and two of those get used.
The equal-tailed interval runs from the 2.5% quantile to the 97.5% quantile: 2.5% of the posterior is left outside at each end. The shortest interval — the highest-posterior-density interval — is the narrowest set holding 95%, wherever that happens to be. Both hold the same probability. On a skewed posterior they are not the same interval, and the difference is not small.
Why the shorter one is shorter
The mechanism is worth stating because it is the whole of the argument, and it is one sentence of calculus.
Move the lower endpoint of any interval right by a small amount and the interval loses probability at a rate equal to the density there. To hold the level, the upper endpoint has to move right too, and it gains probability at a rate equal to the density there. The net change in length is positive when the density at the lower end is higher, negative when it is lower, and zero when the two are equal.
So the shortest interval is the one whose endpoints sit at the same density. That is the definition of a highest-density interval — everything inside is denser than everything outside — and it is the same statement.
For a posterior with a long right tail, the equal-tailed interval does not satisfy it: its lower endpoint sits high on the steep left side, its upper endpoint sits low on the shallow right one. Sliding both to the left trades a little probability at a high density for a little at a low one and comes out ahead on length.
What the trade is worth, counted at every count
The gain is a real number and it is available for every possible dataset, because at twenty trials there are twenty-one of them.
Across the twenty-one counts the saving averages 4.86% and reaches 22.41%, and where it reaches that is the informative part: at the boundary counts, where the posterior is most skewed. At zero of twenty the equal-tailed interval runs from 0 to 0.1166 and the shortest runs from 0 to 0.0905 — a fifth narrower — because the equal-tailed interval is obliged to leave 2.5% below a lower endpoint the posterior has almost no room for.
The saving also falls with the sample size, and it falls fast. At ten trials it averages 7.00%, at forty 3.14%, at eighty 1.94%. That is the posterior becoming symmetric: as the data accumulates, a Beta posterior approaches a normal, and on a symmetric density the two intervals coincide. The shortest interval is worth having exactly where the sample is small, which is also where every other property of an interval is at its least reliable.
What it costs
A width is not worth reading without a coverage beside it, and the shortest interval is the one that misses already made the general case for frequentist intervals: shortness is free to manufacture, and the cheapest way to get it is to stop covering.
The Bayesian version is not the same argument, because neither of these intervals is claiming a coverage. Both claim a posterior probability and both have it, exactly. Coverage is the unadvertised property, and it is checkable by the same finite sum that measures any other interval’s.
Over a hundred and ninety-nine true proportions the equal-tailed interval averages 95.03% and the shortest averages 93.64%. It falls below 90% at 28 of them against the equal-tailed interval’s 4. At a true proportion of 0.1 — the setting where a credible interval’s coverage was first counted here — the equal-tailed interval covers 95.68% and the shortest covers 86.72%.
Nine points of coverage for seven per cent of width is a bad trade at that setting and a defensible one elsewhere: the shortest interval covers more than the equal-tailed one at 30 of the hundred and ninety-nine proportions and less at 72. It is not uniformly worse. It is worse on average, worse at its worst, and worse where proportions are usually small, which is most of the places anybody reports one.
None of that is an argument against the shortest interval on its own terms. It is the observation that a set chosen to be short is not a set chosen to contain the truth, and that the two criteria come apart on exactly the posteriors where the choice matters.
The objection that settles it
The width and the coverage are both arguments about how well each interval performs. There is a third consideration, and unlike the first two it is not a trade at all.
Nothing in the problem says the quantity of interest is the proportion. A reader may want the odds, or the log-odds, or the reciprocal. A parameter is whatever the question is about, and the same posterior supports all of them.
The equal-tailed interval transforms exactly. Its endpoints are quantiles, and a quantile of a monotone function of a random variable is that function of the quantile — so the 2.5% point of the odds is the odds of the 2.5% point, character for character. The interval for the odds is the interval for the proportion, rewritten.
The shortest interval does not. It is defined by a density, and a density changes under a change of variable: the density of is the density of multiplied by . Two points that sat at equal density on one scale do not on the other, so the set that was shortest stops being shortest.
Shortest, on which scale
The third interval in every figure above is the one this suggests: compute the shortest interval on the log-odds scale, and map it back to proportions.
It is a legitimate interval — it holds 95% of the posterior, because probability is what a change of variable preserves — and it is a third set, different from both of the others. At two of twenty it runs from 0.0247 to 0.3057.
The reversal is not a surprise once stated, and the second half of it is worth dwelling on. At zero of twenty the shortest interval on the proportion scale runs from exactly 0 to 0.0905. Zero proportion is minus infinity in log-odds. The interval chosen for being the shortest available is, on the other scale, infinitely long — not merely longer than its rivals but unbounded, while the equal-tailed interval there is 8.60 units wide and the log-odds interval is 7.91.
So “shortest” is not a property of the interval. It is a property of the interval and a scale, and the scale was chosen by whoever wrote the parameter down.
Three scales, and none of them is the right one
The obvious reply is that the log-odds is the natural scale for a proportion, so the interval computed there is the one to use and the difficulty dissolves.
It does not, because there is more than one candidate and they disagree. The log-odds is the canonical link, the scale on which a logistic regression is linear. The arcsine — — is the variance-stabilising transform, the scale on which the sampling variance of the estimate stops depending on where the estimate is, which is the property that makes proportions from studies of different sizes averageable at all. And the proportion itself is what gets reported and what a reader understands. Each is defensible and each names a different interval.
At two successes in twenty, under Jeffreys’ prior, the three shortest intervals and the equal-tailed one are four different sets, and the widths read differently depending on the ruler:
| computed on | the interval | width on p | width in log-odds | width in arcsine |
|---|---|---|---|---|
| the proportion | 0.0093 to 0.2540 | 0.2447 | 3.5935 | 0.4317 |
| the log-odds | 0.0247 to 0.3057 | 0.2810 | 2.8545 | 0.4279 |
| the arcsine | 0.0185 to 0.2721 | 0.2537 | 2.9894 | 0.4125 |
| equal-tailed | 0.0214 to 0.2839 | 0.2625 | 2.8986 | 0.4152 |
Every bold entry is on the diagonal. Each shortest interval is the narrowest on the scale it was computed on, and on neither of the others — three for three, and it holds at every count checked.
Read the rows the other way and the equal-tailed interval is the only one that never wins and never badly loses: second on the log-odds ruler, second on the arcsine, third on the proportion, and never worse than 7.3% off the best. It is not competing, because it was not chosen for width. It is the same set on every row, which is what the other three are not.
What survives
Three claims come out of this and they are of different strengths, which is worth separating because they are usually run together.
Shortness is real and small. Averaged over the counts at twenty trials it is 4.86% of width, and it falls to under two per cent by eighty trials. An interval procedure that buys 4.86% is not nothing, and it is not the difference between a usable study and an unusable one.
Its coverage cost is real and larger. One and four tenths of a point on average, four points at the worst case, nine at the setting proportions are most often reported from. Coverage is not what either interval claims, and a reader who assumes it holds — which the essay that first counted a credible interval’s coverage found to be the usual assumption — is worse served by the shorter one.
And only one of the two is a statement about the parameter. That is not a trade. A reader who asks for a 95% interval for the odds and is handed the equal-tailed interval’s endpoints put through the odds has been given the same statement in other units. A reader handed the shortest interval’s endpoints has been given a set that is not the shortest for the odds and was never chosen for anything about the odds.
The three do not even sit on a frontier, which is the shape a trade would take. At a true proportion of a tenth, and at twenty trials:
| mean coverage | worst coverage | below 90% | mean width | |
|---|---|---|---|---|
| shortest on the proportion | 93.64% | 85.54% | 28 of 199 | 0.2301 |
| equal-tailed | 95.03% | 89.64% | 4 of 199 | 0.2485 |
| shortest on the log-odds | 96.05% | 90.46% | 0 | 0.2716 |
Those three rows are ordered — more width, more coverage, in step — so on this evidence the choice looks like an ordinary trade after all, with the log-odds interval buying its two extra points with nine per cent of width. The reason it is not is that the ordering is an accident of this posterior’s skew. The shortest interval on the proportion scale is displaced towards zero and a proportion of a tenth is a small number; had the truth been near a half, all three would be the same interval, and had it been near one they would be ordered the other way. A trade is a relation that holds across the range. This is one that changes sign inside it, which the coverage figure shows directly: the shortest interval covers more than the equal-tailed one at thirty of the hundred and ninety-nine proportions.
That figure is the honest case for the other side and it should be read as one. At the boundary counts the equal-tailed interval’s rule is doing no work — there is no lower tail to cut — and it pays the full 22.41% anyway. An analyst whose data is a count of zero has the strongest reason available to prefer the shorter interval, and it is the same analyst whose posterior is most skewed and whose coverage is therefore worst.
Where the frequentist version of this argument does not reach
It is tempting to file all of this under the width-against-coverage trade that the interval field has been making since its first essay, and the filing is wrong in a way worth naming.
There, shortness and coverage are the same quantity seen twice: the Wald interval is narrow because its normal approximation understates the uncertainty, so its narrowness is its failure. Widening it repairs both.
Here they are genuinely different. The shortest interval holds exactly 95% of the posterior — its probability statement is correct to machine precision, not approximately — and it is shorter for a reason that has nothing to do with any approximation failing. It is shorter because it is placed where the density is. The coverage it loses is lost through a different door: an interval placed by density is placed closer to zero, and a truth on the other side of the mode is missed more often.
Two intervals can both be exactly right about what they claim and disagree about what they contain. That is not available in the frequentist setting, where being right about what is claimed is covering. It is what keeps turning up wherever these two vocabularies meet: a Bayesian object does what it says, and what it says is not what a reader assumes.
What is claimed here, and what is not
The claim is a comparison of three 95% intervals from one posterior: their widths at every count, the frequentist coverage of each computed exactly rather than simulated, and the fact that two of the three change under a change of variable while one does not. Everything is a finite sum over the twenty-one possible outcomes, so no number here carries a simulation error.
What stays out: intervals at credible levels other than 95%, where the arithmetic is identical and the sizes differ; the shortest region for a posterior with two modes, which is not an interval at all, and which is left undefined here rather than approximated; and decision-theoretic accounts of which interval a stated loss function selects, which is a different question from either of the two asked here. A U-shaped Beta posterior — both shape parameters below one, which Jeffreys’ prior with no data is — has a highest-density region in two pieces, and every figure above would have drawn the wrong object for it.
One limitation deserves more than a clause, because it is the honest boundary of the coverage comparison. Averaging coverage over a hundred and ninety-nine equally spaced true proportions is averaging against a flat weighting, and a flat weighting is itself a prior. The numbers reported as “mean coverage” are therefore what an analyst would experience across a career of studies whose true proportions were spread evenly over the unit interval, which nobody’s career looks like. The worst case needs no such qualification — it is a minimum over the range, not an average against a weighting — which is why the comparison is stated in both forms and why the second is the one carrying the argument. The same objection applies to every average of coverage over a range of truths, and it is worth making once here rather than attaching it to each.
The coverage curves also carry an oscillation that none of this discussion touches. A proportion has finitely many possible outcomes, so coverage is a step function of the truth, and every curve above is a sawtooth rather than a smooth line. That is discreteness rather than noise, it is the same phenomenon at work in all three curves, and it is why the comparisons here are read off worst cases and averages rather than off the value at any single true proportion.
Still open: the interval for something else
The transformation argument was used here to settle a comparison and then dropped, and it is the more useful half. If an interval’s endpoints can be put through a monotone function and the result is the interval for the transformed quantity, then an interval for the odds, the log-odds, a relative risk or a reciprocal is free — and the question is why anybody computes a fresh standard error on the new scale instead.
The answer is that the fresh computation is what the delta method is, that it gives a different interval, and that the difference is not a rounding: at twenty trials and a true proportion of a tenth the transformed interval covers 87.60% and a delta-method interval on the log-odds covers 83.52%. And a point estimate does not transform at all, which is the part that surprises. That is what happens when the quantity itself changes.
The check, and the refusal
Three claims are gated. That the shortest interval holds the probability it claims and that its endpoints sit at equal density, both to within a millionth — an optimisation result stated as the condition that characterises it rather than as the search that found it. That no interval of the same posterior probability is shorter, checked against a sweep of a thousand of them at every lower-tail position. And that the equal-tailed interval’s endpoints put through the odds are the quantiles of the odds’ own posterior, required to agree to , because that one is an identity and a tolerance there would hide a monotonicity bug rather than a numerical one.
A fourth is gated because the three-scale table would otherwise be a table nobody checked: at each of four counts, each scale’s own shortest interval must be the narrowest on that scale and on neither of the others — three for three, twelve times over. Beside it, the symmetric count is required to collapse all four intervals into one, which is the case that says the disagreement needs skew to produce it rather than arising from the arithmetic.
The refusal is the claim the essay turns on, and it is required to fail: the shortest interval must move when the scale changes. Recomputed on the log-odds scale it must be a different set, and it is — the two endpoints move by 0.0672 in total at two of twenty. If that check ever passed, the comparison drawn from it would be measuring the same arithmetic twice, and the essay’s central distinction would be a distinction with nothing behind it. Beside it, the equal-tailed interval computed the two ways must agree to , which is the two-route habit applied to a quantity that is supposed not to move.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- When the prior is confident and wrong — both name coverage, credible interval, interval width, posterior
- A hole no sample size fills — both name coverage, discreteness, jeffreys' prior
- Pooling a proportion — both name discreteness, log-odds, posterior
- The interval that integrates — both name coverage, credible interval, posterior
- What a guaranteed minimum costs — both name coverage, discreteness, interval width
- A coverage table with its own error — both name coverage, discreteness
Named objects
A flat tag is an object no other essay names yet.
CoverageCredible intervalDiscretenessEqual-tailed intervalHighest posterior densityInterval widthJeffreys' priorLog-oddsMonotone transformationOddsPosterior