The interval that integrates
Worth reading first: What a credible interval covers · The weight that decides.
What a credible interval covers took a Bayesian interval for one proportion and put it through the frequentist’s own test: build every possible sample, count how often the interval contains the truth. The answer was that a credible interval on Jeffreys’ prior covers 95.7% at n = 20 and p = 0.1 where the textbook Wald interval covers 87.6% — the Bayesian interval beating the frequentist one on the frequentist’s criterion.
This essay does the same thing one level up, where the object being estimated is a group inside a population and the nuisance parameter is the population’s own spread. The result has the same shape and a larger margin.
What the interval is
Conditional on a known population spread τ, a group’s posterior is normal. That is the content of the weight that decides: the posterior mean is B·μ + (1 − B)·y with B = se²/(se² + τ²), and the posterior variance is (1 − B)·se² plus a term for not knowing μ either.
τ is not known. So the posterior for the group is that normal averaged over the posterior for τ:
p(θ | data) = ∫ N(θ ; centre(τ), spread(τ)) · p(τ | data) dτ
which is a mixture of normals — one component for every value of τ the data allows, weighted by how well that value explains the eight groups. The integral is one-dimensional and a few hundred evaluations of a normal density settle it.
Three consequences follow from the mixture being a mixture, and they are not the same consequence stated three ways.
It is wider
The obvious one, and the smallest of the three.
Thirty-one per cent of width buys sixteen points of coverage, which is a good trade by any accounting. It is worth noticing that the trade is not available the other way round: no amount of widening the plug-in interval by a constant factor would fix it, because the shortfall is not uniform across groups.
What sixteen thousand intervals is worth
The count deserves one qualification, because the intervals are not sixteen thousand independent readings.
Two thousand datasets of eight groups each produce sixteen thousand intervals, and the eight from one dataset share a single estimate of the population spread. So a dataset whose τ̂ came out low produces eight intervals that are all narrow together.
The effective number of independent readings is therefore nearer two thousand than sixteen thousand, and the standard error on a coverage is about 0.49 points rather than the 0.17 the larger count suggests.
That is enough for the comparison being made — the margins here are points rather than tenths of a point — and it is the right number to quote beside them.
It is heavier in the tails
A mixture of normals with a common centre and different spreads is not a normal. It is leptokurtic — more peaked in the middle and heavier in the tails than any single normal with the same variance.
That matters because a 95% interval is a statement about the tails and nothing else. Its endpoints are at the 2.5% and 97.5% points, which are exactly where the mixture differs most from the normal that shares its variance. Scaling up a normal interval until it has the mixture’s variance would still put the endpoints in the wrong places.
The practical version: the integrated interval is not the plug-in interval times 1.31. It is a different shape, and the difference is concentrated where the answer is read.
It has no τ in it at all
The third consequence is the one that changes what can be said.
The plug-in interval is a statement conditional on τ = 1.07. If asked what it would have been at τ = 2, there is an answer, and it is different. The integrated interval has no such conditional attached: τ has been integrated out, and what remains is a statement about the group.
This is the property where the two schools agree identifies as the Bayesian machinery’s actual advantage, and the hierarchical case is where it earns most. A frequentist confidence interval for a group in a hierarchy has to deal with τ as a nuisance parameter, and the standard devices for that — profile likelihood, a variance correction, a bootstrap over groups — are all approximations that need their own coverage checked. Integrating is exact by construction and needs a prior.
Counted, at both ends of the slider
The coverage figure above is at one population spread. The interesting question is where the two intervals differ, and it has the same answer as everything else in this field.
The shape is not the one this field assumed
Running the comparison across the whole slider produces a table that corrects the account the previous three essays have been giving, and it is worth setting out in full. Fifteen hundred datasets at each setting:
| true τ | integrated | plug-in |
|---|---|---|
| 0.25 | 98.6% | 92.3% |
| 0.5 | 97.5% | 83.6% |
| 1 | 95.2% | 78.6% |
| 2 | 94.2% | 88.0% |
| 4 | 94.8% | 93.7% |
The plug-in is worst in the middle. Not at the alike end, which is what “the plug-in fails where τ is hard to estimate” would predict. It recovers at both extremes, and for two different reasons.
At large τ it recovers because eight groups determine τ well and there is nothing left to integrate over. That is the expected reason.
At small τ it recovers for a reason that is almost an accident. When the groups really are nearly identical, the moment estimate collapses to zero — on more than half of datasets at τ = 0.25 — and the plug-in then reports complete pooling: eight identical estimates with a narrow interval around the grand mean. That answer happens to be nearly right, because the eight truths are all bunched near the grand mean too. The interval is far too narrow for the model that generated the data and still contains the truth most of the time, because the truth is where it is pointing.
So the plug-in’s coverage is a poor guide to whether the plug-in is behaving sensibly. At τ = 0.25 it covers 92.3% while reporting that eight distinguishable groups are one group.
And the integrated interval is conservative at small τ, at 98.6%. That is the flat prior on τ doing what a flat prior on a positive quantity does: keeping substantial mass at values of τ larger than the data supports, which widens every interval. A half-Cauchy with a well-chosen scale would be less conservative there, and would be an assumption. The trade is the one a prior on the spread measures.
The honest rule, which is narrower than the one the field started with: the integrated interval covers at least 94% at every setting measured and is conservative where the groups are alike; the plug-in is between five and sixteen points short across the whole middle of the range, and its recovery at the ends is not evidence that it is working.
The failures are correlated, and that changes what the number means
There is a second measurement hiding behind the coverage figure and it took a mistake to find.
The obvious standard error for a counted coverage is √(p(1−p)/n) with n the number of intervals. On sixteen thousand intervals that is 0.17 percentage points, and it is wrong, because the eight intervals from one dataset are not independent — they share the same τ̂ and the same estimated centre.
Computed over the right unit, which is the dataset, the standard error is 0.53 percentage points for the plug-in and 0.18 for the integrated interval. The first is three times the naive figure; the second is barely above it.
That gap is a measurement of something real. The integrated interval’s misses are scattered: one interval in twenty, roughly independently, roughly as a 95% interval should. The plug-in’s misses come in clumps. A study whose eight groups happened to look more alike than they are produces a τ̂ that is too small, which makes every one of its eight intervals too narrow at once. The failure is a property of the study, not of the interval, and nothing inside the study reveals it.
For a reader of one analysis, that is the more important of the two findings. A procedure that misses 5% of the time independently is a procedure whose errors average out across the eight statements in a report. A procedure that misses 21% of the time in clumps produces reports that are either mostly right or mostly wrong.
The clumping, counted directly
The standard errors above are an inference about correlation from the spread of a mean. The correlation can be counted head-on instead, and it is a better number to quote because it needs no reasoning at all: how often does every one of a study’s eight intervals cover?
If the eight failed independently, that share would be the per-interval rate raised to the eighth power. Over the same two thousand studies:
| covers every group | if the eight failed independently | |
|---|---|---|
| integrated | 68.9% | 67.4% |
| plug-in | 39.4% | 14.8% |
The integrated interval is within one and a half points of independence, which is what a procedure whose errors are unrelated to each other looks like. The plug-in is two and a half times above it.
That factor is the whole of the clumping claim, and reading it the friendly way round is instructive: the plug-in produces a fully correct set of eight statements far more often than its 79% per-interval rate would suggest. The price of that is the other tail. On the studies where τ̂ came out too small, the report is not wrong once — it is wrong repeatedly, in the same direction, with every group’s interval too narrow and every group’s estimate too close to the middle, and with no internal inconsistency to give it away.
A reader deciding whether to trust one analysis therefore faces a worse problem than the headline number suggests. Twenty-one per cent of intervals missing sounds like a tolerable error rate spread thinly. What it actually is, at eight groups, is a substantial minority of studies in which almost everything is wrong together.
The two-route check underneath
The whole apparatus rests on the posterior for τ, and the posterior for τ is computed with an algebraic shortcut: with a flat prior on μ, the population mean integrates out of the likelihood in closed form and leaves a one-dimensional function.
That shortcut has a second route to disagree with. The comparison integrates μ numerically on a grid of 1,201 points, states no identity, and shares only the normal density. The two agree to 6.5 × 10⁻¹⁰ on the density at every point of the τ grid, and to seven significant figures on its posterior mean. The residual is the second route’s own quadrature, which is the right direction: the route with more numerical integration in it is the less accurate one.
Without that check the coverage numbers would rest on an unverified piece of algebra doing the single most load-bearing job in the field.
The group the model is wrong about
Coverage averaged over eight groups drawn from the population is the right measurement for the model as stated, and it is not the only question a reader has. The one that usually matters is about a particular group, and often about the group that looks unusual.
Integrating over τ does not repair that. It cannot: the model says the groups come from a common population, the outlying group does not, and no amount of care about a nuisance parameter fixes a model that is wrong about the structure. What integration does do is make the failure visible in the interval rather than only in the estimate.
The mechanism is that an outlying group inflates the observed spread between the groups, which pushes the posterior for τ to the right, which widens every interval. The plug-in does the same thing to τ̂ — this is not an advantage of integration in kind — but the integrated interval also carries the uncertainty about how far right the posterior has been pushed, and one outlier in eight groups makes that uncertainty large. So the interval for the odd group comes out wide, which is the correct response to a group about which the population model is saying something the data disputes.
That is a real but modest defence. The honest summary is that a hierarchical interval is a statement made inside a model, and the counting in this essay establishes that the interval is right about what the model says. Whether the model is right about the groups is a different question, and it is the one when borrowing goes wrong is about.
Where it sits beside the rest of the site
Three intervals on this site have now had their coverage counted rather than asserted, and the pattern across them is worth stating.
For a proportion, coverage can be computed exactly by summing over the twenty-one possible outcomes. For a mean with unknown variance, the t interval covers exactly, by construction, and the check before the standard error is about when its assumptions fail. For a group in a hierarchy there is no exact answer available and the coverage has to be simulated, which is why every number here carries a standard error and why the standard error had to be computed over the right unit.
The direction of the results is consistent. An interval whose stated properties are checked usually turns out to have them; an interval whose stated properties are inherited from a conditional argument usually does not. The Wald interval inherits from the normal approximation. The plug-in hierarchical interval inherits from conditioning on an estimate. Both fail, and both fail worst where they are most needed.
What it costs to run
Nothing about this is expensive, which is worth saying because the reason empirical Bayes is the default is that it is fast.
The posterior for τ is one pass over a grid: at each point, a precision-weighted mean and a sum of eight log-densities. Two hundred and forty points is enough for four-digit accuracy on everything in this essay. The interval for one group is then a bisection on the mixture’s distribution function, which is another sixty passes over the same grid.
Counting the coverage — two thousand datasets, eight groups each, both intervals — takes under a minute, and most of that minute is building the endpoints rather than testing them. Whether an interval covers can be settled with a single evaluation of the mixture’s distribution function at the truth, since a central interval contains θ exactly when the posterior distribution function at θ lies between the two tail probabilities. That observation is worth fifty times the run time, and it is exact rather than an approximation to the interval.
The cost of the integrated interval, in other words, is roughly two hundred normal densities per group. The cost of the plug-in interval is one. That is the entire trade, and it buys sixteen percentage points of coverage and an error structure that does not fail in clumps.
What this field settles
Four essays and one claim, which can now be stated without hedging.
The shrinkage weight needs a population spread. The population spread is estimated from as many numbers as there are groups, which is usually few. Every consequence in this field follows from those two sentences and from the refusal to pretend otherwise:
- The estimate of the spread is the most extreme value available — exactly zero — on 32.6% of eight-group datasets that have a real spread in them, and the chi-square tail that predicts that rate needs no simulation to state.
- Substituting it and proceeding gives intervals that cover 78.8% of the time against a claimed 95%, and that fail together rather than separately: all eight cover on 39.4% of studies where independent failure would give 14.8%.
- Integrating over it instead gives 95.2%, at 31% more width, with an error structure that is within one and a half points of independent.
- The integral needs a prior on the spread, and on eight groups that prior never stops mattering — its influence falls from a factor of two to about 15% as the groups separate, and not to nothing.
The last of those is the honest limit of the repair and it is not a small one. What integration buys is not certainty about τ; it is a statement of uncertainty that is the right size. Where eight groups genuinely cannot settle how different eight groups are, the answer is a wide interval and an acknowledged dependence on what was assumed — which is what the data supports, and is more than the alternative offers.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- When the prior is confident and wrong — both name coverage, credible interval, flat prior, posterior, prior
- Pooling a proportion — both name hierarchical model, posterior, prior, shrinkage
- The p-value a replication gets — both name flat prior, posterior, prior, shrinkage
- What a prior is worth — both name coverage, flat prior, posterior, prior
- A group from the population's own tail — both name empirical bayes, hierarchical model, shrinkage
- An interval for something else — both name coverage, credible interval, posterior
Named objects
A flat tag is an object no other essay names yet.
CoverageCredible intervalEmpirical BayesFlat priorFrequentist interpretationHierarchical modelPosteriorPriorShrinkage