The spread, and its own uncertainty

The interval that integrates

A credible interval for one group in a hierarchy has to average over every value the population spread might take. That averaging is what makes it cover — 95.2% against the plug-in's 78.8% — and it costs 31% more width, a heavier tail, and a mixture rather than a normal.

Worth reading first: What a credible interval covers · The weight that decides.

What a credible interval covers took a Bayesian interval for one proportion and put it through the frequentist’s own test: build every possible sample, count how often the interval contains the truth. The answer was that a credible interval on Jeffreys’ prior covers 95.7% at n = 20 and p = 0.1 where the textbook Wald interval covers 87.6% — the Bayesian interval beating the frequentist one on the frequentist’s criterion.

This essay does the same thing one level up, where the object being estimated is a group inside a population and the nuisance parameter is the population’s own spread. The result has the same shape and a larger margin.

Counted coverage of two 95% intervals, over 2,000 datasets. Each point is one of the eight groups, at its own standard error. The integrated interval covers 95.2% overall against its stated 95%; the plug-in covers 78.8%, and its shortfall grows with the group's standard error — from 86.1% at se 0.5 to 77.0% at se 2.1. The mean widths are 3.48 and 2.66.
Fig. 1 Sixteen thousand intervals, one per group per dataset, over two thousand datasets generated from a known truth. Both intervals claim 95%. The one that integrates over the population spread covers 95.2%; the one that estimates the spread and substitutes it covers 78.8%.

What the interval is

Conditional on a known population spread τ, a group’s posterior is normal. That is the content of the weight that decides: the posterior mean is B·μ + (1 − B)·y with B = se²/(se² + τ²), and the posterior variance is (1 − B)·se² plus a term for not knowing μ either.

τ is not known. So the posterior for the group is that normal averaged over the posterior for τ:

p(θ | data) = ∫ N(θ ; centre(τ), spread(τ)) · p(τ | data) dτ

which is a mixture of normals — one component for every value of τ the data allows, weighted by how well that value explains the eight groups. The integral is one-dimensional and a few hundred evaluations of a normal density settle it.

Three consequences follow from the mixture being a mixture, and they are not the same consequence stated three ways.

It is wider

The obvious one, and the smallest of the three.

The same group, with τ known and with τ integrated overTwo posteriors for one group's value. The narrow one conditions on the moment estimate τ̂ = 1.07 and treats it as known; the wide one integrates over everything the eight groups leave open about τ. The intervals are 4.11 and 6.03 wide, a factor of 1.47, and their centres differ by only 0.257.00.1000.2000.3000.400-2024the value of group 8, whose standard error is 2.1posterior densityplug-in, 4.11 wideintegrated, 6.03 widethe group's own meangroup 8 of 8, se = 2.1widths 4.11 against 6.03
Fig. 2 The two posteriors for one group, drawn on the same axes. The narrow one conditions on τ̂ = 1.07; the wide one integrates. The intervals are 4.11 and 6.03 wide. Averaged over the whole simulation the widths are 2.66 and 3.48, so the integrated interval is 31% wider.

Thirty-one per cent of width buys sixteen points of coverage, which is a good trade by any accounting. It is worth noticing that the trade is not available the other way round: no amount of widening the plug-in interval by a constant factor would fix it, because the shortfall is not uniform across groups.

What sixteen thousand intervals is worth

The count deserves one qualification, because the intervals are not sixteen thousand independent readings.

Two thousand datasets of eight groups each produce sixteen thousand intervals, and the eight from one dataset share a single estimate of the population spread. So a dataset whose τ̂ came out low produces eight intervals that are all narrow together.

The effective number of independent readings is therefore nearer two thousand than sixteen thousand, and the standard error on a coverage is about 0.49 points rather than the 0.17 the larger count suggests.

That is enough for the comparison being made — the margins here are points rather than tenths of a point — and it is the right number to quote beside them.

It is heavier in the tails

A mixture of normals with a common centre and different spreads is not a normal. It is leptokurtic — more peaked in the middle and heavier in the tails than any single normal with the same variance.

That matters because a 95% interval is a statement about the tails and nothing else. Its endpoints are at the 2.5% and 97.5% points, which are exactly where the mixture differs most from the normal that shares its variance. Scaling up a normal interval until it has the mixture’s variance would still put the endpoints in the wrong places.

The practical version: the integrated interval is not the plug-in interval times 1.31. It is a different shape, and the difference is concentrated where the answer is read.

It has no τ in it at all

The third consequence is the one that changes what can be said.

The plug-in interval is a statement conditional on τ = 1.07. If asked what it would have been at τ = 2, there is an answer, and it is different. The integrated interval has no such conditional attached: τ has been integrated out, and what remains is a statement about the group.

This is the property where the two schools agree identifies as the Bayesian machinery’s actual advantage, and the hierarchical case is where it earns most. A frequentist confidence interval for a group in a hierarchy has to deal with τ as a nuisance parameter, and the standard devices for that — profile likelihood, a variance correction, a bootstrap over groups — are all approximations that need their own coverage checked. Integrating is exact by construction and needs a prior.

Counted, at both ends of the slider

The coverage figure above is at one population spread. The interesting question is where the two intervals differ, and it has the same answer as everything else in this field.

Counted coverage of two 95% intervals, over 2,000 datasets. Each point is one of the eight groups, at its own standard error. The integrated interval covers 97.5% overall against its stated 95%; the plug-in covers 83.5%, and its shortfall grows with the group's standard error — from 90.1% at se 0.5 to 81.3% at se 2.1. The mean widths are 3.08 and 2.11.
Fig. 3 The same comparison on data whose groups are more alike. The plug-in has improved — 83.6% against 78.6% — and the integrated interval has gone the other way, to 97.5%. Neither movement is what the argument so far predicts.
Counted coverage of two 95% intervals, over 2,000 datasets. Each point is one of the eight groups, at its own standard error. The integrated interval covers 94.8% overall against its stated 95%; the plug-in covers 93.8%, and its shortfall grows with the group's standard error — from 95.1% at se 0.5 to 93.0% at se 2.1. The mean widths are 4.62 and 4.47.
Fig. 4 And where the groups are clearly different. Both cover close to 95%, because eight groups at this separation determine τ well enough that there is little left to integrate over. The plug-in is a good approximation here, and here is where it is not needed.

The shape is not the one this field assumed

Running the comparison across the whole slider produces a table that corrects the account the previous three essays have been giving, and it is worth setting out in full. Fifteen hundred datasets at each setting:

true τ integrated plug-in
0.25 98.6% 92.3%
0.5 97.5% 83.6%
1 95.2% 78.6%
2 94.2% 88.0%
4 94.8% 93.7%

The plug-in is worst in the middle. Not at the alike end, which is what “the plug-in fails where τ is hard to estimate” would predict. It recovers at both extremes, and for two different reasons.

At large τ it recovers because eight groups determine τ well and there is nothing left to integrate over. That is the expected reason.

At small τ it recovers for a reason that is almost an accident. When the groups really are nearly identical, the moment estimate collapses to zero — on more than half of datasets at τ = 0.25 — and the plug-in then reports complete pooling: eight identical estimates with a narrow interval around the grand mean. That answer happens to be nearly right, because the eight truths are all bunched near the grand mean too. The interval is far too narrow for the model that generated the data and still contains the truth most of the time, because the truth is where it is pointing.

So the plug-in’s coverage is a poor guide to whether the plug-in is behaving sensibly. At τ = 0.25 it covers 92.3% while reporting that eight distinguishable groups are one group.

And the integrated interval is conservative at small τ, at 98.6%. That is the flat prior on τ doing what a flat prior on a positive quantity does: keeping substantial mass at values of τ larger than the data supports, which widens every interval. A half-Cauchy with a well-chosen scale would be less conservative there, and would be an assumption. The trade is the one a prior on the spread measures.

The honest rule, which is narrower than the one the field started with: the integrated interval covers at least 94% at every setting measured and is conservative where the groups are alike; the plug-in is between five and sixteen points short across the whole middle of the range, and its recovery at the ends is not evidence that it is working.

The failures are correlated, and that changes what the number means

There is a second measurement hiding behind the coverage figure and it took a mistake to find.

The obvious standard error for a counted coverage is √(p(1−p)/n) with n the number of intervals. On sixteen thousand intervals that is 0.17 percentage points, and it is wrong, because the eight intervals from one dataset are not independent — they share the same τ̂ and the same estimated centre.

Computed over the right unit, which is the dataset, the standard error is 0.53 percentage points for the plug-in and 0.18 for the integrated interval. The first is three times the naive figure; the second is barely above it.

That gap is a measurement of something real. The integrated interval’s misses are scattered: one interval in twenty, roughly independently, roughly as a 95% interval should. The plug-in’s misses come in clumps. A study whose eight groups happened to look more alike than they are produces a τ̂ that is too small, which makes every one of its eight intervals too narrow at once. The failure is a property of the study, not of the interval, and nothing inside the study reveals it.

For a reader of one analysis, that is the more important of the two findings. A procedure that misses 5% of the time independently is a procedure whose errors average out across the eight statements in a report. A procedure that misses 21% of the time in clumps produces reports that are either mostly right or mostly wrong.

The clumping, counted directly

The standard errors above are an inference about correlation from the spread of a mean. The correlation can be counted head-on instead, and it is a better number to quote because it needs no reasoning at all: how often does every one of a study’s eight intervals cover?

If the eight failed independently, that share would be the per-interval rate raised to the eighth power. Over the same two thousand studies:

covers every group if the eight failed independently
integrated 68.9% 67.4%
plug-in 39.4% 14.8%

The integrated interval is within one and a half points of independence, which is what a procedure whose errors are unrelated to each other looks like. The plug-in is two and a half times above it.

That factor is the whole of the clumping claim, and reading it the friendly way round is instructive: the plug-in produces a fully correct set of eight statements far more often than its 79% per-interval rate would suggest. The price of that is the other tail. On the studies where τ̂ came out too small, the report is not wrong once — it is wrong repeatedly, in the same direction, with every group’s interval too narrow and every group’s estimate too close to the middle, and with no internal inconsistency to give it away.

A reader deciding whether to trust one analysis therefore faces a worse problem than the headline number suggests. Twenty-one per cent of intervals missing sounds like a tolerable error rate spread thinly. What it actually is, at eight groups, is a substantial minority of studies in which almost everything is wrong together.

The two-route check underneath

The whole apparatus rests on the posterior for τ, and the posterior for τ is computed with an algebraic shortcut: with a flat prior on μ, the population mean integrates out of the likelihood in closed form and leaves a one-dimensional function.

That shortcut has a second route to disagree with. The comparison integrates μ numerically on a grid of 1,201 points, states no identity, and shares only the normal density. The two agree to 6.5 × 10⁻¹⁰ on the density at every point of the τ grid, and to seven significant figures on its posterior mean. The residual is the second route’s own quadrature, which is the right direction: the route with more numerical integration in it is the less accurate one.

Without that check the coverage numbers would rest on an unverified piece of algebra doing the single most load-bearing job in the field.

What eight groups say about τ, when the truth is 1. The posterior density for the population spread after eight groups whose standard errors run from 0.5 to 2.1. The shaded band is the central 95% interval, from 1.02 to 4.44; the posterior median is 1.99 and the mean 2.18. The vertical mark at 1.07 is the moment estimate that empirical Bayes substitutes and then treats as known.
Fig. 5 The object both routes compute. Everything above is an average against this curve, which is why it is worth having two ways of producing it.

The group the model is wrong about

Coverage averaged over eight groups drawn from the population is the right measurement for the model as stated, and it is not the only question a reader has. The one that usually matters is about a particular group, and often about the group that looks unusual.

One group 4 population widths from the rest. Squared error for each group under partial pooling, with each group's own mean beside it. Seven of the eight are estimated better by pooling. The eighth, which was never from the population, is estimated 2.9 times worse — 6.5 against 2.3.
Fig. 6 The case pooling handles worst: seven groups from the population and one that is genuinely four population widths away from it. Its own mean is the only honest estimate available, and pooling drags it towards a middle it does not belong to — by 6.1 times the squared error, which when borrowing goes wrong counts in full.

Integrating over τ does not repair that. It cannot: the model says the groups come from a common population, the outlying group does not, and no amount of care about a nuisance parameter fixes a model that is wrong about the structure. What integration does do is make the failure visible in the interval rather than only in the estimate.

The mechanism is that an outlying group inflates the observed spread between the groups, which pushes the posterior for τ to the right, which widens every interval. The plug-in does the same thing to τ̂ — this is not an advantage of integration in kind — but the integrated interval also carries the uncertainty about how far right the posterior has been pushed, and one outlier in eight groups makes that uncertainty large. So the interval for the odd group comes out wide, which is the correct response to a group about which the population model is saying something the data disputes.

That is a real but modest defence. The honest summary is that a hierarchical interval is a statement made inside a model, and the counting in this essay establishes that the interval is right about what the model says. Whether the model is right about the groups is a different question, and it is the one when borrowing goes wrong is about.

One dataset on which the moment estimate collapsed. The eight groups' own means as hollow circles, the integrated posterior's interval as a bar, and the plug-in estimate as a filled mark. The moment estimate of τ came out exactly zero on this dataset, so every plug-in estimate is the same number, -0.08. The integrated posterior, on the same observations, still gives the groups different values, because its interval for τ runs from 0.02 to 2.06 and it never has to pick one.
Fig. 7 The worst case for the plug-in, and it is not rare. On a dataset where the moment estimate of τ came out exactly zero, every plug-in interval is centred at the same point and built on the assumption that the groups are identical. The integrated intervals on the same observations are the bars, and they are not the same object at all.

Where it sits beside the rest of the site

Three intervals on this site have now had their coverage counted rather than asserted, and the pattern across them is worth stating.

What a 95% credible interval covers, n = 30. Computed by summing over all 31 possible counts rather than by simulating them. Jeffreys' prior covers close to 95% across the range; a confident prior centred in the wrong place covers almost nothing where the truth is far from it.
Fig. 8 The proportion case, for comparison: exact coverage of a credible interval summed over all thirty-one possible outcomes rather than simulated, because for a proportion the sample space is finite. The oscillation is discreteness, not noise, and it is the reason coverage for a proportion is summed on this site rather than estimated.

For a proportion, coverage can be computed exactly by summing over the twenty-one possible outcomes. For a mean with unknown variance, the t interval covers exactly, by construction, and the check before the standard error is about when its assumptions fail. For a group in a hierarchy there is no exact answer available and the coverage has to be simulated, which is why every number here carries a standard error and why the standard error had to be computed over the right unit.

The direction of the results is consistent. An interval whose stated properties are checked usually turns out to have them; an interval whose stated properties are inherited from a conditional argument usually does not. The Wald interval inherits from the normal approximation. The plug-in hierarchical interval inherits from conditioning on an estimate. Both fail, and both fail worst where they are most needed.

What it costs to run

Nothing about this is expensive, which is worth saying because the reason empirical Bayes is the default is that it is fast.

The posterior for τ is one pass over a grid: at each point, a precision-weighted mean and a sum of eight log-densities. Two hundred and forty points is enough for four-digit accuracy on everything in this essay. The interval for one group is then a bisection on the mixture’s distribution function, which is another sixty passes over the same grid.

Counting the coverage — two thousand datasets, eight groups each, both intervals — takes under a minute, and most of that minute is building the endpoints rather than testing them. Whether an interval covers can be settled with a single evaluation of the mixture’s distribution function at the truth, since a central interval contains θ exactly when the posterior distribution function at θ lies between the two tail probabilities. That observation is worth fifty times the run time, and it is exact rather than an approximation to the interval.

The cost of the integrated interval, in other words, is roughly two hundred normal densities per group. The cost of the plug-in interval is one. That is the entire trade, and it buys sixteen percentage points of coverage and an error structure that does not fail in clumps.

What this field settles

Four essays and one claim, which can now be stated without hedging.

The shrinkage weight needs a population spread. The population spread is estimated from as many numbers as there are groups, which is usually few. Every consequence in this field follows from those two sentences and from the refusal to pretend otherwise:

  • The estimate of the spread is the most extreme value available — exactly zero — on 32.6% of eight-group datasets that have a real spread in them, and the chi-square tail that predicts that rate needs no simulation to state.
  • Substituting it and proceeding gives intervals that cover 78.8% of the time against a claimed 95%, and that fail together rather than separately: all eight cover on 39.4% of studies where independent failure would give 14.8%.
  • Integrating over it instead gives 95.2%, at 31% more width, with an error structure that is within one and a half points of independent.
  • The integral needs a prior on the spread, and on eight groups that prior never stops mattering — its influence falls from a factor of two to about 15% as the groups separate, and not to nothing.

The last of those is the honest limit of the repair and it is not a small one. What integration buys is not certainty about τ; it is a statement of uncertainty that is the right size. Where eight groups genuinely cannot settle how different eight groups are, the answer is a wide interval and an acknowledged dependence on what was assumed — which is what the data supports, and is more than the alternative offers.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

CoverageCredible intervalEmpirical BayesFlat priorFrequentist interpretationHierarchical modelPosteriorPriorShrinkage