What the plug-in forgets
Worth reading first: The weight that decides.
The shrinkage weight is one line and it has two quantities in it.
B = se² / (se² + τ²)
The first, se, is a group’s own standard error, and it is known: it comes from how many observations the group has and how noisy they are, and nothing about it is in doubt. The second is τ, the spread between the groups in the population they were drawn from, and nothing in the data is τ. It has to be inferred from the only evidence available, which is that the eight group means sit further apart than their own standard errors would explain.
The essay that derived the weight three ways derived all three with τ given. This one is about what happens when it is not.
The usual repair, and what it assumes
The standard answer is empirical Bayes, and it is three steps. Estimate τ from the spread between the observed group means. Substitute that estimate into B. Compute every group’s estimate and interval as though the substituted value were the truth.
The first step is honest and is the one the prior the data estimates is about: the between-group variance of the observed means is the real spread plus the average noise, so subtracting the average noise leaves an estimate of the real spread. The second step is arithmetic. The third step is the one that is not stated out loud, and it is a claim: that a number estimated from eight groups can be treated as a number that was known before the study began.
It cannot, and the size of the error is a count.
The interval, twice
Two intervals for the same group are drawn below. Both are 95% intervals for one group’s true value. Both use the same observations. One conditions on τ̂ = 1.07; the other integrates over everything the eight groups leave open about τ, which the figure above says is the range 1.02 to 4.44 and a long tail past it.
Two features of that picture matter more than the ratio.
The centres nearly agree. Integrating over τ does not move the estimate much, because the posterior mean of a shrunk quantity is roughly the shrunk quantity at the posterior mean of τ. What integrating changes is the spread.
The wide one is not merely a scaled copy of the narrow one. It is a mixture: one normal for every τ the posterior allows, averaged with that posterior as the weight. A mixture of normals with a common centre and different spreads is heavier in the tails than any of its components, and the tail is exactly where a 95% interval’s endpoints are.
Counted
Neither picture settles anything on its own. What settles it is running both procedures against two thousand datasets from a known truth and counting how often each interval contains the value it is an interval for.
79% against a claimed 95%. That is not a rounding problem or a small-sample caveat. One interval in five that says “95% confident” contains nothing of the sort, and the shortfall is systematic rather than noisy: the simulation’s own standard error on that figure is half a percentage point.
The gradient across the groups is the mechanism showing through. A group with a small standard error barely shrinks, so its interval barely depends on τ, so getting τ wrong barely hurts it. A group with a large standard error is mostly being told what to be by the population, and the population’s spread is precisely the thing that was guessed at.
It is worth being clear about what is and is not being claimed. The plug-in estimate is fine. Averaged over these two thousand datasets it is very slightly over-pooled — the moment estimate of τ runs low on average at eight groups, which pulls every group a little further towards the middle than it should go — but the effect on squared error is small, and partial pooling with an estimated τ still comfortably beats both of the obvious alternatives that eight groups, one population counted. The defect is not in the point estimate. It is in every statement of uncertainty attached to it.
The shape that no single number stands for
Look again at the first figure and notice what kind of object it is. It is not a bell. It rises steeply from near zero, peaks below 2, and then trails off to the right for a very long way — the central interval reaches 4.44, with half a per cent of the mass still beyond 6.
That shape is not an accident of the seed, and it is the reason a point estimate is the wrong summary even before the question of coverage comes up. Three perfectly reasonable summaries of that distribution disagree with each other: the posterior median is 1.99, the posterior mean is 2.18, and the value that maximises it is lower than either. The moment estimate is 1.07. An empirical-Bayes analysis picks one of those four numbers, and which one it picks changes every estimate on the page — not by much where the groups are well measured, and by a great deal where they are not.
A distribution that spans a factor of four with a heavy right tail is a distribution that has been asked a question eight numbers cannot answer. What the integrated interval does is decline to answer it: it uses every value of τ, weighted by how well that value explains the eight groups, and never commits to one.
Two things the width is made of
Working out where the missing width went is worth doing, because one of the two pieces is dropped by hand more often than the other.
The first piece is the one this essay is about: τ is uncertain, so the weight B is uncertain, so the estimate is uncertain by more than the conditional calculation says.
The second is easier to miss. Conditional on τ, a group’s posterior is centred at B·μ + (1 − B)·y, and μ is not known either — it is the precision-weighted average of the eight group means, which has its own standard error. So the variance of a group’s estimate has two terms:
Var(θ | τ) = (1 − B)·se² + B²·Var(μ | τ)
The second term is the uncertainty about where the middle is, inherited in proportion to how far the group was shrunk towards it. A group that moves 90% of the way to the middle takes 81% of the middle’s variance with it. Both intervals above carry that term; dropping it as well, which is the other common shortcut, makes the plug-in worse still.
The missing width, priced
The width ratio and the coverage shortfall are two readings of one defect, and they can be checked against each other with no further computation.
An interval that is a factor k too narrow, on a normal quantity, covers 2Φ(1.96/k) − 1. For the worst-measured group at a population spread of 1 the ratio is 1.47, which predicts 1.96/1.47 = 1.33 standard errors and therefore 81.8% coverage. That group is counted at 77.0%. At a population spread of 4 the ratio is 1.03, predicting 94.3%, and the plug-in is counted at 93.7%.
So the width accounts for most of the shortfall and not all of it. Where the two intervals nearly agree the prediction is within six-tenths of a point; where they differ most it is nearly five points optimistic. That residual is the shape rather than the scale, and it is the second feature named above: the integrated interval is a mixture of normals with a common centre and different spreads, which is heavier in the tails than any normal of the same width. Rescaling a normal cannot reproduce it, and the endpoints of a 95% interval are exactly where the difference lives.
The practical form of that is a warning about a tempting repair. Multiplying the plug-in interval by a fudge factor — 1.47, say, calibrated at one setting — would recover most of the coverage and would still be short, by a few points, in precisely the cases the fudge was calibrated on. The two things that are wrong with the plug-in interval are its width and its shape, and only one of them is a number.
What the second term is worth
The variance of a group’s estimate has a term for the group’s own residual noise and a term for uncertainty about where the middle is, and the second is described above as easy to miss. Its size can be written down.
For J groups measured to a common precision, the population mean’s variance conditional on τ is (se² + τ²)/J, and se² + τ² is se²/B. So the second term is B²·se²/(BJ) = B·se²/J, and the whole variance is
se²·[ (1 − B) + B/J ]
At eight groups and a group borrowing 90% of its estimate, that is se²(0.100 + 0.113) = 0.213 se² against the 0.100 se² the first term alone gives. The middle’s uncertainty more than doubles the variance of a heavily shrunk group — it is not a correction at the edge of the arithmetic, it is the larger of the two terms as soon as B passes 0.89.
The two ends of the expression are where the reading is clearest. A group that borrows little, at B = 0.5, picks up only 12.5% extra. A group that borrows everything, at B = 1, has a variance of se²/J rather than zero: complete pooling does not give a group a certain estimate, it gives it the overall mean’s uncertainty.
That last value is what makes the collapsed case in the caption above wrong twice rather than once. When the moment estimate of τ comes out at exactly zero, the analysis both pools completely — which the data did not license — and, if the second term has been dropped as well, reports an interval of essentially the group’s own residual width for a quantity it has no group-specific evidence about. The correct floor is not narrow; it is the width of the population mean, and at eight groups that is se/√8.
The failures come in clumps
There is a second finding underneath the coverage number and it changes what the number means.
The obvious way to state the precision of a counted coverage is the standard error of a proportion: √(p(1−p)/n) with n the number of intervals. On sixteen thousand intervals that gives 0.17 percentage points. It is wrong here, and it is wrong in the direction that matters.
The eight intervals from one dataset are not independent. They share the same τ̂ and the same μ̂, so a dataset on which τ̂ came out low produces eight intervals that are all too short together. Computing the standard error over the right unit — the dataset — gives 0.53 percentage points for the plug-in, three times the naive figure. For the integrated interval it gives 0.18, barely above the naive 0.17.
That difference is not a technicality about error bars. It says the two procedures fail differently. The integrated interval misses a scattered one time in twenty. The plug-in misses in clumps: a study whose groups happened to look more alike than they are produces a full set of estimates that are all over-pooled and a full set of intervals that are all too narrow, and nothing inside that study indicates it.
When it does not matter
An argument that a method is broken everywhere is usually an argument that has not been checked at the ends. The slider on the figures above runs the same comparison at other population spreads, and the two intervals converge.
So the plug-in is a good approximation exactly where it is not needed, and worst where partial pooling has the most to offer: a handful of groups that are somewhat but not obviously different. Run across the whole range, its coverage is 92.3%, 83.6%, 78.6%, 88.0% and 93.7% at population spreads of 0.25, 0.5, 1, 2 and 4 — worst in the middle, recovering at both ends, and the interval that integrates works through why the recovery at the alike end is not the good news it looks like.
That is the shape of the whole field. It is also why the failure has survived: an analyst who checks the method on a dataset with clearly distinct groups will find it works.
The same asymmetry runs through this site’s other fields and is worth naming as a pattern rather than a coincidence. The Wald interval for a proportion is fine in the middle and wrong at the edges. The normal approximation converges in the middle long before the tail. A method that is checked where it is comfortable is a method that has not been checked.
The route that had to exist
Everything above rests on one calculation — the posterior for τ — and that calculation uses an algebraic shortcut. With a flat prior on μ, the population mean can be integrated out of the likelihood in closed form, leaving a one-dimensional function of τ that a grid can handle in milliseconds. Without the shortcut the posterior is two-dimensional and the whole thing is slow enough to change what figures are possible.
A shortcut doing that much work needs something to disagree with, so it has one. The second route integrates μ numerically as well, on a grid of 1,201 points, and states no identity at all. The two share the normal density and nothing else.
They agree to 6.5 × 10⁻¹⁰ on the density at every point of the τ grid, and their posterior means for τ agree to seven significant figures. The residual disagreement is the second route’s own trapezoid, which is the expected direction: the route with more numerical integration in it is the less accurate one.
The grid that had to be found
There is a trap in the other calculation and it is the same trap the non-central t sprang during the previous phase: an integral taken over a range that is right in the middle of its cases and wrong at the ends.
Here it has a specific shape worth naming. The posterior for τ under a flat prior has a polynomial tail, not an exponential one. At large τ each of the J observations contributes a factor of 1/τ and integrating μ out returns one, so the density falls like τ^−(J−1) — at eight groups, τ^−7. That is integrable, and it decays slowly enough that a range set at “eight times the observed spread” leaves real mass outside it.
The obvious check is to look at the density at the top of the range, and it is the wrong check. On one of the datasets here the density at τ = 103 is still 3.2 × 10⁻⁶ of the peak, which sounds alarming and is not: the mass out there is already negligible. The quantity that matters is whether the answer moves. Quadrupling the range changes the posterior mean of τ by 0.1%; stopping the range at the posterior median — which keeps half the mass and looks superficially reasonable — changes it by 62%. The first number is what a good range looks like and the second is what the check has to be able to catch, which is why both are asserted.
The range is therefore not a multiplier. It doubles until the mass beyond its last decade is negligible, and the check compares the answer against the answer on a range four times as wide.
What is left over
Two things are unresolved and each is the next essay in this field.
The first is that everything above used a flat prior on τ, and said so without defending it. A flat prior is a choice, it is not the only proper one, and the whole objection to this repair is that it replaces an estimate with an assumption. That objection can be measured rather than argued about, and where it bites turns out to be exactly where the data has nothing to say.
The second is the moment estimate itself. It appeared in the caption of the first figure as 1.07, and in one of the captions above as exactly zero. An estimate of zero is not a small estimate — it is an instruction to pool completely, to give all eight groups the same number, on data that has said no such thing. That happens on nearly a third of eight-group datasets generated with a real spread in them, and it is worth its own essay.
What is settled is the claim this one was for. Partial pooling with an estimated τ is not partial pooling with a known τ, the difference is not small, and it is measurable to a fraction of a percentage point: 95.2% against 78.8%, on the same data, from the same model, differing only in whether one number was carried or dropped.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- What a prior is worth — both name coverage, flat prior, posterior, posterior mean, prior
- A league table of a hundred — both name hierarchical model, partial pooling, posterior mean, shrinkage
- One population, or two — both name empirical bayes, hierarchical model, partial pooling, shrinkage
- The fewest groups that can borrow — both name empirical bayes, hierarchical model, partial pooling, shrinkage
- The p-value a replication gets — both name flat prior, posterior, prior, shrinkage
- What a credible interval covers — both name coverage, flat prior, posterior, prior
Named objects
A flat tag is an object no other essay names yet.
CoverageEmpirical BayesFlat priorHierarchical modelMarginal likelihoodPartial poolingPlug in estimatePosteriorPosterior meanPriorShrinkage