The spread, and its own uncertainty

What the plug-in forgets

The shrinkage weight needs a population spread, and the population spread has to be estimated from eight numbers. Empirical Bayes estimates it, substitutes it, and proceeds as though it were known — and the interval that comes out covers 79% rather than the 95% it claims.

Worth reading first: The weight that decides.

The shrinkage weight is one line and it has two quantities in it.

B = se² / (se² + τ²)

The first, se, is a group’s own standard error, and it is known: it comes from how many observations the group has and how noisy they are, and nothing about it is in doubt. The second is τ, the spread between the groups in the population they were drawn from, and nothing in the data is τ. It has to be inferred from the only evidence available, which is that the eight group means sit further apart than their own standard errors would explain.

The essay that derived the weight three ways derived all three with τ given. This one is about what happens when it is not.

The weight on the population, σ = 3. Each curve is one population spread τ. A group's estimate moves B = se²/(se² + τ²) of the way to the population mean, where se = σ/√n is what the group's own mean does not know. At τ = 1 a group of 9 observations sits halfway.
Fig. 1 The weight, as a reminder of what τ is doing in it. Every curve is B against the size of a group, and the family moves bodily as the population spread changes. Choosing a τ is choosing which curve the whole analysis sits on, and the four drawn here are not close together.
What eight groups say about τ, when the truth is 1The posterior density for the population spread after eight groups whose standard errors run from 0.5 to 2.1. The shaded band is the central 95% interval, from 1.02 to 4.44; the posterior median is 1.99 and the mean 2.18. The vertical mark at 1.07 is the moment estimate that empirical Bayes substitutes and then treats as known.00.2000.4000.600024τ, the spread between the eight groupsposterior densityτ̂ = 1.07truth 1p(τ | y) with a flat on τ prior, μ integrated out95% of the mass between 1.02 and 4.44
Fig. 2 Eight groups, generated with a real population spread of 1, and everything they say about that spread. The posterior median is 1.99 and its central 95% interval runs from 1.02 to 4.44 — a factor of more than four from end to end. The mark at 1.07 is the moment estimate that empirical Bayes substitutes for the whole of it.

The usual repair, and what it assumes

The standard answer is empirical Bayes, and it is three steps. Estimate τ from the spread between the observed group means. Substitute that estimate into B. Compute every group’s estimate and interval as though the substituted value were the truth.

The first step is honest and is the one the prior the data estimates is about: the between-group variance of the observed means is the real spread plus the average noise, so subtracting the average noise leaves an estimate of the real spread. The second step is arithmetic. The third step is the one that is not stated out loud, and it is a claim: that a number estimated from eight groups can be treated as a number that was known before the study began.

It cannot, and the size of the error is a count.

Eight groups, τ = 1 against a within-group spread of 3. Each row is a group. The hollow circle is the group's own mean, the filled one is the estimate after pooling, and the small mark is the truth the data was generated from. The group of 3 moves 75% of the way to the population mean of 0.10; the group of 40 moves 18%.
Fig. 3 What the substitution buys, and it is real: eight groups, each moved a different share of the way towards the middle according to how well it is measured. Nothing in this picture is wrong. What is missing from it is any indication that the number deciding every one of those distances was itself estimated from these same eight rows.

The interval, twice

Two intervals for the same group are drawn below. Both are 95% intervals for one group’s true value. Both use the same observations. One conditions on τ̂ = 1.07; the other integrates over everything the eight groups leave open about τ, which the figure above says is the range 1.02 to 4.44 and a long tail past it.

The same group, with τ known and with τ integrated over. Two posteriors for one group's value. The narrow one conditions on the moment estimate τ̂ = 1.07 and treats it as known; the wide one integrates over everything the eight groups leave open about τ. The intervals are 4.11 and 6.03 wide, a factor of 1.47, and their centres differ by only 0.257.
Fig. 4 The group with the largest standard error, drawn twice. Conditioning on the moment estimate gives an interval 4.11 wide; integrating over τ gives one 6.03 wide, a factor of 1.47. The two centres differ by 0.26 — a fifth of the difference in width. What is missing from the narrow one is not a better estimate. It is width.

Two features of that picture matter more than the ratio.

The centres nearly agree. Integrating over τ does not move the estimate much, because the posterior mean of a shrunk quantity is roughly the shrunk quantity at the posterior mean of τ. What integrating changes is the spread.

The wide one is not merely a scaled copy of the narrow one. It is a mixture: one normal for every τ the posterior allows, averaged with that posterior as the weight. A mixture of normals with a common centre and different spreads is heavier in the tails than any of its components, and the tail is exactly where a 95% interval’s endpoints are.

Counted

Neither picture settles anything on its own. What settles it is running both procedures against two thousand datasets from a known truth and counting how often each interval contains the value it is an interval for.

Counted coverage of two 95% intervals, over 2,000 datasets. Each point is one of the eight groups, at its own standard error. The integrated interval covers 95.2% overall against its stated 95%; the plug-in covers 78.8%, and its shortfall grows with the group's standard error — from 86.1% at se 0.5 to 77.0% at se 2.1. The mean widths are 3.48 and 2.66.
Fig. 5 Sixteen thousand intervals, one per group per dataset, sorted by the group’s own standard error. The integrated interval covers 95.2% of the time against its stated 95%. The plug-in covers 78.8%, and its shortfall grows with the standard error: 86.1% for the best-measured group and 77.0% for the one with the largest.

79% against a claimed 95%. That is not a rounding problem or a small-sample caveat. One interval in five that says “95% confident” contains nothing of the sort, and the shortfall is systematic rather than noisy: the simulation’s own standard error on that figure is half a percentage point.

The gradient across the groups is the mechanism showing through. A group with a small standard error barely shrinks, so its interval barely depends on τ, so getting τ wrong barely hurts it. A group with a large standard error is mostly being told what to be by the population, and the population’s spread is precisely the thing that was guessed at.

It is worth being clear about what is and is not being claimed. The plug-in estimate is fine. Averaged over these two thousand datasets it is very slightly over-pooled — the moment estimate of τ runs low on average at eight groups, which pulls every group a little further towards the middle than it should go — but the effect on squared error is small, and partial pooling with an estimated τ still comfortably beats both of the obvious alternatives that eight groups, one population counted. The defect is not in the point estimate. It is in every statement of uncertainty attached to it.

The shape that no single number stands for

Look again at the first figure and notice what kind of object it is. It is not a bell. It rises steeply from near zero, peaks below 2, and then trails off to the right for a very long way — the central interval reaches 4.44, with half a per cent of the mass still beyond 6.

That shape is not an accident of the seed, and it is the reason a point estimate is the wrong summary even before the question of coverage comes up. Three perfectly reasonable summaries of that distribution disagree with each other: the posterior median is 1.99, the posterior mean is 2.18, and the value that maximises it is lower than either. The moment estimate is 1.07. An empirical-Bayes analysis picks one of those four numbers, and which one it picks changes every estimate on the page — not by much where the groups are well measured, and by a great deal where they are not.

A distribution that spans a factor of four with a heavy right tail is a distribution that has been asked a question eight numbers cannot answer. What the integrated interval does is decline to answer it: it uses every value of τ, weighted by how well that value explains the eight groups, and never commits to one.

Two things the width is made of

Working out where the missing width went is worth doing, because one of the two pieces is dropped by hand more often than the other.

The first piece is the one this essay is about: τ is uncertain, so the weight B is uncertain, so the estimate is uncertain by more than the conditional calculation says.

The second is easier to miss. Conditional on τ, a group’s posterior is centred at B·μ + (1 − B)·y, and μ is not known either — it is the precision-weighted average of the eight group means, which has its own standard error. So the variance of a group’s estimate has two terms:

Var(θ | τ) = (1 − B)·se² + B²·Var(μ | τ)

The second term is the uncertainty about where the middle is, inherited in proportion to how far the group was shrunk towards it. A group that moves 90% of the way to the middle takes 81% of the middle’s variance with it. Both intervals above carry that term; dropping it as well, which is the other common shortcut, makes the plug-in worse still.

The missing width, priced

The width ratio and the coverage shortfall are two readings of one defect, and they can be checked against each other with no further computation.

An interval that is a factor k too narrow, on a normal quantity, covers 2Φ(1.96/k) − 1. For the worst-measured group at a population spread of 1 the ratio is 1.47, which predicts 1.96/1.47 = 1.33 standard errors and therefore 81.8% coverage. That group is counted at 77.0%. At a population spread of 4 the ratio is 1.03, predicting 94.3%, and the plug-in is counted at 93.7%.

So the width accounts for most of the shortfall and not all of it. Where the two intervals nearly agree the prediction is within six-tenths of a point; where they differ most it is nearly five points optimistic. That residual is the shape rather than the scale, and it is the second feature named above: the integrated interval is a mixture of normals with a common centre and different spreads, which is heavier in the tails than any normal of the same width. Rescaling a normal cannot reproduce it, and the endpoints of a 95% interval are exactly where the difference lives.

The practical form of that is a warning about a tempting repair. Multiplying the plug-in interval by a fudge factor — 1.47, say, calibrated at one setting — would recover most of the coverage and would still be short, by a few points, in precisely the cases the fudge was calibrated on. The two things that are wrong with the plug-in interval are its width and its shape, and only one of them is a number.

What the second term is worth

The variance of a group’s estimate has a term for the group’s own residual noise and a term for uncertainty about where the middle is, and the second is described above as easy to miss. Its size can be written down.

For J groups measured to a common precision, the population mean’s variance conditional on τ is (se² + τ²)/J, and se² + τ² is se²/B. So the second term is B²·se²/(BJ) = B·se²/J, and the whole variance is

se²·[ (1 − B) + B/J ]

At eight groups and a group borrowing 90% of its estimate, that is se²(0.100 + 0.113) = 0.213 se² against the 0.100 se² the first term alone gives. The middle’s uncertainty more than doubles the variance of a heavily shrunk group — it is not a correction at the edge of the arithmetic, it is the larger of the two terms as soon as B passes 0.89.

The two ends of the expression are where the reading is clearest. A group that borrows little, at B = 0.5, picks up only 12.5% extra. A group that borrows everything, at B = 1, has a variance of se²/J rather than zero: complete pooling does not give a group a certain estimate, it gives it the overall mean’s uncertainty.

That last value is what makes the collapsed case in the caption above wrong twice rather than once. When the moment estimate of τ comes out at exactly zero, the analysis both pools completely — which the data did not license — and, if the second term has been dropped as well, reports an interval of essentially the group’s own residual width for a quantity it has no group-specific evidence about. The correct floor is not narrow; it is the width of the population mean, and at eight groups that is se/√8.

The failures come in clumps

There is a second finding underneath the coverage number and it changes what the number means.

The obvious way to state the precision of a counted coverage is the standard error of a proportion: √(p(1−p)/n) with n the number of intervals. On sixteen thousand intervals that gives 0.17 percentage points. It is wrong here, and it is wrong in the direction that matters.

The eight intervals from one dataset are not independent. They share the same τ̂ and the same μ̂, so a dataset on which τ̂ came out low produces eight intervals that are all too short together. Computing the standard error over the right unit — the dataset — gives 0.53 percentage points for the plug-in, three times the naive figure. For the integrated interval it gives 0.18, barely above the naive 0.17.

That difference is not a technicality about error bars. It says the two procedures fail differently. The integrated interval misses a scattered one time in twenty. The plug-in misses in clumps: a study whose groups happened to look more alike than they are produces a full set of estimates that are all over-pooled and a full set of intervals that are all too narrow, and nothing inside that study indicates it.

When it does not matter

An argument that a method is broken everywhere is usually an argument that has not been checked at the ends. The slider on the figures above runs the same comparison at other population spreads, and the two intervals converge.

The same group, with τ known and with τ integrated over. Two posteriors for one group's value. The narrow one conditions on the moment estimate τ̂ = 4.25 and treats it as known; the wide one integrates over everything the eight groups leave open about τ. The intervals are 7.48 and 7.69 wide, a factor of 1.03, and their centres differ by only 0.098.
Fig. 6 The same group, on data generated with a population spread of 4. The intervals are now 7.48 and 7.69 wide — a ratio of 1.03. With the groups this far apart eight of them pin τ down well enough that there is almost nothing left to integrate over, and the plug-in is very nearly right.
The same group, with τ known and with τ integrated over. Two posteriors for one group's value. The narrow one conditions on the moment estimate τ̂ = 0.00 and treats it as known; the wide one integrates over everything the eight groups leave open about τ. The intervals are 1.26 and 4.92 wide, a factor of 3.90, and their centres differ by only 0.575.
Fig. 7 And the same group when the groups are nearly alike. The moment estimate of τ is exactly zero on this dataset, so the plug-in interval is 1.26 wide and centred where complete pooling puts it; the integrated interval is 4.92 wide, a factor of 3.9. The narrow interval is not an estimate of anything — it is what the arithmetic returns when it is told, wrongly, that the groups are identical.

So the plug-in is a good approximation exactly where it is not needed, and worst where partial pooling has the most to offer: a handful of groups that are somewhat but not obviously different. Run across the whole range, its coverage is 92.3%, 83.6%, 78.6%, 88.0% and 93.7% at population spreads of 0.25, 0.5, 1, 2 and 4 — worst in the middle, recovering at both ends, and the interval that integrates works through why the recovery at the alike end is not the good news it looks like.

That is the shape of the whole field. It is also why the failure has survived: an analyst who checks the method on a dataset with clearly distinct groups will find it works.

The same asymmetry runs through this site’s other fields and is worth naming as a pattern rather than a coincidence. The Wald interval for a proportion is fine in the middle and wrong at the edges. The normal approximation converges in the middle long before the tail. A method that is checked where it is comfortable is a method that has not been checked.

The route that had to exist

Everything above rests on one calculation — the posterior for τ — and that calculation uses an algebraic shortcut. With a flat prior on μ, the population mean can be integrated out of the likelihood in closed form, leaving a one-dimensional function of τ that a grid can handle in milliseconds. Without the shortcut the posterior is two-dimensional and the whole thing is slow enough to change what figures are possible.

A shortcut doing that much work needs something to disagree with, so it has one. The second route integrates μ numerically as well, on a grid of 1,201 points, and states no identity at all. The two share the normal density and nothing else.

They agree to 6.5 × 10⁻¹⁰ on the density at every point of the τ grid, and their posterior means for τ agree to seven significant figures. The residual disagreement is the second route’s own trapezoid, which is the expected direction: the route with more numerical integration in it is the less accurate one.

The grid that had to be found

There is a trap in the other calculation and it is the same trap the non-central t sprang during the previous phase: an integral taken over a range that is right in the middle of its cases and wrong at the ends.

Here it has a specific shape worth naming. The posterior for τ under a flat prior has a polynomial tail, not an exponential one. At large τ each of the J observations contributes a factor of 1/τ and integrating μ out returns one, so the density falls like τ^−(J−1) — at eight groups, τ^−7. That is integrable, and it decays slowly enough that a range set at “eight times the observed spread” leaves real mass outside it.

The obvious check is to look at the density at the top of the range, and it is the wrong check. On one of the datasets here the density at τ = 103 is still 3.2 × 10⁻⁶ of the peak, which sounds alarming and is not: the mass out there is already negligible. The quantity that matters is whether the answer moves. Quadrupling the range changes the posterior mean of τ by 0.1%; stopping the range at the posterior median — which keeps half the mass and looks superficially reasonable — changes it by 62%. The first number is what a good range looks like and the second is what the check has to be able to catch, which is why both are asserted.

The range is therefore not a multiplier. It doubles until the mass beyond its last decade is negligible, and the check compares the answer against the answer on a range four times as wide.

What is left over

Two things are unresolved and each is the next essay in this field.

The first is that everything above used a flat prior on τ, and said so without defending it. A flat prior is a choice, it is not the only proper one, and the whole objection to this repair is that it replaces an estimate with an assumption. That objection can be measured rather than argued about, and where it bites turns out to be exactly where the data has nothing to say.

The second is the moment estimate itself. It appeared in the caption of the first figure as 1.07, and in one of the captions above as exactly zero. An estimate of zero is not a small estimate — it is an instruction to pool completely, to give all eight groups the same number, on data that has said no such thing. That happens on nearly a third of eight-group datasets generated with a real spread in them, and it is worth its own essay.

What is settled is the claim this one was for. Partial pooling with an estimated τ is not partial pooling with a known τ, the difference is not small, and it is measurable to a fraction of a percentage point: 95.2% against 78.8%, on the same data, from the same model, differing only in whether one number was carried or dropped.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

CoverageEmpirical BayesFlat priorHierarchical modelMarginal likelihoodPartial poolingPlug in estimatePosteriorPosterior meanPriorShrinkage