Empirical Bayes — where it appears
Named by 10 essays across 3 fields — each of them below, with the objects they name alongside it.
What the plug-in forgets
The shrinkage weight needs a population spread, and the population spread has to be estimated from eight numbers. Empirical Bayes estimates it, substitutes it, and proceeds as though it were known — and the interval that comes out covers 79% rather than the 95% it claims.
A group from the population's own tail
Partial pooling halves the total squared error when a group's own standard error equals the spread between groups. Every group whose true effect sits more than 1.73 population widths from the centre — 8.33% of a perfectly normal population — does worse than it would have with its own mean, and its loss grows without bound. Among eight groups with the spread estimated, the most extreme is worse off in 61.6% of datasets. Capping the shift at one standard error keeps the total at 0.528 of the unpooled error and holds every group under twice it.
A prior on the spread
Integrating over the population spread means putting a prior on it, which sounds like the objection rather than the repair. The prior's effect is measurable, it is invisible where the groups are clearly different, and the reflex choice for a scale parameter turns out not to have a posterior at all.
Estimates that are too alike
Posterior means give each group its least-error estimate, and as a set they are too alike: with each group's standard error equal to the population's spread, they spread 0.707 as widely as the truth. Beyond two population widths lie 2.28% of the true effects, 7.86% of the groups' own means, and 0.234% of the posterior means — a tenth of the truth. Rescaling the estimates to the right spread counts the tail exactly and costs 17% more squared error; summing each group's posterior chance of being beyond the line counts it without changing any estimate.
The prior the data estimates
A hierarchical model needs a population spread, and it does not ask for one. It reads τ off the distance between the group means — biased six per cent low, exactly zero on 53% of datasets where the groups are identical — and the prior stops being a belief.
When the spread estimates to zero
The usual estimate of a population spread is a difference of two positive quantities, clamped at zero. On a third of eight-group datasets with a real spread in them the difference comes out negative, the estimate is exactly zero, and every group is pooled completely on data that said no such thing.
The interval that integrates
A credible interval for one group in a hierarchy has to average over every value the population spread might take. That averaging is what makes it cover — 95.2% against the plug-in's 78.8% — and it costs 31% more width, a heavier tail, and a mixture rather than a normal.
The fewest groups that can borrow
At three groups the estimator that shrinks towards its own data's mean returns the group means untouched, on every dataset, because its constant is J − 3. At two it expands instead of shrinking. And the number of groups at which partial pooling starts to be worth doing is five, or two, or never — it depends on how far apart the groups are.
Where the borrowing goes
Pooling cuts the total squared error across eight groups by 56%. Two of the eight take 61% of that reduction, the four best-measured groups share 11% between them, and the largest group gets 1.5% of what the smallest does. The headline is a fact about the groups nobody was asking about.
One population, or two
Group effects from two clusters rather than one bell, with the same total spread. The analysis recovers the same population spread, uses the same weight for every group, and reports nothing unusual — while 46% of its estimates land in a region holding 6.6% of the truths.
Named alongside it
The objects these essays reach for when they reach for this one.
Hierarchical modelPartial poolingShrinkageVariance componentsMean squared errorMethod of momentsPosteriorPriorFlat priorPosterior meanCoverageJames–Stein