Concept

Empirical Bayes — where it appears

Estimating a prior from the same data it is then used to shrink, which buys most of what a hierarchical model buys and understates the resulting uncertainty. What it leaves out is the uncertainty in the estimated prior itself, which matters most when there are few groups.

Named by 10 essays across 3 fields — each of them below, with the objects they name alongside it.

What eight groups say about τ, when the truth is 1. The posterior density for the population spread after eight groups whose standard errors run from 0.5 to 2.1. The shaded band is the central 95% interval, from 1.02 to 4.44; the posterior median is 1.99 and the mean 2.18. The vertical mark at 1.07 is the moment estimate that empirical Bayes substitutes and then treats as known.

What the plug-in forgets

The shrinkage weight needs a population spread, and the population spread has to be estimated from eight numbers. Empirical Bayes estimates it, substitutes it, and proceeds as though it were known — and the interval that comes out covers 79% rather than the 95% it claims.

fullbayes · Shrinkage
Which groups partial pooling serves, standard error 1 population width. Pooling's expected squared error for a group, divided by its own mean's, against how far the group truly sits from the centre. It is ×0.25 at the centre and crosses ×1 at 1.732 population widths, beyond which 8.33% of a normal population lies; capping the shift at one standard error holds every group under ×2.

A group from the population's own tail

Partial pooling halves the total squared error when a group's own standard error equals the spread between groups. Every group whose true effect sits more than 1.73 population widths from the centre — 8.33% of a perfectly normal population — does worse than it would have with its own mean, and its loss grows without bound. Among eight groups with the spread estimated, the most extreme is worse off in 61.6% of datasets. Capping the shift at one standard error keeps the total at 0.528 of the unpooled error and holds every group under twice it.

borrowed · Shrinkage
Three priors on the spread, at a true τ of 0.5. The posterior for τ under a flat prior (mean 1.66), a half-Cauchy of scale 1 (1.32) and one of scale 0.25 (1.15). The three answers differ by 30% of the widest. The prior does visible work when eight groups cannot separate a small spread from none, and almost none when they can.

A prior on the spread

Integrating over the population spread means putting a prior on it, which sounds like the objection rather than the repair. The prior's effect is measurable, it is invisible where the groups are clearly different, and the reflex choice for a scale parameter turns out not to have a posterior at all.

fullbayes · Prior
How the truth, the raw means, the posterior means and the constrained estimates spread, standard error 1. Beyond two population widths above the centre lie 2.28% of the true values, 7.86% of the raw means, 0.234% of the posterior means, and 2.28% of the constrained estimates.

Estimates that are too alike

Posterior means give each group its least-error estimate, and as a set they are too alike: with each group's standard error equal to the population's spread, they spread 0.707 as widely as the truth. Beyond two population widths lie 2.28% of the true effects, 7.86% of the groups' own means, and 0.234% of the posterior means — a tenth of the truth. Rescaling the estimates to the right spread counts the tail exactly and costs 17% more squared error; summing each group's posterior chance of being beyond the line counts it without changing any estimate.

borrowed · Shrinkage
τ̂ across 2,000 datasets of 12 groups, true τ = 1.5. The population spread is not supplied to a hierarchical model — it is estimated from how far apart the group means are, after subtracting the noise that would separate them anyway. It averages 1.41 here against a true 1.5, and comes out exactly zero on 5% of datasets.

The prior the data estimates

A hierarchical model needs a population spread, and it does not ask for one. It reads τ off the distance between the group means — biased six per cent low, exactly zero on 53% of datasets where the groups are identical — and the prior stops being a belief.

hierarchical · Prior
How often "there is no spread between the groups" is reported about data that has one. Every dataset here was generated with a real population spread of 1. The moment estimator is the difference between the observed spread and what noise alone would produce, clamped at zero, and the difference comes out negative often: at eight groups it reports exactly zero on 32.6% of datasets, which is an instruction to pool completely and give all eight groups the same estimate. The rate falls to 4.2% at 48 groups.

When the spread estimates to zero

The usual estimate of a population spread is a difference of two positive quantities, clamped at zero. On a third of eight-group datasets with a real spread in them the difference comes out negative, the estimate is exactly zero, and every group is pooled completely on data that said no such thing.

fullbayes · Pooling
Counted coverage of two 95% intervals, over 2,000 datasets. Each point is one of the eight groups, at its own standard error. The integrated interval covers 95.2% overall against its stated 95%; the plug-in covers 78.8%, and its shortfall grows with the group's standard error — from 86.1% at se 0.5 to 77.0% at se 2.1. The mean widths are 3.48 and 2.66.

The interval that integrates

A credible interval for one group in a hierarchy has to average over every value the population spread might take. That averaging is what makes it cover — 95.2% against the plug-in's 78.8% — and it costs 31% more width, a heavier tail, and a mixture rather than a normal.

fullbayes · Credible
What each estimator costs, against how many groups there are. Each point is 4,000 datasets, with the population spread estimated from the data rather than supplied. Partial pooling first beats BOTH of the estimators it sits between at 5 groups; below that, complete pooling — which estimates nothing at all — is the better answer. Its own cost falls from 1.751 at 2 groups to 0.795 at 40.

The fewest groups that can borrow

At three groups the estimator that shrinks towards its own data's mean returns the group means untouched, on every dataset, because its constant is J − 3. At two it expands instead of shrinking. And the number of groups at which partial pooling starts to be worth doing is five, or two, or never — it depends on how far apart the groups are.

hierarchical · Pooling
What each group gains from being pooled, τ = 1. Eight groups whose sizes span a factor of 13.3. The smallest gains 2.076 of squared error, which is 36.8% of the total reduction; the largest gains 0.030, which is 0.5%. 2 of the eight account for half of everything pooling buys.

Where the borrowing goes

Pooling cuts the total squared error across eight groups by 56%. Two of the eight take 61% of that reduction, the four best-measured groups share 11% between them, and the largest group gets 1.5% of what the smallest does. The headline is a fact about the groups nobody was asking about.

hierarchical · Pooling
Twelve groups from two clusters, τ = 1. Every group's truth is in one of two clusters, and the population's total spread is exactly τ = 1, so the analysis recovers τ̂ = 1.515 and every shrinkage weight is what it would be for a single normal population. The estimates are pulled towards the grand mean, which is the middle of the gap — a place 0 of the 12 truths are and 6 of the estimates end up.

One population, or two

Group effects from two clusters rather than one bell, with the same total spread. The analysis recovers the same population spread, uses the same weight for every group, and reports nothing unusual — while 46% of its estimates land in a region holding 6.6% of the truths.

hierarchical · Pooling

Named alongside it

The objects these essays reach for when they reach for this one.

Hierarchical modelPartial poolingShrinkageVariance componentsMean squared errorMethod of momentsPosteriorPriorFlat priorPosterior meanCoverageJames–Stein

All concepts