Partial pooling — where it appears
Named by 16 essays across 4 fields — each of them below, with the objects they name alongside it.
Eight groups, one population
Eight hospitals are neither one hospital nor eight unrelated problems. The two obvious answers cost 2.23 and 1.15 in squared error; the estimate between them costs 0.88, and the weight it uses is not a matter of taste.
The slope that borrows
Pooling a mean makes it look as though how much a group borrows depends on how much data it has. Pool a slope instead and the illusion breaks — ten groups with ten observations each can borrow anything from 28% to 91%, decided entirely by where those ten observations were placed.
What the plug-in forgets
The shrinkage weight needs a population spread, and the population spread has to be estimated from eight numbers. Empirical Bayes estimates it, substitutes it, and proceeds as though it were known — and the interval that comes out covers 79% rather than the 95% it claims.
A group from the population's own tail
Partial pooling halves the total squared error when a group's own standard error equals the spread between groups. Every group whose true effect sits more than 1.73 population widths from the centre — 8.33% of a perfectly normal population — does worse than it would have with its own mean, and its loss grows without bound. Among eight groups with the spread estimated, the most extreme is worse off in 61.6% of datasets. Capping the shift at one standard error keeps the total at 0.528 of the unpooled error and holds every group under twice it.
Pooling a proportion
A proportion cannot be shrunk on its own scale — an estimate would leave the interval, and how much information a count carries depends on where it sits. Move to log-odds and the approximation works, at the price of a group that saw nothing having no estimate at all until the correction supplies one.
The weight that decides
B = se²/(se² + τ²) is not a compromise between two answers. It is exactly the posterior mean's weight, it agrees with a numerical integration to ten digits, and an argument that mentions no population at all arrives at almost the same estimator.
Estimates that are too alike
Posterior means give each group its least-error estimate, and as a set they are too alike: with each group's standard error equal to the population's spread, they spread 0.707 as widely as the truth. Beyond two population widths lie 2.28% of the true effects, 7.86% of the groups' own means, and 0.234% of the posterior means — a tenth of the truth. Rescaling the estimates to the right spread counts the tail exactly and costs 17% more squared error; summing each group's posterior chance of being beyond the line counts it without changing any estimate.
Borrowing towards a line
A group shrunk towards the average of all groups is being compared with groups it has nothing in common with. Fit a group-level predictor and it is shrunk towards what the predictor says a group like it should be — which halves the spread left to borrow against and takes a quarter off the squared error.
The prior the data estimates
A hierarchical model needs a population spread, and it does not ask for one. It reads τ off the distance between the group means — biased six per cent low, exactly zero on 53% of datasets where the groups are identical — and the prior stops being a belief.
When the spread estimates to zero
The usual estimate of a population spread is a difference of two positive quantities, clamped at zero. On a third of eight-group datasets with a real spread in them the difference comes out negative, the estimate is exactly zero, and every group is pooled completely on data that said no such thing.
A league table of a hundred
A hundred groups with sizes from 4 to 400, and a top ten to publish. Ranked by their own means, small groups fill 62.1% of the top ten against their 36.3% share of the true top ten. Ranked by posterior means they fill 13.4%. The ranking built from each group's chance of being in the top ten recovers 5.47 of the true ten, the best of three and barely half; and the group ranked first could hold any rank from 1 to 31.
When borrowing goes wrong
Partial pooling wins on the total and can lose badly on one group. Placed six population widths out, the group that was never from the population is estimated six times worse than by its own mean — and nothing in the output says so.
The fewest groups that can borrow
At three groups the estimator that shrinks towards its own data's mean returns the group means untouched, on every dataset, because its constant is J − 3. At two it expands instead of shrinking. And the number of groups at which partial pooling starts to be worth doing is five, or two, or never — it depends on how far apart the groups are.
Where the borrowing goes
Pooling cuts the total squared error across eight groups by 56%. Two of the eight take 61% of that reduction, the four best-measured groups share 11% between them, and the largest group gets 1.5% of what the smallest does. The headline is a fact about the groups nobody was asking about.
One population, or two
Group effects from two clusters rather than one bell, with the same total spread. The analysis recovers the same population spread, uses the same weight for every group, and reports nothing unusual — while 46% of its estimates land in a region holding 6.6% of the truths.
What a two-unit study should report
The fixed-effect interval covers the mean of the two sites in hand 96.37% of the time and the population mean 54.77%. The random-effects interval covers the population mean 94.96% — exactly its level, from one degree of freedom — and is 11.6 times wider.
Named alongside it
The objects these essays reach for when they reach for this one.
Hierarchical modelShrinkageEmpirical BayesVariance componentsPriorRandom-effectsMean squared errorMethod of momentsPosteriorPosterior meanSample sizeExchangeability