What partial pooling does to one group, to the set, and to a ranking

Estimates that are too alike

Posterior means give each group its least-error estimate, and as a set they are too alike: with each group's standard error equal to the population's spread, they spread 0.707 as widely as the truth. Beyond two population widths lie 2.28% of the true effects, 7.86% of the groups' own means, and 0.234% of the posterior means — a tenth of the truth. Rescaling the estimates to the right spread counts the tail exactly and costs 17% more squared error; summing each group's posterior chance of being beyond the line counts it without changing any estimate.

Worth reading first: The weight that decides.

A group from the population’s own tail judged partial pooling one group at a time and found the groups it serves badly. This essay judges it on a question that uses all the estimates at once. How many schools are more than two standard deviations above average? What does the distribution of hospital effects look like? How much do regions really differ?

Those questions are asked of the same table of estimates that answers “how good is this school”, and a table that is optimal for the second is systematically wrong for the first.

How the truth, the raw means, the posterior means and the constrained estimates spread, standard error 1Beyond two population widths above the centre lie 2.28% of the true values, 7.86% of the raw means, 0.234% of the posterior means, and 2.28% of the constrained estimates.00.2000.4000.600-4-2024estimate, in population widths from the centredensitythe truth: 2.28% beyondraw means: 7.86% beyondposterior means: 0.234% beyondconstrained: 2.28% beyondtwo widthsB = 0.50, closed formeach estimate is best alone; together they are too alike
Fig. 1 The spread of four sets of values across a population of groups, with each group’s standard error equal to the population’s spread: the true effects, the groups’ own means, the posterior means, and constrained estimates rescaled to the right spread. The vertical line is two population widths above the centre; the legend gives each set’s share beyond it. The slider sets each group’s standard error.

Why the posterior means are too alike

Each group’s posterior mean is its own mean shrunk towards the centre by the weight BB. Every group is shrunk, so the whole set is shrunk, and its spread can be written down. With the centre and the spread τ\tau known, the groups’ own means have variance τ2+se2\tau^2 + \mathrm{se}^2, and multiplying each by 1B1 - B gives

Var(posterior means)  =  (1B)2(τ2+se2)  =  (1B)τ2\operatorname{Var}(\text{posterior means}) \;=\; (1 - B)^2(\tau^2 + \mathrm{se}^2) \;=\; (1 - B)\,\tau^2

because 1B=τ2/(τ2+se2)1 - B = \tau^2/(\tau^2 + \mathrm{se}^2). The set of posterior means has less variance than the truth, by exactly the factor 1B1 - B. With the standard error equal to the spread that factor is one half, and the posterior means spread 0.707 times as widely as the effects they estimate.

The groups’ own means err the other way: their spread is τ2+se2\sqrt{\tau^2 + \mathrm{se}^2}, 1.414 times the truth at the same setting, because each carries its own noise on top of its real effect. So the three sets bracket the truth from both sides, and neither the unpooled nor the pooled table has the spread of the population it describes.

This is not an error in the posterior means. Each one is the best estimate of its own group’s effect given its data, and being the best estimate means being pulled towards the middle by exactly the amount the data cannot pin down. A set of best guesses is a set of cautious guesses, and cautious guesses are closer together than the things they guess at.

Counting the tail

The shortfall in spread becomes dramatic as soon as the estimates are used to count anything in a tail.

The share of groups beyond 2 population widths, counted by four sets of estimates. The truth has 2.28% of groups beyond 2 widths. The raw means count 7.86%, the posterior means 0.234% and the constrained estimates 2.28%; their mean squared errors are 1.000, 0.500 and 0.586 of a population variance.
Fig. 2 The share of groups beyond two population widths above the centre, counted from the truth and from three sets of estimates, with each set’s mean squared error beside it.

With the standard error equal to the spread, in closed form:

threshold the truth own means posterior means constrained estimates
1.5 widths 6.68% 14.44% 1.69% 6.68%
2 widths 2.28% 7.86% 0.234% 2.28%
2.5 widths 0.621% 3.85% 0.020% 0.621%

Beyond two widths the posterior means count a tenth of the true share; beyond two and a half, a thirtieth. The groups’ own means count three and a half times the true share at two widths and six times it at two and a half. Neither is a bad estimate of any individual group. Both are bad estimates of how many groups are extreme, in opposite directions, and the further into the tail the question reaches the worse both become.

The practical consequence falls on exactly the reports that tail counts feed: the number of hospitals flagged as outliers, the number of schools labelled exceptional, the share of regions above a target. Computed from the pooled estimates, those counts are far too small; computed from the raw means, far too large.

What the arrows do to the set

The standard picture of partial pooling makes the narrowing visible without naming it.

Eight groups, τ = 1 against a within-group spread of 3. Each row is a group. The hollow circle is the group's own mean, the filled one is the estimate after pooling, and the small mark is the truth the data was generated from. The group of 3 moves 75% of the way to the population mean of 0.10; the group of 40 moves 18%.
Fig. 3 Eight groups of unequal size: each group’s own mean, the pooled estimate at the head of its arrow, and the true effect the data were generated from. Every arrow points towards the centre.

Every arrow in the eight-group picture points inward. That is the whole mechanism of the gain, and it is also, read across the rows rather than along them, a statement that the set of pooled estimates occupies a narrower band than the set of own means — and, since the own means are wider than the truth and the arrows overshoot the truth’s width, narrower than the truth as well. The picture shows the true effects as small marks, and across the eight rows those marks are more spread out than the heads of the arrows.

With groups of unequal size the narrowing is uneven. The weight that decides showed that small groups move most of the way to the centre and large groups barely move, so a set of pooled estimates has its small groups bunched in the middle and its large groups spread out near their own means. A histogram of such a set is not merely too narrow; its tails are populated almost entirely by large groups, because only large groups are allowed to stay far out. That is the seed of the ranking problem a league table of a hundred groups measures, and it is already present in the histogram.

How it changes with the noise

The shortfall depends on the weight, which depends on how noisy each group’s own mean is.

How the truth, the raw means, the posterior means and the constrained estimates spread, standard error 2. Beyond two population widths above the centre lie 2.28% of the true values, 18.55% of the raw means, 0.000% of the posterior means, and 2.28% of the constrained estimates.
Fig. 4 The same four spreads when each group’s standard error is twice the population’s spread. The posterior means are packed into less than half the true width and the own means spread over more than twice it.
each group’s standard error posterior means’ spread own means beyond two widths posterior means beyond two widths
half a width 0.894 3.68% 1.267%
one width 0.707 7.86% 0.234%
two widths 0.447 18.55% 0.0004%

When the groups are noisy the posterior means have almost no tail at all — four in a million beyond two widths, against a true two in a hundred — while the raw means put nearly one group in five there. The noisier the data, the more a table of pooled estimates looks like a population of identical groups, and the more a table of raw estimates looks like a population of wildly different ones. The truth is in neither table, and a histogram of either is a histogram of the method.

Thresholds further out

The ratio between the posterior means’ count and the truth worsens steadily as the threshold moves out, because the posterior means’ distribution is a narrower normal and the ratio of two normal tails with different widths grows without limit.

The share of groups beyond 2.5 population widths, counted by four sets of estimates. The truth has 0.62% of groups beyond 2.5 widths. The raw means count 3.85%, the posterior means 0.020% and the constrained estimates 0.62%; their mean squared errors are 1.000, 0.500 and 0.586 of a population variance.
Fig. 5 The share of groups beyond two and a half population widths, counted four ways, with the standard error equal to the spread. The posterior means’ bar is too short to see.

Beyond two and a half widths the truth has 0.621% of groups and the posterior means have 0.020%: out of a thousand groups, six genuinely exceptional ones against a count of none. The own means count thirty-nine. A screening programme that flags groups beyond a threshold and uses pooled estimates to do it will, at this setting, flag almost nobody; one that uses raw means will flag six times too many, and most of the ones it flags will be small groups measured high that fall back on the next measurement.

The tail is also where a model’s shape assumption does the most work. Every number here assumes the group effects are normal. A level with no data in it is the general warning about reading far past where observations reach, and it applies with extra force to a count of extreme groups: if the real population has a heavier tail than the normal, every method’s count is too low, the posterior means’ lowest of all.

Estimates rescaled to the right spread

If the posterior means are too narrow by a known factor, they can be widened by it. Scaling each group’s posterior mean away from the centre by 1/1B1/\sqrt{1 - B} gives a set whose spread matches the population’s exactly. These are Louis’s constrained Bayes estimates, and with the centre and spread known they count every tail correctly: 2.28% beyond two widths, 0.621% beyond two and a half, the truth in every row of the table above.

The price is individual accuracy. The constrained estimate for each group is 1B\sqrt{1 - B} times its own mean, which is less shrunk than the posterior mean and therefore further from the least-error answer:

each group’s standard error posterior means’ error constrained estimates’ error own means’ error
half a width 0.200 0.211 0.250
one width 0.500 0.586 1.000
two widths 0.800 1.106 4.000

(Errors are mean squared errors in units of the population’s variance.) At one standard error per width the constrained estimates cost 17% more squared error than the posterior means and still less than three fifths of the unpooled error; at two widths they cost 38% more and a quarter of the unpooled. The rescaling buys a correct histogram with a modest amount of each group’s accuracy.

No single set of point estimates can be best at both. The posterior means minimise each group’s error; the constrained estimates get the distribution right; a third target — getting the ranks right — is met by neither. Estimates designed to compromise among all three exist, and they are compromises in the literal sense.

Counting without estimates

There is a better answer for the tail count specifically, and it does not require choosing between sets of point estimates at all.

The question “how many groups are beyond two widths” does not need any group’s estimate; it needs each group’s chance of being beyond two widths. Under the pooled model each group has a posterior distribution — normal, centred at its posterior mean, with variance (1B)se2(1 - B)\,\mathrm{se}^2 — and the chance it is beyond the line is a tail area of that distribution. Summing those chances over the groups gives the expected count.

Its average over datasets is the true count, exactly, by the law of iterated expectation: each group’s posterior chance averages to the prior chance, and the prior chances sum to the true share. It uses the posterior means, not a rescaled version of them, so each group’s reported estimate stays the least-error one; and it answers the collective question with a quantity built for collective questions.

With the spread estimated from fifty groups, as an analyst would have it, over two thousand simulated populations:

share beyond two widths
the truth 2.30%
own means 7.82%
posterior means 0.68%
constrained estimates 2.63%
sum of posterior chances 2.65%

Estimating the spread from fifty groups loosens everything slightly — the posterior means count 0.68% rather than 0.23%, because an estimated spread is often a little too large and shrinks less — but the ordering holds, and the two methods designed for the question land within four tenths of a point of the truth while the posterior means are short by a factor of three.

The spread is already in the model

The question “how much do the groups really differ” has a direct answer that needs no table of estimates at all: the model’s own estimate of the population spread, τ^\hat\tau. It is the quantity the pooling weights are built from, it is reported by any hierarchical analysis, and it is the spread of the true effects rather than of any set of estimates of them.

It has problems of its own, and they are already measured. At eight groups the moment estimate collapses to zero on a large share of datasets, and an interval for a group that plugs it in covers well short of its label. With enough groups it is a reasonable estimate of the right quantity, and with few it should be reported with its own uncertainty rather than as a number.

What should not happen is that the spread is estimated as τ^\hat\tau inside the model and then reported as the standard deviation of the pooled estimates outside it. The first is an estimate of τ\tau. The second is an estimate of 1Bτ\sqrt{1 - B}\,\tau, and at the noise levels where pooling is most useful it can be less than half the first.

Where the collective question is asked by accident

The failure is easy to fall into because the collective question is often asked of a table built for the individual one.

A report publishes pooled estimates for each hospital, correctly. A reader, or a later analysis, draws a histogram of the published numbers to show how much hospitals vary, and reports its spread as the variation between hospitals. The spread of the published numbers is 1B\sqrt{1 - B} times the real one. The hospitals look more alike than they are, by a factor that is largest exactly when each hospital’s data were thinnest — and the analysis that would have given the real figure, the estimate of τ\tau itself, was sitting in the model that produced the table.

The same happens in reverse with raw estimates. One population or two found that pooling can blur a population with two clusters into one bell without the output saying so; the histogram of raw means blurs in the other direction, spreading a single bell into something that looks like several. Both are the histogram of an estimator being read as the histogram of the population.

What a table of pooled estimates should carry

The measurements point to a short list of columns, each answering one of the questions a table of group estimates is actually asked.

The posterior mean, for the question about one group. It is the least-error estimate of that group’s effect and nothing in this essay argues against printing it.

The group’s posterior interval, built by integrating over the uncertainty in the spread rather than plugging in an estimate of it, for the question of how sure that estimate is.

The posterior chance of being beyond any threshold the report uses — above a target, below a safety limit, in the top tenth — for the questions that classify groups. Summed down the column it gives the expected count, which is the collective number the posterior means get wrong, and it does so without asking any reader to choose a different set of point estimates.

The estimated population spread with its own interval, reported once, for the question of how much the groups really differ.

A table with those four columns cannot be misread into a histogram of hospital effects, because the fourth column already is the answer and the first is visibly labelled as an estimate for one row. A table with only the first column invites every collective reading this essay measures, and every one of them comes out too alike.

What the formulas establish, and what they do not

Posterior means count less than a fifth of the true tail and the groups’ own means more than three times it, beyond two population widths with the standard error equal to the spread: 0.234%, 2.28% and 7.86%. The constrained estimates’ tail is the truth’s, exactly, with the spread known.

With the spread estimated from fifty groups, the constrained count is still far nearer the truth than the posterior mean’s — within four tenths of a point, against short by more than two thirds.

The posterior means have the least error, the constrained estimates more, the groups’ own means most, at every threshold and noise level drawn.

The sum of posterior chances averages to the true count by iterated expectation when the spread is known, and with it estimated the counted versions close on the truth as the number of groups grows: at two hundred groups over five hundred simulated populations, 2.40% against a true 2.21%, with the constrained estimates at 2.42% and the posterior means still at 0.36%. The small remaining overcount is the estimated spread, which is a little too large on average at every size measured.

What does not survive is a count of the groups beyond a threshold taken from the posterior means. Beyond two widths it is a tenth of the truth.

Not claimed: that any of this holds at eight groups. With the spread estimated from eight means, over eight thousand simulated sets, the true share beyond two widths is 2.26% and the posterior means happen to count 2.05% on average — not because they work, but because the estimated spread is so often far too large or collapses to zero that two large errors offset — while the constrained estimates count 3.76% and the sum of posterior chances 3.98%. At that size the estimate of the spread, not the choice of estimator, decides the count. The closed forms assume normal effects with a known centre; a population with a heavier tail than the model’s is undercounted further by every method here.

Still open: the ranking

The third collective use of a table of estimates is to order it. A league table ranks the groups, names the top ten, and implies that the tenth is ahead of the eleventh. Ranks are neither individual estimates nor a distribution, and neither the posterior means nor the constrained estimates are built to get them right.

With groups of unequal size the problem changes character, because the raw means put the smallest groups at both ends of the table and the posterior means pull them all to the middle. Which groups each method places in the top ten, how many of the true top ten any method recovers, and how wide the range of ranks a single group could plausibly hold is, are measured in a league table of a hundred groups.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Constrained bayesEmpirical BayesHierarchical modelMean squared errorPartial poolingPosterior meanShrinkageTail probability