What partial pooling does to one group, to the set, and to a ranking

The small groups one centre protects

In a league table of a hundred groups where larger groups do better on average, pooling every group towards one centre overstates the small groups by half a population width — 0.501 — because it pulls them towards an average they are not part of. The top ten barely notices. The bottom ten, the list that flags poor performers, holds 3.19 of the true bottom ten when pooled to one centre and 4.28 when the groups are ranked by their own noisy means. Pooling towards a centre that rises with size restores 4.35 and puts the small groups back at the bottom where 80.6% of the true bottom ten sit.

Worth reading first: The weight that decides.

A league table of a hundred ranked groups of 4 to 400 by their own means and by partially pooled posterior means, and found the pooled ranking far better: small groups, whose means are noisy, filled 62.1% of the raw top ten against their 36.3% share of the true one, and pooling cut that to 13.4%. The simulation made size unrelated to the true effect, so that any association between size and rank was manufactured by the estimator. It ended on the case where that is false. In the applications league tables are most used for — surgeons, hospitals, schools — larger units tend to do better, and pooling every unit towards one centre then misstates the smallest units systematically rather than merely noisily.

This essay measures how, and finds that the damage is at the end of the table nobody was looking at.

The bottom ten of a league table of a hundred when larger groups do better, size–effect slope 0.6Groups of 4 to 400, true effects rising with log size at a slope of 0.6 population widths. The true bottom ten is 80.6% small groups. Ranked by their own means the bottom ten holds 4.28 of the true ten; pooled to one centre, 3.19, with 42.6% small groups; pooled to a centre that rises with size, 4.35, with 83.0%.true bottom ten — small groups80.6%true bottom ten held10.00 of 10groups' own means — small groups84.8%true bottom ten held4.28 of 10pooled to one centre — small groups42.6%true bottom ten held3.19 of 10pooled to a centre by size — small groups83.0%true bottom ten held4.35 of 10600 tables of a hundred groups, 4 to 400 eachone centre shields the worst small groups
Fig. 1 The bottom ten of a league table of a hundred groups whose true effects rise with log size: the share of small groups in it and how many of the true bottom ten it holds, for the truth, the groups’ own means, pooling to one centre, and pooling to a centre that rises with size. The slider sets how strongly size predicts effect.

The table with no size relation

How many of a league table's top ten are small groups, ranked four ways. A hundred groups with sizes from 4 to 400, of which 36% have twenty units or fewer. Small groups make up 36.3% of the true top ten, 62.1% of the top ten by raw means, 13.4% by posterior means and 22.9% by the posterior chance of being in the top ten. The three rankings recover 4.43, 5.38 and 5.47 of the true top ten.
Fig. 2 The league table of the earlier essay, where size and effect are unrelated: the share of small groups in the top ten when ranked by the truth, by the groups’ own means, by posterior means and by each group’s probability of being in the top ten.

The earlier result is the baseline. With size unrelated to effect, small groups are 36.3% of the table and should be 36.3% of any list chosen by the truth; ranked by their own means they are 62.1% of the top ten, because a noisy mean is more often extreme; pooled, 13.4%, because pooling pulls them in harder than their true spread warrants. Every departure from 36.3% there was the estimator’s doing, and the estimator that departed least — pooling — was the one to recommend.

That recommendation carried an assumption it did not need to state, since it was true by construction: that knowing a group’s size tells nothing about its effect. A reader who takes the recommendation to a real table, where size tells a good deal, carries the assumption along. Under one centre the pooled estimates still depart least from the truth on average across all groups — their total error is far below the raw means’ — and still misplace the small groups as a class, which the average across all groups does not reveal. Eight groups, one population made the exchangeability assumption explicit at the start of the hierarchical argument; a size relation is what it looks like when it quietly fails.

Larger groups, better on average

The table has the same hundred groups as before, sized from 4 to 400 on a logarithmic spacing, each estimated with a standard error of five population widths divided by the square root of its size. What changes is the truth. Each group’s true effect is now

θi=βzi+τrεi,\theta_i = \beta z_i + \tau_r \varepsilon_i,

where ziz_i is the group’s log size standardised across the table and β2+τr2=1\beta^2 + \tau_r^2 = 1, so the true effects still have a spread of one. At β=0.6\beta = 0.6 size explains 36% of the variation in true performance — a strong volume–outcome relation, but not an exotic one — and the groups of twenty or fewer have an average true effect of −0.662, two thirds of a population width below the table’s centre.

Two pooled estimates are compared, both with their hyperparameters known so that the comparison is between models rather than between estimators of their parameters. One centre is the model of the essay on the league table: every group is pulled towards zero, the table’s average, by a weight set by its standard error and the population spread of one. A centre by size pulls each group towards βzi\beta z_i, the effect its size predicts, by a weight set by the smaller residual spread τr=0.8\tau_r = 0.8 — a regression at the level of groups, which is what the hierarchical model becomes when a group-level predictor is added.

What one centre does to a small group

A small group’s own mean is noisy and the pooled estimate leans heavily on the centre. Under one centre the centre is zero, and a small group’s true effect averages −0.662. So its pooled estimate is pulled up towards an average it is not part of.

One league table of a hundred, estimates against group size, pooled to one centre. Hollow points are the groups' own means, clipped at three population widths; filled points are the estimates pooled to zero; the line is that prediction. For the groups of twenty or fewer the pooled estimate sits on average 0.456 above the truth.
Fig. 3 One league table drawn against group size: the groups’ own means as rings, the estimates pooled to one centre as filled points, and the line that size predicts. The small groups on the left are pulled towards zero, above the line their size puts them on.

Over six hundred tables, the groups of twenty or fewer are overstated by 0.501 population widths on average under one centre, and by 0.001 under a centre by size. That is not noise, which pooling was designed to reduce; it is a bias that pooling created, of three quarters of the small groups’ true distance below the centre. Their root-mean-square error is 0.911 under one centre and 0.719 under a centre by size, a fifth less, and across all hundred groups the errors are 0.663 and 0.564.

The picture makes the mechanism plain. Under one centre the filled points for the small groups on the left sit in a band around zero, flattened by the pooling; under a centre by size they sit around the rising line, flattened towards it. Both models shrink the small groups hard, because their own means say little. What differs is what they shrink them to, and only one of the two targets is where the small groups actually are.

The same table, pooled to its line

One league table of a hundred, estimates against group size, pooled to a centre that rises with size. Hollow points are the groups' own means, clipped at three population widths; filled points are the estimates pooled to the line size predicts; the line is that prediction. For the groups of twenty or fewer the pooled estimate sits on average 0.070 below the truth.
Fig. 4 The same league table with every group pooled towards the effect its size predicts. The small groups on the left now sit around the rising line, below zero, where their true effects are.

Pooled towards the line, the same table looks different in exactly one place. The large groups on the right hardly move between the two pictures, since their own means are precise and neither model pulls them far. The small groups on the left drop from a band around zero to a band around the line, which at a size of four sits about a population width below the table’s centre. On this one table the groups of twenty or fewer are understated by 0.070 on average under the centre by size and overstated by 0.456 under one centre — the six-hundred-table averages, 0.001 and 0.501, seen in a single draw.

The move is the ordinary logic of the weight that decides: a group’s estimate is a weighted average of its own mean and a centre, with weight set by how much its own mean can be trusted. That logic does not say what the centre should be. It says only that a noisy group will end up close to it, and so the choice of centre becomes, for a noisy group, the choice of answer.

The top of the table does not notice

The obvious place to look for the consequence is the top ten, which is what a league table is published for, and there it is nearly invisible.

At β=0.6\beta = 0.6 the true top ten contains small groups only 3.8% of the time, because small groups are worse on average. Ranked by their own means the top ten is 36.4% small groups, the noise effect of the earlier essay; pooled to one centre, 2.2%; pooled to a centre by size, 0.1%. And the number of the true top ten each ranking holds is 4.91 by own means, 7.09 by one centre and 7.14 by a centre by size — a difference of a twentieth of a group. The top of the table is populated by large groups under every sensible model, and large groups are barely pooled, so the choice of centre hardly matters there.

That is why the problem is easy to miss. A league table checked by asking whether its top ten looks right will pass under the wrong model.

The bottom of the table does

The bottom ten is the list that triggers an inspection, a review, a withdrawal of a licence, and under a volume–outcome relation it is where the small groups belong: 80.6% of the true bottom ten are groups of twenty or fewer.

ranking small groups in the bottom ten true bottom ten held
the truth 80.6% 10.00
groups’ own means 84.8% 4.28
pooled to one centre 42.6% 3.19
pooled to a centre by size 83.0% 4.35

Pooling to one centre halves the small groups’ share of the bottom ten, from 80.6% in the truth to 42.6%, and identifies 3.19 of the true bottom ten — fewer than ranking by the groups’ own means, 4.28. The model that was adopted to stop small groups being flagged by noise stops them being flagged at all, including the ones that are genuinely worst. The pooled centre by size identifies 4.35 and has the small groups back at 83.0% of the bottom ten.

The single-centre model’s failure here is the mirror of its success in the earlier essay. There, small groups were over-represented at the top because their noisy means occasionally landed high, and pooling fixed that by pulling them towards the middle. Here the small groups genuinely belong at the bottom, and pulling them towards the middle removes them from the only list on which they should appear. A model that treats every small group as a noisy estimate of the table’s average will, when small groups are not average, protect them from the verdict their data support.

How the size relation sets the damage

The bottom ten of a league table of a hundred when larger groups do better, size–effect slope 0.3. Groups of 4 to 400, true effects rising with log size at a slope of 0.3 population widths. The true bottom ten is 58.0% small groups. Ranked by their own means the bottom ten holds 4.29 of the true ten; pooled to one centre, 4.25, with 25.0% small groups; pooled to a centre that rises with size, 4.58, with 40.6%.
Fig. 5 The same comparison with a weaker size relation, a slope of 0.3. The true bottom ten is 58.0% small groups; one centre still halves that, and a centre by size recovers most of it.

With no size relation, both pooled models are the same model and both hold 5.44 of the true bottom ten — the improvement over raw means that the earlier essay found at the top. At a slope of 0.3 the true bottom ten is 58.0% small groups; one centre gives 25.0% and holds 4.25 of the true ten, now below raw means at 4.29, and a centre by size gives 40.6% and 4.58. At a slope of 0.6 the gap is the one in the table. The stronger the relation between size and performance, the more a single-centre model misplaces the small groups, and the point at which it falls below the raw means comes early: at a slope of 0.3, where size explains only 9% of the variation in performance, it is already there.

The same trap wherever the groups differ in kind

The volume–outcome table is one instance of a general hazard. Partial pooling borrows strength on the assumption that the groups are exchangeable — that before the data are seen, nothing distinguishes one group’s expected effect from another’s. One population or two asked what happens when the groups come from two populations and the model assumes one; a size relation is the continuous version of the same failure, with every group belonging to a population indexed by its size.

It is also regression to the mean working in a direction nobody intended. Regression to the mean pulls an extreme observation back towards the mean of the population it came from, and is correct when the mean is the right one. Pooled to the table’s average, a small group is regressed towards a mean that is not its population’s; the regression is real and its target is wrong, and the resulting estimate is biased in exactly the way a naive observed mean would be biased in the other direction.

What the plug-in forgets found that pooling with an estimated population spread understates its own uncertainty; a misplaced centre adds a bias that no interval built around it will reflect, since every interval is centred on the pooled estimate. The error is in the location, and the reported uncertainty is about the spread.

What a centre by size requires

The repair is a group-level regression, and it needs two things the single-centre model does not.

It needs the size relation to be estimated. Here β\beta was known; in practice it is estimated from the same table, by regressing the groups’ means on their log size with weights — and its uncertainty then enters the pooled estimates, which is the question left open below. The residual spread τr\tau_r is then estimated from the scatter about that line, as the population spread was before. Estimating both costs a little of the regression model’s advantage and none of its direction.

It needs size to be a legitimate predictor. A centre by size says that a small group is expected to do worse, and a small group that does badly is then shrunk less towards acceptability than it would be under one centre. That is the right statistics if size is associated with performance and a questionable policy if the association is itself the thing under dispute — if small units are worse because they are under-resourced, a model that expects them to be worse can make the expectation look like a finding. Estimates that are too alike warned that posterior means are not a count of anything; a posterior mean pulled towards a size-predicted centre is also not a verdict on the group independent of its size. It should be reported as what it is: the group’s performance given its size.

What cannot be defended is the single-centre model applied where the size relation is known to exist. It is not neutral about size; it assumes the relation is zero, and when that assumption is false it is biased by the full size of the relation, in the direction that hides the groups most in need of attention.

What a table of units should print

The practical consequence for anyone publishing such a table is a short list. Estimate the size relation and print it, since whether it exists decides which pooled model is defensible. Rank the bottom of the table under the model that includes it, because that is where the single-centre model fails, and flag a small unit on its performance given its size. Say what the ranking conditions on: a unit’s place in a size-adjusted table answers “how does it do compared with units of its size?”, which is the question an inspector of a small unit needs, and not “how does it do compared with everyone?”, which is the question a patient choosing a unit asks. The single-centre table answers neither for the small units; the raw table answers the second with too much noise to be useful for any unit of four.

What one centre does to small groups, at the top and at the bottom

With true effects rising with log size at a slope of 0.6, pooling to one centre overstates groups of twenty or fewer by 0.501 population widths on average; pooling to a centre by size, by 0.001. Their root-mean-square error falls from 0.911 to 0.719.

The top ten holds 7.09 and 7.14 of the true top ten under the two models, a negligible difference. The bottom ten holds 3.19 under one centre, 4.28 by the groups’ own means and 4.35 under a centre by size, and its share of small groups is 42.6%, 84.8% and 83.0% against 80.6% in the truth.

Every figure is a count over six hundred seeded tables of a hundred groups, with the hyperparameters of both models set to their true values; the group sizes and standard errors are those of the earlier league table.

Not claimed: that volume–outcome relations are linear in log size, or that a slope of 0.6 is typical. The phenomenon depends on the relation’s strength and not its form, and a centre by size fitted with the wrong form would be biased by the part of the relation it missed. Not claimed either that the bottom ten should be published at all; a group from the population’s own tail found pooling costly for exactly the groups at the extremes, and the bottom of a table is an extreme.

Still open: the relation estimated from the table it corrects

Every number here uses the true slope. A real table estimates the slope from its own groups, and the groups most informative about the slope are the large ones, whose means are precise — so the estimate of how badly small groups perform is driven by how the large groups spread, extrapolated down to sizes where it is barely observed.

If the relation bends at small sizes — flattening, or steepening, below some volume — the extrapolated line misplaces the small groups’ centre, and the regression model inherits a bias of its own, smaller than the single centre’s and of unknown sign. How large that bias is for plausible bends, and whether a model with a flexible size relation estimated from a hundred groups keeps the regression model’s advantage at the bottom of the table, is the measurement that would make the repair safe to recommend without knowing the shape of the relation in advance.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Empirical BayesHierarchical modelLeague tablePartial poolingPosterior meanRankingRegression to the meanShrinkage