What partial pooling does to one group, to the set, and to a ranking

A league table of a hundred

A hundred groups with sizes from 4 to 400, and a top ten to publish. Ranked by their own means, small groups fill 62.1% of the top ten against their 36.3% share of the true top ten. Ranked by posterior means they fill 13.4%. The ranking built from each group's chance of being in the top ten recovers 5.47 of the true ten, the best of three and barely half; and the group ranked first could hold any rank from 1 to 31.

Worth reading first: The weight that decides.

Estimates that are too alike found that a set of posterior means is narrower than the population it estimates, and that with groups of unequal size the narrowing is uneven: small groups are pulled to the middle and only large ones are allowed to stay far out. That was a statement about a histogram. A league table is a stronger use of the same numbers. It orders the groups, names the best ten, and implies that the tenth is better than the eleventh.

Every real league table ranks groups of unequal size — hospitals by case volume, schools by roll, surgeons by operations performed — so the unevenness is not a detail.

How many of a league table's top ten are small groups, ranked four ways. A hundred groups with sizes from 4 to 400, of which 36% have twenty units or fewer. Small groups make up 36.3% of the true top ten, 62.1% of the top ten by raw means, 13.4% by posterior means and 22.9% by the posterior chance of being in the top ten. The three rankings recover 4.43, 5.38 and 5.47 of the true top ten.
Fig. 1 The share of a league table’s top ten made up of small groups — twenty units or fewer — in truth and under three rankings, over six hundred simulated tables of a hundred groups. Beside each ranking, how many of the true top ten it recovers on average.

The table

A hundred groups, whose sizes run evenly on a logarithmic scale from 4 units to 400. Each group has a true effect drawn from a normal population with spread τ\tau; each unit’s outcome varies around its group’s effect with a spread five times τ\tau. So a group of four has a standard error of 2.5 population widths and a group of four hundred of a quarter. The population’s spread is taken as known, which is the most favourable case for pooling, and every number below is averaged over six hundred tables drawn this way.

Thirty-six of the hundred groups have twenty units or fewer. Because the true effects do not depend on size, small groups make up 36.3% of the true top ten on average — their fair share.

Three rankings are compared against the truth:

  • by each group’s own mean, which is what a table of raw results does;
  • by each group’s posterior mean, which is what a table of pooled estimates does;
  • by each group’s posterior chance of being in the top ten, found by drawing the whole table from the posterior two hundred times per dataset and counting how often each group lands in the top ten.

Who fills the top ten

ranking small groups’ share of the top ten true top ten recovered
the truth 36.3% 10
own means 62.1% 4.43
posterior means 13.4% 5.38
chance of top ten 22.9% 5.47

Ranked by their own means, small groups take nearly two thirds of the top ten. A small group’s mean is noisy, and a noisy mean is often far from the truth in either direction, so the top of a raw table fills with small groups that were measured high — and the bottom fills with small groups that were measured low. This is the league-table form of regression to the mean: next year’s table will demote most of this year’s small leaders, with nothing changed.

Ranked by posterior means, small groups almost vanish from the top ten. Pooling pulls a group of four most of the way to the centre whatever its mean was, so its posterior mean can almost never exceed that of a large group whose mean was merely good. The pooled table’s top ten is dominated by large groups for the opposite reason the raw table’s is dominated by small ones: only large groups are allowed to be extreme.

Both rankings misrepresent who is really at the top, in opposite directions, by roughly the same factor.

Group by group

The composition can be read one size band at a time. In truth a group of any size is in the top ten one time in ten.

How often a group of each size is placed in a league table's top ten, by four rankings. In truth a group of any size is in the top ten one time in ten. Ranked by raw means, the smallest groups (4–10 units) are placed there 20.0% of the time and the largest (101–400) 4.4%; ranked by posterior means, 2.1% and 15.6%; by the posterior chance of being in the top ten, 4.8% and 12.5%.
Fig. 2 For groups in four size bands, how often each ranking places one of them in the top ten. The rule marks the one in ten that every band reaches in truth.
size band in truth own means posterior means chance of top ten
4–10 units 10.1% 20.0% 2.1% 4.8%
11–40 units 9.9% 11.4% 8.0% 9.9%
41–100 units 9.8% 5.9% 12.8% 12.0%
101–400 units 10.1% 4.4% 15.6% 12.5%

A group of four to ten units is twice as likely as it should be to be ranked in the top ten by its own mean, and a fifth as likely by its posterior mean. A group of more than a hundred is under half as likely by its own mean and half as likely again as it should be by its posterior mean. No ranking gets every band to one in ten, and the one that comes closest — the chance of top-ten membership — still underrepresents the smallest groups by half.

The reason none can succeed is that the information is not there. A group of four units has a standard error two and a half times the whole population’s spread; its data cannot distinguish a group in the true top ten from an average one, and every honest ranking must place it nearer the middle than its truth would warrant, because it genuinely might be anywhere.

How much of the true top ten any ranking finds

The raw ranking recovers 4.43 of the true top ten. The pooled one recovers 5.38. Ranking by the chance of top-ten membership recovers 5.47, which is the best of the three — and the best available, since ordering by the posterior probability of membership maximises the expected number of true members in a published list of ten.

Even the best ranking is wrong about more than four of the ten. That is not a failure of the method. It is the information in the data: with this spread of group sizes and this much noise per unit, a table of a hundred can identify about half of its top ten, and every published top ten carries four or five groups that are there by luck and are missing four or five that belong.

The ordering from worst to best is itself informative. Pooling improves recovery by about one whole group over the raw ranking, and switching from posterior means to posterior probabilities adds a tenth of a group more. Most of the gain from modelling comes from pooling at all; the refinement for ranking specifically is small, and it matters most for the smallest groups, whose chance of reaching the top ten it more than doubles.

The rank a group could hold

A rank in a published table is reported as a single number. Its uncertainty can be computed the same way the top-ten probabilities were: draw the whole table from the posterior many times and record each group’s rank each time.

The twenty leading groups of one league table of a hundred by posterior mean, with the ranks each could holdGroups ordered by posterior mean. The bar spans the 5th to 95th percentile of each group's rank over 1,000 draws of the whole table from the posterior. The leading group, of 39 units, could rank anywhere from 1 to 31; the twentieth from 4 to 65.110255075100size 39, raw rank 7size 65, raw rank 11size 36, raw rank 8size 99, raw rank 15size 173, raw rank 19size 348, raw rank 20size 26, raw rank 9size 59, raw rank 16size 218, raw rank 24size 131, raw rank 26size 11, raw rank 3size 19, raw rank 10size 13, raw rank 5size 276, raw rank 29size 13, raw rank 6size 79, raw rank 27size 8, raw rank 4size 21, raw rank 18size 5, raw rank 2size 30, raw rank 25rank in the tableone table of a hundred, 1,000 posterior drawsa rank is an interval too
Fig. 3 One table of a hundred groups: the twenty with the highest posterior means, each with the range of ranks — 5th to 95th percentile — it takes across a thousand draws of the whole table from the posterior. Each row is labelled with the group’s size and its rank by its own mean.

In this one table the group with the highest posterior mean could plausibly be ranked anywhere from first to thirty-first, and the group twentieth by posterior mean anywhere from fourth to sixty-fifth. The intervals overlap almost completely: nearly any pair of the leading twenty could swap places without the data objecting.

The labels show the other half of the story. The leading group by posterior mean has 39 units and was seventh by its own mean; the groups above it in the raw table were small groups measured high, pulled back by pooling. Several of the twenty have hundreds of units and sit well down the raw table, because a moderately high mean on a large group is more convincing than a very high mean on a small one.

A table that printed these intervals instead of the ranks would say something true — that the top twenty are indistinguishable from each other and distinguishable from the bottom — and would not invite the reading that the group ranked third is better than the group ranked seventh.

The raw leaders, and their intervals

The same calculation applied to the groups a raw table would publish at the top shows why those groups are the least safe to name.

The twenty leading groups of one league table of a hundred by their own means, with the ranks each could hold. Groups ordered by their own means. The bar spans the 5th to 95th percentile of each group's rank over 1,000 draws of the whole table from the posterior. The leading group, of 4 units, could rank anywhere from 2 to 80; the twentieth from 5 to 22.
Fig. 4 The same table of a hundred, with its twenty leading groups chosen by their own means rather than by posterior means, each with the range of ranks it takes across a thousand draws from the posterior. Small groups are drawn in a different colour.

The first six places in the raw table go to groups of 4, 5, 11, 8, 13 and 13 units. The group ranked first by its own mean, a group of four, could plausibly hold any rank from 2 to 80, and its median rank across the draws is 28. The group ranked second, of five units, has the same median, 28, and an interval from 2 to 78. Each is more likely to be outside the top ten than in it, and both would be published at the very top.

The twentieth group in the raw table, by contrast, has an interval from 5 to 22: a larger group, measured more precisely, whose raw rank means something. A raw league table therefore puts its least informative groups in the most prominent places, and a reader has no way to tell from the ranks alone which of the leaders are real.

Next year’s table

The cost of a noisy ranking is felt when the table is published again. Measuring the same hundred groups a second time, with their true effects unchanged, and comparing each ranking’s top ten with its own top ten a year later:

ranking of this year’s top ten, still in next year’s
own means 3.55
posterior means 5.28

A raw league table replaces more than six of its top ten every year with nothing having changed. The groups that leave are the small groups that were measured high; the groups that arrive are other small groups measured high the second time. Readers of successive tables see a turnover that looks like a lively competition and is almost entirely measurement.

The pooled table is steadier, keeping more than half its top ten, and its stability is earned rather than imposed: it keeps the groups whose data were strong enough to hold them there, which are mostly the large ones. Its failure is the one measured above, that it rarely lets a genuinely excellent small group into the top ten at all.

Neither table’s turnover is evidence about the groups. A school that falls out of a raw top ten has, most likely, been measured again. That is the practical content of regression to the mean for anyone who publishes rankings annually, and it is a number that can be stated in advance for any table whose group sizes are known.

The same effect in eight rows

The eight-group picture that opens this collection’s account of pooling already contains the reordering.

Eight groups, τ = 1 against a within-group spread of 3. Each row is a group. The hollow circle is the group's own mean, the filled one is the estimate after pooling, and the small mark is the truth the data was generated from. The group of 3 moves 75% of the way to the population mean of 0.10; the group of 40 moves 18%.
Fig. 5 Eight groups of unequal size, smallest at the top: each group’s own mean, its pooled estimate, and its true effect. The small groups move most of the way to the centre and the large ones barely move, so the order of the pooled estimates is not the order of the own means.

In eight groups, one population the arrows are read one row at a time, as how far each estimate moves. Read down the column instead, they are a re-ranking. A small group whose own mean was the most extreme of the eight ends, after pooling, inside the range of the large groups, because its arrow is long; a large group with a moderately extreme mean keeps most of its distance. The weight that decides is the whole mechanism: the order of the pooled estimates is the order of the own means reweighted by each group’s precision, and any ranking of groups of unequal size is therefore a statement about sizes as much as about effects.

What pooling cannot fix

The measurements above are the most favourable case for any ranking. The population’s spread was known, the effects were exactly normal, the unit-level noise was the same in every group, and nothing about a group’s size was related to its effect. Each of those is routinely false.

When the spread is estimated, the pooling weights inherit its error, and a group from the population’s own tail found that the truly exceptional groups are exactly the ones an estimated weight serves worst. When size is related to effect — high-volume hospitals doing better, which is the usual finding — pooling towards a single centre drags small groups towards the large groups’ mean and misstates both. When the effects have a heavier tail than the normal, the posterior means are too cautious about the genuinely exceptional groups at the top, which is exactly the part of the table a league table is for.

So the honest conclusion is not that pooled rankings are right. It is that raw rankings are wrong in a known direction by a measurable amount, that pooled rankings are wrong in the other direction by a similar amount, and that no ranking of groups this unequal in size can identify more than about half of its top ten. A p-value on its own needs its second number beside it to be read; a rank needs its interval.

A picture that does not rank

Some institutions that publish comparisons of groups have stopped ranking altogether and plot instead: each group’s estimate against its size, with limits that fan out as size falls, drawn at the spread the estimates would have if every group had the same true effect. A group outside the funnel is flagged; a group inside it is not placed at all.

The funnel is the table’s information without the order. It shows directly what the composition tables above had to count — that small groups scatter widely and large ones do not — and it makes the flag a statement about one group against the noise its own size implies, rather than a position relative to ninety-nine others. It does not escape the problems measured here. The limits are drawn for a population with no real spread, so they flag real differences and chance together; and with a hundred groups, several fall outside a 95% funnel by chance alone, which is the figure of many groups reading of error bars in a different shape.

What it does avoid is the claim a rank makes. A funnel says a group is unusual or not; a rank says it is better than the group below it. The measurements here say the data can support the first claim for a few groups and the second for almost none.

What a league table should print

Three changes follow directly from the measurements, and none needs a new method.

Print each group’s chance of being in the top tenth, not its rank. It is the quantity that maximises the expected accuracy of a published list, it carries its own uncertainty, and a reader who sees 0.41 and 0.38 beside two schools does not conclude that one is better.

Print each group’s size beside its estimate. The composition tables show that size decides who reaches the top of either kind of table, and a reader can correct for that only if the size is visible.

Print rank intervals if ranks are printed at all. A leading group whose rank could be anything from first to thirty-first is not first in any sense a reader should act on, and the interval says so in the units the table is already using.

What the counts establish, and what they do not

Ranking by raw means overfills the top ten with small groups and ranking by posterior means underfills it, against the truth: 62.1% and 13.4% against 36.3%, over six hundred tables. In truth every size band reaches the top ten one time in ten, and the smallest groups reach it twice as often by raw means and less than half as often by posterior means.

Every rank interval is well formed, and even the leading group’s reaches past tenth place in the table drawn, by either ordering.

Measured again with nothing changed, a raw top ten keeps 3.55 of its members and a pooled top ten 5.28, over the same six hundred tables.

Not claimed: that the chance-of-top-ten ranking is always best by the margin shown. It is best in expectation by construction, and the margin over posterior means — 5.47 against 5.38 recovered — is small enough that on any single table either could recover more. The rank intervals come from one table, chosen by its seed and not selected, and a different table would have different leaders with intervals of similar width. Nothing here addresses a relationship between size and effect, which is the usual situation for hospitals and would change every composition figure.

The simulation made group size unrelated to the true effect, so that any association between size and rank was manufactured by the estimator. In the applications league tables are most used for, size and effect are related — higher-volume surgeons and hospitals have better outcomes on average — and pooling towards a single centre then misstates the smallest groups systematically, not merely noisily.

The repair is to pool each group towards a centre that depends on its size, which makes the model a regression at the group level. How much of the top-ten recovery that restores, whether it changes the composition of the pooled top ten back towards the truth, and how badly a single-centre model misranks the small groups when volume and outcome are genuinely linked, are measurements the same table can make and has not.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Hierarchical modelLeague tablePartial poolingPosterior meanRankingRegression to the meanSample sizeShrinkage