What partial pooling does to one group, to the set, and to a ranking

The relation the table has to estimate

Pooling a league table's groups towards a centre that rises with size repaired the small groups' bias — when the size relation was known. Fitted from the table itself, a straight line in log size loses almost nothing: the small groups' error is 0.750 against 0.719 with the true relation. When the relation bends below the median size, the fitted line extrapolates the large groups' slope into the region where it is barely observed and misplaces the small groups by −0.173 if the relation flattens and +0.264 if it steepens — smaller than one centre's error, of either sign. A quadratic in log size removes nearly all of it, −0.025 and +0.014, and costs 0.025 in error when the relation was straight.

Worth reading first: The weight that decides.

The small groups one centre protects found that when larger groups in a league table do better on average, pooling every group towards one centre overstates the small ones by half a population width, because it pulls them towards an average they are not part of; and that pooling towards a centre that rises with size — a regression on log size — removes the bias and puts the small groups back where the true bottom ten is. Every number there used the true slope of that relation.

The essay ended on the step every real table has to take: the slope is estimated from the table it corrects. The groups most informative about it are the large ones, whose means are precise, so the estimated relation is driven by how the large groups spread and then extended down to sizes where it is barely observed. If the relation bends at small sizes, the extended line misplaces the small groups’ centre, and the regression model inherits a bias of its own.

How far each estimate of the small groups sits from their truth, when the size relation is flattening below the median sizeThe small groups' average true effect is −0.259. Their estimates are off on average by −0.012 (own means), +0.276 (one fitted centre), −0.173 (fitted line in log size), −0.025 (fitted quadratic in log size), −0.003 (the true relation).own means−0.012one fitted centre+0.276fitted line in log size−0.173fitted quadratic in log size−0.025the true relation−0.003← understatedoverstated →600 league tables of 100 groups, sizes 4 to 400average error for groups of 20 or fewer
Fig. 1 How far five estimates of the small groups — twenty patients or fewer — sit from their true effects on average, in league tables whose size relation flattens below the median size: the groups’ own means, pooling to one fitted centre, to a fitted line in log size, to a fitted quadratic, and to the true relation. The slider sets the shape of the relation.

The table, and three relations

A hundred groups, sized from four to four hundred on a logarithmic spread, each with a true effect and a measured mean whose noise shrinks with the square root of its size. The true effects are centred on a relation with size and scattered around it. Three relations are compared, all agreeing above the median size, where the large groups sit:

  • straight — the effect falls in proportion to log size all the way down, as in the essay on one centre;
  • flattening — the effect falls with log size above the median and is level below it, so small groups are no worse than middling ones;
  • steepening — the effect falls twice as fast below the median, so small groups are much worse.

The spread around each relation is set so that the true effects have the same total variance in all three. The analyst sees only the hundred means and their standard errors, and fits a hierarchical regression: a centre that is a polynomial in log size, a between-group spread estimated from the table, and each group pulled towards its fitted centre in proportion to its own noise.

When the relation is straight, estimating it costs little

With a straight relation, the small groups’ average true effect is −0.662. Pooled to the true relation, their estimates are off by −0.003 on average and by 0.719 in root-mean-square error. Pooled to a line fitted from the table, the bias is −0.010 and the error 0.750. The fitted line is almost as good as the true one, because a straight relation is exactly what a line extrapolates correctly: the large groups fix the slope, and the slope holds.

A quadratic in log size, which does not need the relation to be straight, costs a little where it is: the small groups’ error is 0.774, against the line’s 0.750. It spends a parameter on curvature that is not there, and the noise in that parameter reaches the small groups, whose fitted centre depends on it most because they sit at the end of the range. The small groups’ error rises by a little over two hundredths — the premium for flexibility where none was needed.

One centre is the poor choice regardless: pooled to a single fitted centre, with the spread also estimated from the table, the small groups are overstated by +0.712 — more than the half population width the essay on one centre found with the centre at the population average, because a fitted centre is a precision-weighted average, dominated by the large groups that do better, and so sits higher still above the small groups it pulls.

One league table of a hundred, estimates against group size, pooled to one centre. Hollow points are the groups' own means, clipped at three population widths; filled points are the estimates pooled to zero; the line is that prediction. For the groups of twenty or fewer the pooled estimate sits on average 0.456 above the truth.
Fig. 2 The essay on one centre’s picture of the problem: estimates against group size under the groups’ own means, one centre and a centre that rises with size, with the true relation straight and known. Everything below asks what changes when that relation has to be fitted.

When the relation bends, the line inherits a bias

The line’s success depends on the relation being a line, and nothing in the table’s large groups can say whether it continues below them.

One league table whose size relation flattens below the median size, with the line and the quadratic fitted to it. The fitted line has slope 0.255 and the between-group spread it estimates is 1.150; the quadratic's curvature is 0.236. At the smallest size the true relation is −0.259, the line predicts −0.448 and the quadratic 0.214.
Fig. 3 One league table whose size relation flattens below the median: each group’s own mean against its standardised log size, the true relation, and the line and quadratic fitted to the table. The small groups on the left are the noisiest and the ones the fit extrapolates to.

When the relation flattens below the median size, the small groups’ true effects sit at −0.259 on average, higher than a straight extrapolation from the large groups would put them. The fitted line, pulled by the large groups’ slope and by the small groups only weakly, runs below the truth at the small end, and pooling towards it drags the small groups’ estimates down: they are understated by 0.173 on average. That is a bias of the opposite sign to one centre’s, which overstates them by +0.276 here, and a little over half its size.

When the relation steepens below the median, the small groups’ true effects fall to −1.065 on average, below the line. The fitted line now runs above them at the small end, and pooling towards it overstates them by 0.264 — the same sign as one centre’s +1.253 and a fifth of its size.

How far each estimate of the small groups sits from their truth, when the size relation is steepening below the median size. The small groups' average true effect is −1.065. Their estimates are off on average by −0.012 (own means), +1.253 (one fitted centre), +0.264 (fitted line in log size), +0.014 (fitted quadratic in log size), −0.002 (the true relation).
Fig. 4 The same five estimates when the relation steepens below the median size. One centre’s overstatement is now more than a population width; the fitted line’s is a fifth of that, and the quadratic’s nearly nothing.

So the essay on one centre’s guess was right on both counts. The regression model’s bias is smaller than one centre’s, and its sign is not known in advance: it depends on which way the relation bends in a region the table barely observes.

A flexible centre sees the bend

Fitting a quadratic in log size lets the centre bend with the data. Where the relation flattens, the small groups’ bias under the quadratic is −0.025; where it steepens, +0.014 — each within a few hundredths of the true relation’s. The large groups still dominate the fit, but a curvature term lets the small groups’ own means bend the centre where they sit, and there are enough small groups in a table of a hundred to do it.

The cost is the one the straight case showed. In root-mean-square error the quadratic is slightly worse than the line when the relation is straight — 0.774 against 0.750 — about the same when it flattens, 0.866 against 0.871, and clearly better when it steepens, 0.490 against 0.539. None of the fitted centres reaches the true relation’s error, 0.817 and 0.398 in the two bent cases, because every fitted centre spends the table’s noise on estimating what the oracle was given.

The small groups' root-mean-square error under three shapes of the size relation, by how the centre was found. straight: line 0.750, quadratic 0.774, true relation 0.719; flattening below the median size: line 0.871, quadratic 0.866, true relation 0.817; steepening below the median size: line 0.539, quadratic 0.490, true relation 0.398.
Fig. 5 The small groups’ root-mean-square error under each shape of the size relation, for pooling to a fitted line, a fitted quadratic and the true relation.

The asymmetry is the argument for flexibility. Where the extra curvature is not needed it costs about three hundredths of error for the small groups; where it is needed, it removes a bias of 0.17 to 0.26 of a population width. A model that allows the relation to bend is a small insurance premium against a failure that would otherwise be invisible, since the fitted line looks equally reasonable on the table whether or not the truth bends beneath it.

What the bottom ten sees

A league table is read for its extremes, and the small groups one centre protects measured the damage in the bottom ten: one centre recovered fewer of the true bottom ten than the groups’ own noisy means did.

With the relation estimated, the pattern is the same and the differences between the fitted centres are small. With a straight relation, the fitted line’s bottom ten holds 4.19 of the true bottom ten on average, the quadratic’s 4.08 and the true relation’s 4.23, against 2.73 for one centre and 4.22 for the groups’ own means. With a steepening relation, where the true bottom ten is concentrated among the small groups, the fitted line holds 6.52, the quadratic 6.56 and the true relation 6.62, against 1.27 for one centre.

So the ranking, which is what a published table is used for, is protected by any centre that rises with size, and barely distinguishes a line from a quadratic. The bias the bend introduces is a bias in the small groups’ reported effects — the number a hospital or a school is given about itself — rather than in which groups appear at the bottom. That is not a small thing, but it is a different thing, and it is the number a flexible centre repairs.

Why the small end is where the extrapolation lives

The structure of a league table explains why the bend hides. Small groups are many in number and individually uninformative; large groups are few and individually precise. A hierarchical regression weights each group by the inverse of its noise plus the between-group spread, so the large groups set the fitted relation’s slope and the small groups mostly just sit under it. Where the relation is the same everywhere that is exactly right. Where it changes shape at small sizes, the model has fitted the large groups’ relation and applied it to groups it does not describe.

Borrowing towards a line measured what a group-level predictor buys when it describes the groups: it halves the spread left to borrow against. The borrowing is safe exactly as far as the predictor’s fitted form describes the groups it is applied to. When borrowing goes wrong found a single group from outside the population estimated six times worse than by its own mean; a bent size relation makes a whole region of the table into groups the fitted centre does not describe, and the error is correspondingly systematic rather than confined to one.

Where the fitted centre is least certain

The small groups sit at one end of the size range, and a fitted relation is least certain at the ends of the range it was fitted on. That is the leverage the points a slope rests on measured for a design: a fitted line’s variance at a point grows with the point’s squared distance from the centre of the design, and the smallest groups are as far from the centre as the table goes. Here the design is the spread of log sizes — even, from four to four hundred — so the leverage is moderate, and the fitted line is estimated well even at the small end when it is the right line. The bias the bend introduces is not variance at the end of the range; it is the right estimate of the wrong function.

A quadratic pays in the currency the straight case showed. It adds a term whose coefficient is estimated mostly from the ends of the range, so at the small end its fitted centre is noisier than the line’s. On any one table the quadratic can bend the wrong way — the single table drawn above bends upward at the smallest sizes, above the flat truth — and averaged over tables it bends the right way by the right amount. That is the trade a flexible model always makes, and a table of a hundred groups has enough small groups for the average to be what matters.

Two uses of one table

A league table serves two readers, and the fitted centre matters differently to each. A regulator reading the bottom ten wants the right groups flagged; a group reading its own row wants the right number. Estimates that are too alike found that posterior means, which give each group its least-error estimate, are too compressed as a set to count how many groups sit beyond a threshold; a group from the population’s own tail found that a group far from the centre is estimated worse by pooling than by its own mean. A bent size relation adds one more way the two readers can be served differently: the ranking is almost untouched by the choice of centre, and a small group’s own number is shifted by the fitted curve’s error at the small end.

The number a small group is told

The weight a small group’s own data gets is set by how noisy that data is against the spread around the centre, and for the smallest groups it is small. With the relation straight, a group of ten patients gets 79.6% of its reported estimate from the fitted centre and a fifth from its own mean; a group of four, 90.7% from the centre. Where the relation flattens and the spread around it is wider, the group of ten still takes 73.8% from the centre; where it steepens and the spread is narrow, 93.7%.

So for the groups a league table most often singles out, the reported number is mostly the fitted relation evaluated at their size. A bias of 0.173 or 0.264 in that relation at the small end is passed to them almost whole, and nothing in their own data can move them away from it: the data were down-weighted precisely because they are noisy. The only defence is in the choice of centre, made before the table is computed, and the evidence for it is in the small groups’ residuals taken together, which a table of a hundred has plenty of and a table of twenty barely has at all.

What a league table built this way should report

The fitted size relation, with its curvature allowed. A quadratic, or a spline with a few knots, in log size costs little where the relation is straight and removes the extrapolation bias where it is not. A table built on a straight line should say that it assumed one.

The small groups’ reliance on the fit. Each group’s estimate is a weighted average of its own mean and its fitted centre, and for the smallest groups the centre carries most of the weight. Reporting that weight beside the estimate tells a small group how much of its number is its own data and how much is the relation fitted to larger groups — the same disclosure a league table of a hundred argued for its ranking probabilities.

A check of the fitted relation at the small end. The small groups’ own means are noisy individually and informative together: their average residual from the fitted centre is an estimate of the extrapolation’s bias, and a residual consistently above or below zero at small sizes is the table’s own evidence that the relation bends.

What is measured here and what is not

With a straight size relation, pooling to a fitted line gives the small groups an error of 0.750 against 0.719 with the true relation, and one fitted centre overstates them by 0.712.

When the relation flattens below the median size the fitted line understates the small groups by 0.173, and when it steepens it overstates them by 0.264; a fitted quadratic leaves −0.025 and +0.014.

Every number is counted over six hundred simulated league tables of a hundred groups, sizes four to four hundred, with measurement noise five population widths per patient, true effects of unit total variance, and a hierarchical regression whose centre, slope and spread are estimated from each table by weighted least squares and a moment equation.

Not measured: relations that bend in the middle of the size range rather than at the small end, where the large groups would see the bend themselves; tables with few small groups, where a quadratic would have too little to bend with; and full Bayesian fits that carry the uncertainty in the centre into each group’s interval rather than plugging the fitted centre in, which what the plug-in forgets found matters for coverage.

Still open: the interval a small group is given

Every number here is a point estimate. A league table also reports intervals, and a small group’s interval under a fitted centre inherits two uncertainties the point estimate hides: the fitted centre’s own uncertainty at the small end, which is large because it is an extrapolation, and the between-group spread’s. An interval that plugs in the fitted centre as though it were known will be too narrow exactly for the small groups, where the centre carries most of the weight.

How much too narrow, whether the quadratic’s wider centre uncertainty makes its intervals honest where the line’s are not, and how often a small group’s interval excludes its own true effect under each model, are coverage measurements this site knows how to make and has not made for a fitted size relation.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Bias-varianceExtrapolationHierarchical modelLeague tableModel misspecificationPartial poolingRankingShrinkage