What partial pooling does to one group, to the set, and to a ranking

The interval a small group is given

A league table's 95% interval for a small group, pooled towards a size relation fitted from the table, covers 92.8% of the time when the relation is straight and 83.1% when it steepens below the median — and 77.4% for the groups of eight patients or fewer. Adding the fitted centre's own uncertainty, weighted by how much the group borrows from it, brings the line to 88.4% and a fitted quadratic to 93.7%, and the quadratic's interval covers 94.6% of the smallest groups at every shape tried. One centre, plugged in, covers the smallest groups 35.2% of the time when the relation steepens, and no added variance repairs a centre in the wrong place.

Worth reading first: The weight that decides.

The relation the table has to estimate fitted a league table’s size relation from the table itself and pooled every group towards it. A straight line in log size lost almost nothing when the relation was straight; when it bent below the median size, the line carried the large groups’ slope down into the region where it was barely observed and misplaced the small groups by −0.173 or +0.264; a quadratic removed nearly all of it. Every number there was a point estimate.

A league table also publishes intervals, and the essay ended on what they inherit. A small group’s estimate is mostly its fitted centre, because its own mean is noisy, so its interval should carry the centre’s uncertainty, and the centre is least certain exactly at the small end, where it is an extrapolation. An interval that treats the fitted centre as known will be too narrow for the small groups, the essay predicted — and the question was how much, and whether the quadratic’s wider uncertainty at the ends makes its intervals honest where the line’s are not.

How often seven 95% intervals cover the small groups of a league table whose size relation is steepening below the median sizeGroups of twenty patients or fewer, 600 tables of a hundred groups. Coverage: one centre, plugged in 58.0% (average width 2.81); one centre, its uncertainty added 58.5% (average width 2.83); a fitted line, plugged in 83.1% (average width 1.58); a fitted line, its uncertainty added 88.4% (average width 1.77); a fitted quadratic, plugged in 85.5% (average width 1.53); a fitted quadratic, its uncertainty added 93.7% (average width 1.92); the true relation and spread 94.8% (average width 1.55).one centre, plugged in58.0%width 2.81one centre, its uncertainty added58.5%width 2.83a fitted line, plugged in83.1%width 1.58a fitted line, its uncertainty added88.4%width 1.77a fitted quadratic, plugged in85.5%width 1.53a fitted quadratic, its uncertainty added93.7%width 1.92the true relation and spread94.8%width 1.55600 tables, groups of 20 or fewer; nominal 95%an interval borrows the centre's error too
Fig. 1 How often seven 95% intervals cover the true effects of the small groups — twenty patients or fewer — in league tables whose size relation steepens below the median size, with each interval’s average width. The slider sets the shape of the relation.

Three intervals for one group

The table is the one the earlier essays built: a hundred groups from four to four hundred patients on a logarithmic spread, true effects centred on a relation with log size and scattered around it, each group’s mean measured with noise that shrinks with its size. The analyst fits a hierarchical regression — a centre that is one mean, a line or a quadratic in log size, and a between-group spread estimated from the table — and pulls each group towards its fitted centre by the weight Bi=sei2/(sei2+τ^2)B_i = se_i^2/(se_i^2 + \hat\tau^2).

The plug-in interval is the posterior mean plus or minus 1.96 times Biτ^2\sqrt{B_i\hat\tau^2}: the posterior spread if the fitted centre and spread were the truth. It is the interval a table usually prints.

The centre-aware interval adds the fitted centre’s own variance, multiplied by Bi2B_i^2 because the group’s estimate takes a share BiB_i of its centre. The centre’s variance is the regression’s prediction variance at that group’s size, xi⊤(X⊤WX)−1xix_i^{\top}(X^{\top}WX)^{-1}x_i, which is smallest in the middle of the size range and grows at the ends, fastest for a quadratic.

The oracle uses the true relation and the true spread, and shows what a 95% interval can be when nothing has to be estimated.

Where the centre’s uncertainty lives

Before any coverage is counted, the construction says where the extra term matters. A group’s centre-aware variance is its plug-in variance Biτ^2B_i\hat\tau^2 plus Bi2B_i^2 times the centre’s prediction variance at its size, and both factors grow at the small end: the small groups borrow the most, so BiB_i is near one, and they sit at the edge of the size range, where a fitted polynomial is least certain.

How much of each group's interval is the fitted centre's uncertainty, by the group's size. One table with a straight size relation. The centre's share of the centre-aware interval's variance for the group of four: one centre 1.3%, a line 10.4%, a quadratic 26.0%; for the group of four hundred, 0.1%, 0.4% and 0.7%.
Fig. 2 In one table with a straight size relation, the share of each group’s centre-aware interval variance that comes from the fitted centre, against the group’s size, for one centre, a fitted line and a fitted quadratic.

In the table drawn above, the fitted centre accounts for a small share of a large group’s interval under every model, because a group of four hundred barely borrows from its centre. For the group of four the share is 1.3% under one centre, whose single mean is estimated from all hundred groups; 10.4% under a line; and 26.0% under a quadratic, whose curvature is estimated mostly from groups far from the small end and extrapolated to it. That is the plug-in interval’s omission, located: almost nothing for most of the table, and a quarter of the variance for the groups a reader is most likely to look up.

When the relation is straight

With a straight relation, the fitted line’s plug-in interval covers the small groups’ true effects in 92.8% of cases, and the centre-aware interval in 93.9%. The quadratic’s plug-in interval covers 92.1% and its centre-aware one 94.2%, against the oracle’s 94.9%. For the smallest group in the table, of four patients, the quadratic’s plug-in interval covers 90.0% and its centre-aware interval 95.3%.

So the prediction holds and the size is modest when the relation is what the model assumes: a plug-in interval loses two or three points of coverage for the small groups, most of it for the smallest, and adding the centre’s variance restores nearly all of it. The quadratic’s plug-in interval does a little worse than the line’s, because its centre is more uncertain at the ends and the plug-in ignores exactly that; with the uncertainty added the quadratic does a little better, because it now carries its own honesty.

One centre is a different case. Pooled towards one fitted mean, the small groups’ plug-in interval covers 86.9%, and for the group of four, 76.5%. Adding the centre’s variance changes that by a few tenths, to 87.2%, because the single mean is precisely estimated — the problem is not its uncertainty but its position. The small groups one centre protects found that one centre overstates the small groups by half a population width; an interval around an estimate that is half a width off is centred in the wrong place, and no widening short of doubling it will cover.

When the relation steepens

The failure the earlier essay found for point estimates reaches the intervals in full when the relation steepens below the median.

The fitted line’s plug-in interval covers the small groups 83.1% of the time and the groups of eight or fewer 77.4%. Its centre-aware interval covers 88.4% and 84.6% — better, since the line’s extrapolated centre is uncertain as well as biased, but still short, because the line’s centre is biased at the small end by +0.264 and no symmetric widening removes a bias. The quadratic’s plug-in interval covers 85.5% of small groups, and its centre-aware interval 93.7% — and 94.6% of the smallest groups. The quadratic can follow the steepening, so its centre is nearly unbiased at the small end, and its prediction variance there is large enough to cover what it cannot follow.

How often each group's 95% interval covers its true effect, by the group's size, relation steepening below the median size. Over 600 tables. The smallest group, of four patients: fitted line plugged in 71.2%, with its uncertainty 80.8%, fitted quadratic with its uncertainty 94.8%, true relation 95.0%. The largest, of four hundred: 93.3%, 94.3%, 93.5% and 95.7%.
Fig. 3 How often each group’s 95% interval covers its true effect, against the group’s size, when the relation steepens below the median: the fitted line plugged in and with its uncertainty, the fitted quadratic with its uncertainty, and the true relation.

The line’s remaining shortfall has a simple size. Its centre-aware interval for the small groups averages 1.77 wide, a standard error of about 0.45, and its estimates of those groups are off by +0.264 on average, which is 0.58 of that standard error. An interval of the right width centred 0.58 standard errors from the truth covers 91.0% of the time, not 95%; the counted 88.4% is a little lower because the bias is larger for the smallest groups than for the average small one. No variance term can close that gap, because it is not variance: the interval is the right width around the wrong centre, and only a centre that can follow the bend puts it back.

The coverage by size shows where the failure lives. For groups above the median size every interval covers close to its nominal rate: the large groups’ own means carry them, and the centre hardly matters. Below it, the plug-in line falls steadily, to 71.2% for the group of four; its centre-aware version falls less, to 80.8%; the quadratic with its uncertainty stays near 95% all the way down, at 94.8% for the group of four. The gap between the line’s two intervals is the centre’s variance, and the gap between the line’s aware interval and the quadratic’s is the line’s bias.

One centre is at its worst here. Its plug-in interval covers the small groups 58.0% of the time, the groups of eight or fewer 35.2%, and the group of four 16.2%. A table that pools its smallest groups towards the overall mean when the smallest groups are genuinely worse publishes intervals that miss their truth two times in three and point confidently at a centre they are not part of.

When the relation flattens

The flattening relation is the mild case for intervals, for a reason worth noticing. When the relation flattens below the median, the small groups are no worse than middling ones, and the between-group spread around any fitted centre is larger than with a straight relation — the three shapes were set to have the same total variance, and a flattening relation leaves more of it unexplained. The intervals are correspondingly wider, 3.18 for the line’s plug-in against 2.78 with a straight relation, and wide intervals tolerate a misplaced centre.

So the line’s plug-in interval covers the small groups 92.9% of the time despite a bias of −0.173, its centre-aware interval 93.7%, and the quadratic’s 94.3%. Even one centre covers 93.9%, because under a flattening relation the overall mean is not far from where the small groups sit. The same bias that moved the point estimates by a sixth of a standard deviation costs the intervals a point or two, and the order of the methods is unchanged.

What honesty costs

The smallest groups' intervals in one league table whose size relation steepens below the median. The 17 groups of eight patients or fewer. A fitted line with its centre treated as known misses 4 of their true effects; a fitted quadratic with its centre's uncertainty added misses 0.
Fig. 4 One league table whose relation steepens below the median: the seventeen groups of eight patients or fewer, each with its 95% interval from a fitted line plugged in (upper) and from a fitted quadratic with its uncertainty added (lower), and its true effect.

In the table drawn above, the line’s plug-in intervals miss four of the seventeen smallest groups’ true effects, each by sitting above it, and the quadratic’s centre-aware intervals miss none. They do it by being wider. Averaged over the small groups under a steepening relation, the line’s plug-in interval is 1.58 wide, its aware interval 1.77, the quadratic’s plug-in 1.53 and its aware interval 1.92, against the oracle’s 1.55. The honest interval is about a quarter wider than the one the truth would allow, and the whole of the extra width is the price of not knowing the relation’s shape at the small end.

With a straight relation the price is smaller: 3.00 for the quadratic’s aware interval against the oracle’s 2.80 and the line’s 2.89. Flexibility costs width where it was not needed, as it cost point-estimate error in the earlier essay, and it is the insurance that keeps the steepening case honest.

The whole trade in one picture

Coverage against width for seven intervals for a league table's small groups, under three shapes of the size relation. Each interval's coverage of groups of twenty or fewer and its average width, 600 tables per shape. straight: one centre, plugged in 86.9% at 3.19, one centre, its uncertainty added 87.2% at 3.22, a fitted line, plugged in 92.8% at 2.78, a fitted line, its uncertainty added 93.9% at 2.89, a fitted quadratic, plugged in 92.1% at 2.78, a fitted quadratic, its uncertainty added 94.2% at 3.00, the true relation and spread 94.9% at 2.80. flattening below the median size: one centre, plugged in 93.9% at 3.32, one centre, its uncertainty added 94.1% at 3.34, a fitted line, plugged in 92.9% at 3.18, a fitted line, its uncertainty added 93.7% at 3.28, a fitted quadratic, plugged in 92.8% at 3.16, a fitted quadratic, its uncertainty added 94.3% at 3.35, the true relation and spread 94.8% at 3.19. steepening below the median size: one centre, plugged in 58.0% at 2.81, one centre, its uncertainty added 58.5% at 2.83, a fitted line, plugged in 83.1% at 1.58, a fitted line, its uncertainty added 88.4% at 1.77, a fitted quadratic, plugged in 85.5% at 1.53, a fitted quadratic, its uncertainty added 93.7% at 1.92, the true relation and spread 94.8% at 1.55.
Fig. 5 Every interval’s coverage of the small groups against its average width, under the three shapes of the size relation: open marks are plugged-in intervals, filled marks add the centre’s uncertainty or use the truth.

Read across the three shapes, the intervals fall into three groups. The plug-in intervals sit left of 95% everywhere, a few points short with a straight or flattening relation and far short with a steepening one. Adding the centre’s uncertainty moves each of them right by widening it, and for the quadratic the move reaches the nominal rate in every shape. One centre stays where it is whatever is added, a vertical pair of marks at 87% and 58% coverage, because its failure is location rather than width.

The true relation’s interval sits on the 95% line at the narrowest width each shape allows, and the distance from it to the quadratic’s aware interval, measured along the width axis, is the cost of not knowing the relation. A flattening relation puts every interval higher on the width axis because more of the groups’ spread is left unexplained, which is also why every method looks good there: a table whose between-group spread is large forgives a misplaced centre, and one whose spread is small does not.

A group from the population’s own tail found that pooling does worse than a group’s own mean for every group far enough from its centre; the intervals here show the same boundary from the coverage side. A small group under a steepening relation is far from any centre that ignores the steepening, and its interval, pulled towards that centre, covers the truth only when the interval is also wide enough to reach back to where the group actually sits.

What is still missing

The centre-aware intervals fall short of 95% by about a point even when the relation is exactly what the model fits — 93.9% for the line and 94.2% for the quadratic with a straight relation. The missing point is the between-group spread’s own uncertainty: every interval here plugs in τ^2\hat\tau^2, and a small group’s interval depends on it through both its width and its weight. What the plug-in forgets and the interval that integrates measured that omission for eight groups, where it cost far more — the plug-in covering 78.8% against the integrated interval’s 95.2% — and a hundred groups estimate the spread well enough that it costs one point here rather than sixteen.

So there are three sources of uncertainty in a small group’s interval, and a league table’s usual interval carries one of them. The group’s own noise is in every interval. The centre’s uncertainty, which this essay adds, matters at the small end and matters most when the relation’s shape there is not known. The spread’s uncertainty matters when the groups are few.

What a league table should print for its small groups

An interval that includes the fitted centre’s uncertainty. It is one extra term, Bi2B_i^2 times the prediction variance at the group’s size, and it restores most of the coverage the plug-in interval loses at the small end.

A centre flexible enough at the small end. A straight line in log size is right when the relation is straight and confidently wrong when it bends below the median. A quadratic with its uncertainty carried covered the small groups 94.2%, 94.3% and 93.7% under the three shapes tried; the line, carried the same way, 93.9%, 93.7% and 88.4%.

Never one centre, when size matters. One centre’s interval covers the groups of eight or fewer 35.2% of the time when the relation steepens, and no variance term repairs it.

Posterior means as a set, not as a ranking. Estimates that are too alike found that pooled estimates spread less than the truth, so a set of honest intervals will overlap more than a reader expects; that is the table being candid about how little separates its small groups.

And the widths, beside the intervals. A small group’s honest interval is wide because the table knows little about it; a league table of a hundred found small groups crowding its extremes on their own means, and an honest interval is the statement of why they should not. Borrowing towards a line established why a group should be pooled towards what its size predicts; the width of its interval is how much the table trusts that prediction.

What was counted

Counted, over six hundred tables of a hundred groups: with a straight relation, the small groups’ coverage of 92.8% and 93.9% for the line plugged in and with its centre’s uncertainty, and 92.1% and 94.2% for the quadratic; with a steepening relation, 83.1% and 88.4% for the line, 85.5% and 93.7% for the quadratic, and 58.0% for one centre; for the groups of eight or fewer under the steepening relation, 77.4% for the plugged-in line, 94.6% for the aware quadratic and 35.2% for one centre. Every width quoted is averaged over the small groups of all six hundred tables, and every coverage counts each small group of each table once.

Not claimed: that a quadratic is flexible enough for every bend. It follows a relation that changes slope once, which is what the steepening and flattening shapes do; a relation with a threshold, or one that turns back, would need more, and its prediction variance at the ends would be larger still. Not claimed either that the moment estimate of the spread is the best available; a restricted-likelihood estimate would move the intervals slightly and would not remove the missing point.

Still open: a ranking with its uncertainty

A league table’s intervals are for effects, and its readers want ranks. The earlier essays found a small group’s rank dominated by its noise and pulled back by pooling; with a fitted centre, a small group’s rank also inherits the centre’s uncertainty and bias, and an interval for a rank — which rank positions the group could plausibly hold — is a different object from an interval for its effect, because ranks depend on every group’s estimate at once.

Whether rank intervals built from a fitted size relation, with the centre’s uncertainty carried, cover the small groups’ true ranks, how wide they are for a group of four patients in a table of a hundred, and whether a bent relation moves the small groups’ ranks more than their effects, are measurements this family of results points at and has not made.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Credible intervalExtrapolationHierarchical modelLeague tableParameter uncertaintyPartial poolingPlug in estimateShrinkage