What partial pooling does to one group, to the set, and to a ranking

A ranking with its uncertainty

A league table's reader wants a rank, and a rank interval inherits a fitted size relation's errors differently from an effect interval. When the relation steepens below the median size, a fitted line's 95% interval for a small group's effect covers 83.1% of the time; its 90% interval for the same group's rank covers 95.6%, because the line's error lifts the small groups into the crowded middle of the table, where rank intervals are wide. The error surfaces instead in the large groups' ranks, 87.6%, and in width: 36.9 places against the 27.3 the true relation needs. A quadratic's uncertainty has to be drawn once per table, since it moves every small group together — added group by group it over-covers at 92.8%.

Worth reading first: The weight that decides.

The interval a small group is given measured the 95% intervals a league table gives its small groups when every group is pulled towards a size relation fitted from the table. When the relation steepens below the median size a fitted line misplaces the small groups, and its intervals, which treat the line as known, cover 83.1% of them and 77.4% of the smallest. Adding the centre’s own uncertainty, weighted by how much a group borrows from it, repairs most of that, and a fitted quadratic carried that way covers the small groups as it should.

It ended on the number a league table’s reader actually looks at. A table is read as a ranking, and a rank is a different object from an effect: it depends on every group’s estimate at once, so an interval for it has to come from draws of the whole table. The essay asked whether rank intervals built on a fitted relation, with the centre’s uncertainty carried, cover the small groups’ true ranks, how wide they are for a group of four patients in a table of a hundred, and whether a bent relation moves the small groups’ ranks more than their effects. The answers run against what the effect intervals would predict.

Every group's 90% rank interval in one league table whose size relation is steepening below the median size, with its true rankA hundred groups from four to four hundred patients, ranked from their posterior under a fitted quadratic size relation whose uncertainty is drawn once per table. The groups of twenty patients or fewer have rank intervals 30.1 places wide on average, and 34 of their 36 true ranks fall inside; the groups of over a hundred, 32.1 places, with 27 of 30 inside.4102050100200400the group's size, patients (log scale)rank among a hundred, 1 the largest effect125507510090% rank interval, twenty patients or fewer90% rank interval, larger groupsone seeded table; a dot is the group's true rankred dots fall outside their interval
Fig. 1 One league table of a hundred groups from four to four hundred patients, each group’s 90% rank interval drawn against its size, with its true rank as a dot; rank 1 is the largest effect. The intervals come from draws of the whole table under a fitted quadratic size relation whose uncertainty is drawn once per table. The slider sets the shape of the relation.

A rank interval, and three ways to build one

The table is the one used since a league table of a hundred: a hundred groups whose sizes run log-evenly from four to four hundred, true effects centred on a relation with log size and scattered around it, each group’s mean measured with noise that shrinks with its size. The analyst fits a hierarchical regression — a line or a quadratic in log size, and a between-group spread estimated from the table — and each group’s posterior is normal, centred on its pooled estimate.

A rank interval draws the whole table from that posterior two hundred times, ranks the hundred groups in each draw, and takes each group’s 5th and 95th percentile rank. Three versions differ only in what the draws treat as uncertain:

  • plugged in, where the fitted centre and spread are known and each group varies only around its own pooled estimate;
  • inflated group by group, where each group’s variance is increased by its borrowing weight squared times the variance of its fitted centre — the repair that worked for effect intervals, applied to each group separately;
  • centre drawn once, where the relation’s coefficients are drawn from their own sampling distribution once per table, so that an error in the centre moves every group that borrows from it at the same time, and each group is then drawn around the shifted centre.

The second and third give every group the same marginal variance. They differ in whether the centre’s error is one error or a hundred, and for an effect interval that difference cannot matter, because an effect interval reads one group at a time. For a rank it is the whole question.

Where the effect intervals and the rank intervals part

How far the small groups' rank intervals and effect intervals sit from their nominal coverage, relation steepening below the median size. Groups of twenty patients or fewer. Coverage of the 90% rank interval, with its width in places, and of the 95% effect interval: fitted line, plugged in 95.6% over 36.9 places and 83.1%; fitted quadratic, plugged in 87.5% over 28.1 places and 85.5%; quadratic, each group inflated 92.8% over 32.8 places and 93.7%; quadratic, centre drawn once 89.1% over 30.0 places and 93.7%; the true relation 91.6% over 27.3 places and 94.8%.
Fig. 2 How far the small groups’ intervals sit from their nominal coverage — the 90% rank interval and the 95% effect interval built on the same posterior — for five constructions, when the size relation steepens below the median size. The right column is the rank interval’s average width in places of the hundred.

The fitted line, plugged in, is the construction whose effect intervals failed the small groups: 83.1% against a nominal 95%. Its rank intervals for the same groups cover 95.6% against a nominal 90% — too many, not too few. The fitted quadratic, plugged in, fails both, 85.5% for effects and 87.5% for ranks. Inflating each group’s variance repairs the effect intervals to 93.7% and over-repairs the rank intervals, to 92.8%. Drawing the centre once per table leaves the effect intervals where the inflation put them and brings the rank intervals to 89.1%. The true relation, known exactly, gives 94.8% and 91.6%.

The rank figures all sit a little above their targets for a reason that has nothing to do with the centre: a rank is an integer, and an interval from the 5th to the 95th percentile of an integer includes both end values, so it covers slightly more than 90% even when everything else is right. The true relation’s 91.6% is that bias. Against it, the plugged-in line over-covers by four points and the quadratic drawn once under-covers by about two and a half.

Why a misplaced group’s rank interval covers

A fitted line under a steepening relation puts the small groups too high. Their true effects are well below the line’s extrapolation, so their pooled estimates are lifted by about a quarter of a population width, and an effect interval centred there misses a truth it was built to contain.

A rank interval centred there does something else. The lifted groups move into the middle of the table, where there are many groups of similar estimated effect, and a group’s rank there is uncertain across a wide stretch of places: a small shift in its draw moves it past many neighbours. So the rank intervals widen — 36.9 places for the line against 27.3 for the true relation — and a wide interval reaches down far enough to include a true rank near the bottom of the table. The small groups’ ranks are covered because their intervals are wide, and their intervals are wide because they have been put in the wrong place.

The width is the tell. A rank interval that covers because it is a third wider than the truth needs is not a good interval; it is an uninformative one that happens to include the answer. A reader of the table sees a small group with a plausible range of 37 places and concludes the table knows little about it, when the table actually knows it is near the bottom, and would say so if the centre were right.

Where the line’s error goes instead

The error does not disappear. Ranks are a zero-sum quantity — a group lifted into the middle of the table pushes the groups there down — and the line’s misplacement of the small groups surfaces in the ranks of groups it never touched.

How often each group's 90% rank interval covers its true rank, by the group's size, relation steepening below the median size. Over 300 tables. Groups of twenty or fewer: fitted line plugged in 95.6%, quadratic with its centre drawn once 89.1%, true relation 91.6%. Groups of over a hundred: 87.6%, 89.5% and 91.6%.
Fig. 3 How often each group’s 90% rank interval covers its true rank, against the group’s size, when the relation steepens below the median size: the fitted line plugged in, the fitted quadratic with its centre drawn once per table, and the true relation.

The line’s rank intervals over-cover the small groups and under-cover the large ones: 87.6% for groups of over a hundred patients, against 91.6% under the true relation. The large groups’ own estimates are excellent, since they borrow almost nothing and their centre is well determined, so their effect intervals are right. Their ranks are wrong because the small groups have been inserted among them. An analyst who checks the table’s calibration on the large groups, where the data are best, would find the ranks under-covering and would not look at the small end, where the cause is.

That is the sense in which a bent relation moves the small groups’ ranks less than their effects and moves the table more. The effect error stays with the group that has it. The rank error is shared out across the groups that group was placed among.

One error in the centre, not a hundred

The quadratic’s centre is nearly right on average and uncertain at the small end, where it extrapolates. Carrying that uncertainty into an effect interval needs only each group’s own share of it, and the effect intervals built that way cover 93.7% of the small groups, close to 95%.

Carried into a rank interval the same way, group by group, it over-covers at 92.8% and widens the intervals to 32.8 places. The reason is that independent inflation makes the small groups’ errors independent of each other, so in each draw they scatter among themselves and reorder: a group that was fifth-lowest is sometimes twentieth. The centre’s real error is one number, or two or three coefficients, that moves all the small groups together. It shifts the whole small end of the table relative to the large groups and leaves the small groups’ order among themselves nearly untouched. Drawn that way, once per table, the rank intervals are narrower, 30.0 places, and cover 89.1% — below the integer-inflated 91.6% of the true relation by about two and a half points, which is what remains of the quadratic’s extrapolation after its own uncertainty is honestly carried.

The smallest groups, eight patients or fewer, show the same ordering more sharply, because they borrow most from the centre and so carry most of its error. Under the steepening relation their rank intervals cover 96.2% with the line plugged in, 94.8% with the quadratic inflated group by group, 89.9% with its centre drawn once, and 92.5% under the true relation. The group-by-group inflation over-covers them by more than it over-covers the small groups as a whole, since the share of their variance it scatters independently is the largest; the joint draw sits the same two and a half points below the true relation that it sits for the small groups generally, which says its shortfall is the quadratic’s extrapolation error and not a feature of how the error is drawn.

The distinction is invisible for effects and decisive for ranks, and it is the general one for any quantity that depends on several groups at once: a difference between two small groups, the share of small groups below a threshold, the chance a small group is in the bottom ten. Every such quantity needs the posterior’s joint structure, and a shared centre makes the joint structure strongly correlated at the small end. What the plug-in forgets found the estimated spread’s uncertainty missing from each group’s interval; here the estimated centre’s uncertainty is present in each group’s interval and missing from their relation to each other unless it is drawn as one error.

How many places a small group’s rank can be narrowed to

How many of a hundred places a group's 90% rank interval spans, by the group's size. Average width over 300 tables. Under a straight relation and the true centre, the group of four spans 56.5 places and the group of four hundred 18.8; under a steepening relation, 16.0 and 16.9, and 25.8 for the group of four under a fitted line.
Fig. 4 The average width of a group’s 90% rank interval, in places of the hundred, against the group’s size: under a straight relation and under a steepening one, with the true centre, and under a steepening relation with a fitted line.

When the relation is straight, a small group’s rank is close to unknown. The group of four patients has a 90% rank interval 56.5 places wide on average, and the groups of ten to twenty are wider still, over sixty places — more than half the table. The largest groups’ intervals are 18.8 places. A league table that prints a small group at position 43 is printing one draw from a range that runs from about 15th to about 75th, which is the measured form of a group from the population’s own tail’s warning that a pooled estimate is a statement about the population as much as about the group.

When the relation steepens below the median, the same small groups have narrow rank intervals — the group of four spans 16.0 places — because the relation itself puts them near the bottom of the table and the bottom of a table is a boundary: a group whose effect is certainly low cannot rank lower than a hundredth, and the interval is compressed against the end. The size relation, in other words, is most of what a small group’s rank is known from, and a table that fits it correctly can rank its small groups better than their own data could. A fitted line under the same relation gives the group of four 25.8 places, and gives back most of that gain.

When the relation flattens instead

The steepening relation is the case where the small groups belong at the bottom. The opposite bend — small groups no worse than middling ones, the relation flattening below the median — puts them in the middle of the table, and the middle is where a rank is least certain.

Under that relation the true relation’s rank intervals for the small groups span 75.3 places, and the group of four patients’ spans 81.0: a small group in a table whose relation flattens can honestly be placed anywhere from about the tenth place to about the ninetieth. The fitted line, which now understates the small groups by carrying the large groups’ slope down into a region where the relation has stopped falling, pushes them towards the bottom of the table and narrows their intervals a little, to 71.9 places, and those narrower intervals under-cover: 88.1% of the small groups and 86.8% of the smallest, against 90.3% and 90.1% for the true relation. So the two bends fail in opposite directions at the level of the rank, as they did at the level of the effect — a line misplaces the small end towards the middle under one bend and towards the edge under the other — but the rank interval’s width responds to the misplacement, widening in the crowded middle and narrowing near an edge, so coverage moves in the direction the width pushes it. Over-coverage by a wide, misplaced interval under one bend, under-coverage by a narrower misplaced one under the other.

The quadratic with its centre drawn once per table covers 89.3% of the small groups under flattening, 89.4% of the smallest and 90.5% of the large, within about a point of the true relation at every size. It is the one construction measured here that is close to right under all three shapes, which is the rank-interval version of what borrowing towards a line and the essays after it found for point estimates: a centre flexible enough to follow the relation, with its own uncertainty carried, is the only centre that can be trusted at the small end, because the small end is exactly where the table has least to say about the relation’s shape.

It also answers, for ranks, the question the small groups one centre protects raised for effects — whether pooling towards a single centre ever serves the small groups — from the other side. A single centre would put every small group in the middle of the table whatever the true relation, and a small group’s rank interval there is as wide as the table allows. The rank would be covered; it would say almost nothing.

What a league table can print about its small groups

A rank interval, not a rank, for every group below about fifty patients. Under a straight size relation their 90% intervals span more than half the table.

Rank intervals drawn from the whole table with the centre’s error drawn once per table. Inflating each group separately makes the intervals a tenth wider than they should be and over-covers; treating the centre as known under-covers when it is a quadratic and hides the line’s misplacement behind wide intervals.

A check of the large groups’ rank coverage is a check of the small end. When the centre is wrong at the small end, the large groups’ ranks under-cover and their effects do not; a table whose large groups’ ranks are miscalibrated has a problem where its data are thinnest.

Each coverage is over three hundred simulated tables of a hundred groups, two hundred posterior draws a table; the effect-interval coverages are the earlier essay’s, over six hundred tables. The rank intervals are the 5th to 95th percentiles of integer ranks and include both ends, which is why the true relation’s rank intervals cover 91.6% rather than 90%; every comparison here is against that figure rather than against 90%. Not measured: a ranking by the posterior probability of being in the bottom ten, which a league table of a hundred set beside ranking by posterior means for the top ten, under a fitted and bent relation; relations whose shape is uncertain enough that a choice between the line and the quadratic is itself made from the table; and tables in which the between-group spread changes with size.

Still open: a flag that knows its centre was chosen

A league table’s most consequential output is usually not a rank at all but a flag — a group placed in the bottom ten, or below a threshold, and investigated. The earlier essays found that pooling protects small groups from being flagged by noise, and this one that the protection depends on the shape the table was pooled towards. In practice the shape is chosen from the table: a line if it looks straight, a quadratic if it bends.

That choice is made from the same data the flag is computed from, which is the situation the relation the table has to estimate left for later. Whether a table that picks its centre by a fit statistic flags the right small groups as often as one told the shape in advance, and whether the joint draw of the centre can carry the uncertainty of the choice as well as the coefficients, is measurable on the same tables and has not been measured. It would say whether a published bottom ten built this way can be defended to the small groups in it.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Credible intervalHierarchical modelLeague tableParameter uncertaintyPartial poolingPlug in estimateRankingShrinkage