What partial pooling does to one group, to the set, and to a ranking

A centre chosen by its own table

A league table that picks its centre's shape from its own data — a line unless a curvature test says otherwise — picks the line on 85.5% of tables whose relation flattens below the median, because the bend lives where the groups are smallest and noisiest. Its bottom ten then flags like the line's: 2.61 innocent small groups a table against 2.19 for a quadratic told the shape. Always fitting the quadratic costs 0.07 of a correctly flagged group when the relation is straight and matches the told answer when it bends, and a flag set by probability carries the choice through a mixture of the two.

Worth reading first: The weight that decides.

A league table’s most consequential output is a flag. A hospital in the bottom ten of a hundred is investigated, a school below a line is put under review, and the small units are where a flag is most likely to be wrong, because their means are the noisiest numbers in the table. The small groups one centre protects found that pooling every group towards one average protects the small ones from noise and then flags them for being small, and that pooling towards a centre that rises with size puts them back where they belong. The relation the table has to estimate fitted that centre from the table: a line in log size if the relation is straight, a quadratic if it bends where the small groups are.

Every one of those essays was told which shape to fit. A ranking with its uncertainty ended on what an analysis actually does: it looks at the table, and if the relation looks straight it fits a line, and if it bends it fits a quadratic. The decision is made from the same hundred means the flag is computed from. Whether a bottom ten built that way flags the right small groups as often as one built on the true shape is the question a group on the list would ask, and it can be counted.

The table and the choice

The tables are the earlier ones. A hundred groups with sizes running log-evenly from four patients to four hundred, true effects centred on a relation with log size and scattered around it, each group’s mean measured with noise that shrinks with its size. Three true relations agree above the median size and differ below it, where the small groups are: straight; flattening, so that small groups are no worse than middling ones; and steepening, so that they fall away twice as fast. The analyst fits a hierarchical regression with a between-group spread estimated from the table and pools each group towards the fitted centre. The ten lowest pooled estimates are the flag.

The choice is made the way a careful analysis would make it. The quadratic is fitted, and its curvature coefficient is tested against its own standard error at 5%; if it is significant the quadratic is kept, and otherwise the line. That is the chosen centre. Beside it sit three rules that make no choice: always a line, always a quadratic, and told the shape — the line when the truth is straight and the quadratic when it bends, the analysis an oracle would run.

How often the table sees the bend

How often a league table's own data choose a bent centre over a straight one. Over 400 tables of a hundred groups each. The quadratic's curvature coefficient is significant at 5% on 5.3% of tables whose relation is straight, 14.5% of tables whose relation is flattening below the median, 29.3% of tables whose relation is steepening below the median; its average AIC weight is 0.374, 0.456, 0.558.
Fig. 1 How often a table’s own curvature test keeps the quadratic, and the quadratic’s average weight under Akaike’s criterion, for the three true shapes. Four hundred tables of a hundred groups each.

On a straight relation the test keeps the quadratic on 5.3% of tables, as a 5% test should. On a relation that flattens below the median it keeps it on 14.5%; on one that steepens, on 29.3%. Both bends are large — each moves the smallest groups’ centre by 0.76 of the true effects’ standard deviation, one way or the other — and the table sees them less than a third of the time.

The reason is where the bend is. The curvature coefficient is estimated from all hundred groups, but the bend lives below the median size, among the fifty groups whose means are measured with the most noise: a group of four patients has a standard error ten times a group of four hundred’s. The precise groups lie on the part of the relation where all three shapes agree. A test of curvature in that table is a test run mostly on the data that cannot see the curvature, and the fitted centre’s own uncertainty, which the interval a small group is given measured, is largest exactly where the bend is.

So a chosen centre is, in practice, a line. On the flattening relation the table chooses the line on 85.5% of tables, and on the steepening one on 70.8%. The choice does not add a little error to the told answer; on most tables it is the wrong answer, made with a test’s confidence.

What the chosen centre flags

Small groups a published bottom ten flags though they are not in it, by the centre the table was pooled towards. Per table of a hundred, over 400 tables. straight: always a line 4.59 (holding 4.25 of the true ten), chosen by the curvature test 4.57 (holding 4.22 of the true ten), always a quadratic 4.42 (holding 4.18 of the true ten), told the shape 4.59 (holding 4.25 of the true ten); flattening below the median: always a line 2.86 (holding 4.48 of the true ten), chosen by the curvature test 2.61 (holding 4.51 of the true ten), always a quadratic 2.19 (holding 4.50 of the true ten), told the shape 2.19 (holding 4.50 of the true ten); steepening below the median: always a line 3.44 (holding 6.55 of the true ten), chosen by the curvature test 3.44 (holding 6.55 of the true ten), always a quadratic 3.40 (holding 6.59 of the true ten), told the shape 3.40 (holding 6.59 of the true ten).
Fig. 2 Per table, the groups of twenty patients or fewer that the published bottom ten names although they are not in the true bottom ten, for four centres under each true shape, with how many of the true ten each flag holds.

The flag’s first number barely moves. Whatever the centre, the bottom ten holds about 4.2 of the true bottom ten when the relation is straight, 4.5 when it flattens and 6.6 when it steepens. That is the table’s noise speaking: a hundred groups, half of them small, and a flag that is right less than half the time on two of the three shapes whatever is done to the centre.

The second number is the one a small group cares about, and it moves where the line is wrong. When the relation flattens below the median, the true bottom ten holds 5.17 small groups on average, and a line fitted through the large groups’ slope extends it into the small end and places the small groups too low. Its flag names 2.86 small groups a table that are not in the true bottom ten. The quadratic, told or not, names 2.19. The line puts two thirds of an innocent small group on the list per table more than the told shape does: across a programme that publishes a table a year, two more small groups wrongly flagged every three years.

The line’s error is not simply that it lists more small groups. It lists 4.58 small groups a table, fewer than the 5.17 the true bottom ten holds, and 1.72 of them belong there; the told quadratic lists 3.65 and 1.46 of them belong. The line moves the flag towards the small end and then picks the wrong members of it, because below the median it ranks small groups by how far its extrapolated slope carries them rather than by their own data. A small group pooled towards a line is judged partly on the large groups’ trend, which is the weight that decides doing exactly what it should for a centre in the wrong place.

The chosen centre names 2.61. Because the test keeps the quadratic on only 14.5% of these tables, the choice takes back about a third of the line’s excess and leaves the rest. On a steepening relation the line places the small groups too high, which shelters them, and the flags of all four centres agree to within a twentieth of a group: when the true bottom ten holds 9.99 small groups on average, the centre’s shape decides little about which small groups fill it. The straight relation is where a quadratic could cost, and it costs little — it holds 4.18 of the true ten against the line’s 4.25, and it names 4.42 innocent small groups against the line’s 4.59.

That last pair is the essay’s practical finding. Fitting the quadratic always costs 0.07 of a correctly flagged group when the relation is straight, flags no more innocent small groups there than the line, and matches the told answer exactly when the relation bends either way. The chosen centre does worse than it on the flattening relation and no better on the other two. The choice is not a compromise between the line and the quadratic; it is the line most of the time, with a test’s authority attached.

Why choosing is worse than not choosing

Borrowing towards a line introduced the centre that rises with size as a repair, and every essay since has treated its shape as something the analyst knows. The repair was always conditional on that, and the condition is the one an analysis of a real table cannot meet without looking.

The relation the table has to estimate found that the quadratic’s extra flexibility costs 0.025 of squared error at the small end when the relation is straight — a curvature estimated from noisy groups is slightly noisy. That small price is the whole case for choosing: pay it only when the bend is there.

The trouble is that the test cannot tell when it is there. Its power at the bends that matter is between a seventh and a third, so it avoids the small price on nearly every straight table and on most bent ones, where the price was worth paying. A rule that saves a small cost when it is not needed and forgoes a large saving when it is has its asymmetry backwards. This is the shape the interval after the choice priced for intervals: a test used to decide between two analyses answers the question it was asked, about the coefficient, and not the question the decision depends on, about the groups.

There is a second, quieter cost. On the tables where the test does keep the quadratic, it keeps it because the curvature came out large, and a curvature estimate selected for being large overstates the bend. On the flattening relation the fitted curvature averages 0.147 over every table and 0.372 over the tables that keep it — two and a half times as bent; on the steepening one, −0.163 and −0.294. So the chosen centre is a line on most tables and an exaggerated quadratic on the rest. The two errors push the small groups in opposite directions, and on any one table a reader cannot know which of them it carries.

Flags set by a chance rather than a count

A bottom ten is a count, and the essay before this one argued for something better than a count: each group’s posterior chance of being in the bottom ten, from draws of the whole table with the centre’s uncertainty drawn once per draw, and a flag for every group whose chance exceeds a half. That flag is not obliged to name ten groups. It names the ones the table is more sure of than not.

That construction also offers a way to carry the choice instead of making it. The table’s draws can come from the line on some draws and the quadratic on others, in proportion to the two models’ weights under Akaike’s criterion — each model’s fit to the hundred means, charged one unit for each coefficient it estimates. Those weights do not jump: the quadratic’s averages 0.374 on straight tables, 0.456 on flattening ones and 0.558 on steepening ones, so a mixture spends more than a third of its draws on the curved centre even where the test would almost never keep it, and more than half where the relation steepens.

One league table's flag chances when its own test keeps a straight centre the truth bends away from. A steepening relation; the curvature statistic is −1.32, so the test keeps the line, and the quadratic's AIC weight is 0.46. Flagging every group whose chance exceeds a half, the line flags 6 groups, 4 of them truly in the bottom ten; the mixture flags 11, 6 of them truly there.
Fig. 3 One steepening table on which the curvature test keeps the line. Each group’s chance of being in the bottom ten, by size, under the line drawn jointly (open) and under the mixture of line and quadratic (filled); the dashed line is the flag at a half.

The table drawn above is ordinary for its kind. Its curvature statistic is −1.32, so the test keeps the line; its quadratic’s weight is 0.46, so the mixture is close to even. Under the line, six groups have a better-than-even chance of being in the bottom ten, and four of the six truly are. Under the mixture, eleven do, and six of them truly are. The five groups that cross the line are all of five or six patients, each held just below a half by the line — between 0.476 and 0.491 — because it placed the small end too high. The mixture lets the curved centre lower them on half its draws and lifts every one of them just over. Two of the five are truly in the bottom ten and three are not: on this table the mixture finds two more of the true ten and names three more innocent groups, which is the price of flagging a cluster of groups whose chances sit on the line.

Groups flagged for a chance above a half of being in the bottom ten, and how many of them truly are. Per table of a hundred, over 400 tables of 200 draws each. straight: the chosen centre, drawn flags 2.49, 1.31 truly in the bottom ten, the choice carried by AIC weights flags 2.55, 1.34 truly in the bottom ten, the told centre, drawn flags 2.40, 1.33 truly in the bottom ten; flattening below the median: the chosen centre, drawn flags 3.12, 2.04 truly in the bottom ten, the choice carried by AIC weights flags 3.15, 2.07 truly in the bottom ten, the told centre, drawn flags 3.44, 2.22 truly in the bottom ten; steepening below the median: the chosen centre, drawn flags 7.29, 5.04 truly in the bottom ten, the choice carried by AIC weights flags 7.93, 5.47 truly in the bottom ten, the told centre, drawn flags 8.35, 5.68 truly in the bottom ten.
Fig. 4 Groups flagged for a chance above a half, and how many of them are in the true bottom ten, from the chosen centre drawn jointly, from the mixture weighted by Akaike’s criterion, and from the told centre — per table, under each true shape.

Over four hundred tables the pattern holds. On the steepening relation the chosen centre’s probability flag names 7.29 groups a table and 5.04 of them are truly in the bottom ten; the mixture names 7.93 and 5.47; the told centre 8.35 and 5.68. Carrying the choice through the weights recovers two thirds of what choosing loses. On the flattening relation the three agree to within a fifth of a group — 2.04, 2.07 and 2.22 true flags — and on the straight one to within a twentieth. Where the choice did not matter, carrying it costs nothing.

The probability flag is also the honest one. It names three groups a table when the relation flattens and two and a half when it is straight, against a count’s ten, because those are the groups a table of a hundred can name with better-than-even confidence. Between a half and seven in ten of its flags are right, against four or five in ten for a count, and every flag carries its chance with it. A count of ten flags whatever the evidence; a chance flags only what the evidence supports, and the mixture makes that support include the analyst’s uncertainty about the centre’s shape.

What a published bottom ten should be built on

Fit the flexible centre and do not test it away. A quadratic in log size costs a few hundredths of a correctly flagged group on a straight relation and matches a told shape on bent ones. The curvature test it would be tested by has between a seventh and a third of the power the decision needs, because the bend is where the noise is.

If two centres are plausible, draw from both. Weights under Akaike’s criterion are a smooth version of the choice, and they keep weight on the curved centre when the table cannot reject it — which, at the small end of a league table, is nearly always. A flag that knows its centre was chosen is a flag computed from a mixture.

Flag by chance, not by count. A count of ten names the bottom ten whether the table can see them or not, and the shape of the centre decides which innocent small groups fill the places it cannot. A chance above a half names fewer groups, is right about two times in three, and lets the centre’s uncertainty, including the uncertainty of its shape, reach the decision.

Read a small group’s chance, not its place. On the table drawn above, five groups of five or six patients sit between 0.476 and 0.491 under the line and between 0.501 and 0.576 under the mixture — flagged or not by the choice of centre alone. A chance that close to a half is the honest report for those groups, and a league table of a hundred showed how wide such chances are for the groups smaller than twenty.

What was counted, and what was not

Each number is over four hundred simulated league tables of a hundred groups at each true shape, the probability flags from two hundred draws of each table and the single table from a thousand. The curvature test is the quadratic coefficient against its standard error at the fitted between-group spread; Akaike’s criterion uses each centre’s marginal likelihood with the spread estimated, charging three parameters for the line and four for the quadratic. Small groups are those of twenty patients or fewer.

Not claimed: that a quadratic is flexible enough for every relation, or that a relation can only bend below the median. A relation that bends among the large groups is seen easily by any test, and is not the case a league table’s small groups are exposed to. Not measured: a test of curvature using the large groups alone, which has no power at the small end by construction, nor a smoother whose flexibility is itself estimated, which is the continuous version of this choice. The between-group spread is the same at every size, as in every earlier essay on these tables.

Still open: a spread that changes with size

Everything above lets the centre bend and holds the scatter around it fixed. Real league tables suggest the opposite as well: small units are often more variable among themselves than large ones, not only noisier in measurement — a small hospital’s true quality varies more from its neighbours’ than a large one’s does. A between-group spread that grows as size falls changes what pooling does to the small end. Each small group borrows less, because the population it borrows from is wider, and its flag chance rises towards what its own mean says.

Whether a spread modelled as a function of size, estimated from the same table, protects the innocent small groups the way a told centre does or undoes that protection, and whether a table can tell a bent centre from a widening spread at all — both put the small groups lower, by different routes — is measurable on the same hundred groups and has not been measured. It is the next decision a league table makes about its small end without being told.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Hierarchical modelInformation criterionLeague tableModel selectionParameter uncertaintyPartial poolingRankingShrinkage