The interval a small group is given
Worth reading first: The weight that decides.
The relation the table has to estimate fitted a league table’s size relation from the table itself and pooled every group towards it. A straight line in log size lost almost nothing when the relation was straight; when it bent below the median size, the line carried the large groups’ slope down into the region where it was barely observed and misplaced the small groups by −0.173 or +0.264; a quadratic removed nearly all of it. Every number there was a point estimate.
A league table also publishes intervals, and the essay ended on what they inherit. A small group’s estimate is mostly its fitted centre, because its own mean is noisy, so its interval should carry the centre’s uncertainty, and the centre is least certain exactly at the small end, where it is an extrapolation. An interval that treats the fitted centre as known will be too narrow for the small groups, the essay predicted — and the question was how much, and whether the quadratic’s wider uncertainty at the ends makes its intervals honest where the line’s are not.
Three intervals for one group
The table is the one the earlier essays built: a hundred groups from four to four hundred patients on a logarithmic spread, true effects centred on a relation with log size and scattered around it, each group’s mean measured with noise that shrinks with its size. The analyst fits a hierarchical regression — a centre that is one mean, a line or a quadratic in log size, and a between-group spread estimated from the table — and pulls each group towards its fitted centre by the weight .
The plug-in interval is the posterior mean plus or minus 1.96 times : the posterior spread if the fitted centre and spread were the truth. It is the interval a table usually prints.
The centre-aware interval adds the fitted centre’s own variance, multiplied by because the group’s estimate takes a share of its centre. The centre’s variance is the regression’s prediction variance at that group’s size, , which is smallest in the middle of the size range and grows at the ends, fastest for a quadratic.
The oracle uses the true relation and the true spread, and shows what a 95% interval can be when nothing has to be estimated.
Where the centre’s uncertainty lives
Before any coverage is counted, the construction says where the extra term matters. A group’s centre-aware variance is its plug-in variance plus times the centre’s prediction variance at its size, and both factors grow at the small end: the small groups borrow the most, so is near one, and they sit at the edge of the size range, where a fitted polynomial is least certain.
In the table drawn above, the fitted centre accounts for a small share of a large group’s interval under every model, because a group of four hundred barely borrows from its centre. For the group of four the share is 1.3% under one centre, whose single mean is estimated from all hundred groups; 10.4% under a line; and 26.0% under a quadratic, whose curvature is estimated mostly from groups far from the small end and extrapolated to it. That is the plug-in interval’s omission, located: almost nothing for most of the table, and a quarter of the variance for the groups a reader is most likely to look up.
When the relation is straight
With a straight relation, the fitted line’s plug-in interval covers the small groups’ true effects in 92.8% of cases, and the centre-aware interval in 93.9%. The quadratic’s plug-in interval covers 92.1% and its centre-aware one 94.2%, against the oracle’s 94.9%. For the smallest group in the table, of four patients, the quadratic’s plug-in interval covers 90.0% and its centre-aware interval 95.3%.
So the prediction holds and the size is modest when the relation is what the model assumes: a plug-in interval loses two or three points of coverage for the small groups, most of it for the smallest, and adding the centre’s variance restores nearly all of it. The quadratic’s plug-in interval does a little worse than the line’s, because its centre is more uncertain at the ends and the plug-in ignores exactly that; with the uncertainty added the quadratic does a little better, because it now carries its own honesty.
One centre is a different case. Pooled towards one fitted mean, the small groups’ plug-in interval covers 86.9%, and for the group of four, 76.5%. Adding the centre’s variance changes that by a few tenths, to 87.2%, because the single mean is precisely estimated — the problem is not its uncertainty but its position. The small groups one centre protects found that one centre overstates the small groups by half a population width; an interval around an estimate that is half a width off is centred in the wrong place, and no widening short of doubling it will cover.
When the relation steepens
The failure the earlier essay found for point estimates reaches the intervals in full when the relation steepens below the median.
The fitted line’s plug-in interval covers the small groups 83.1% of the time and the groups of eight or fewer 77.4%. Its centre-aware interval covers 88.4% and 84.6% — better, since the line’s extrapolated centre is uncertain as well as biased, but still short, because the line’s centre is biased at the small end by +0.264 and no symmetric widening removes a bias. The quadratic’s plug-in interval covers 85.5% of small groups, and its centre-aware interval 93.7% — and 94.6% of the smallest groups. The quadratic can follow the steepening, so its centre is nearly unbiased at the small end, and its prediction variance there is large enough to cover what it cannot follow.
The line’s remaining shortfall has a simple size. Its centre-aware interval for the small groups averages 1.77 wide, a standard error of about 0.45, and its estimates of those groups are off by +0.264 on average, which is 0.58 of that standard error. An interval of the right width centred 0.58 standard errors from the truth covers 91.0% of the time, not 95%; the counted 88.4% is a little lower because the bias is larger for the smallest groups than for the average small one. No variance term can close that gap, because it is not variance: the interval is the right width around the wrong centre, and only a centre that can follow the bend puts it back.
The coverage by size shows where the failure lives. For groups above the median size every interval covers close to its nominal rate: the large groups’ own means carry them, and the centre hardly matters. Below it, the plug-in line falls steadily, to 71.2% for the group of four; its centre-aware version falls less, to 80.8%; the quadratic with its uncertainty stays near 95% all the way down, at 94.8% for the group of four. The gap between the line’s two intervals is the centre’s variance, and the gap between the line’s aware interval and the quadratic’s is the line’s bias.
One centre is at its worst here. Its plug-in interval covers the small groups 58.0% of the time, the groups of eight or fewer 35.2%, and the group of four 16.2%. A table that pools its smallest groups towards the overall mean when the smallest groups are genuinely worse publishes intervals that miss their truth two times in three and point confidently at a centre they are not part of.
When the relation flattens
The flattening relation is the mild case for intervals, for a reason worth noticing. When the relation flattens below the median, the small groups are no worse than middling ones, and the between-group spread around any fitted centre is larger than with a straight relation — the three shapes were set to have the same total variance, and a flattening relation leaves more of it unexplained. The intervals are correspondingly wider, 3.18 for the line’s plug-in against 2.78 with a straight relation, and wide intervals tolerate a misplaced centre.
So the line’s plug-in interval covers the small groups 92.9% of the time despite a bias of −0.173, its centre-aware interval 93.7%, and the quadratic’s 94.3%. Even one centre covers 93.9%, because under a flattening relation the overall mean is not far from where the small groups sit. The same bias that moved the point estimates by a sixth of a standard deviation costs the intervals a point or two, and the order of the methods is unchanged.
What honesty costs
In the table drawn above, the line’s plug-in intervals miss four of the seventeen smallest groups’ true effects, each by sitting above it, and the quadratic’s centre-aware intervals miss none. They do it by being wider. Averaged over the small groups under a steepening relation, the line’s plug-in interval is 1.58 wide, its aware interval 1.77, the quadratic’s plug-in 1.53 and its aware interval 1.92, against the oracle’s 1.55. The honest interval is about a quarter wider than the one the truth would allow, and the whole of the extra width is the price of not knowing the relation’s shape at the small end.
With a straight relation the price is smaller: 3.00 for the quadratic’s aware interval against the oracle’s 2.80 and the line’s 2.89. Flexibility costs width where it was not needed, as it cost point-estimate error in the earlier essay, and it is the insurance that keeps the steepening case honest.
The whole trade in one picture
Read across the three shapes, the intervals fall into three groups. The plug-in intervals sit left of 95% everywhere, a few points short with a straight or flattening relation and far short with a steepening one. Adding the centre’s uncertainty moves each of them right by widening it, and for the quadratic the move reaches the nominal rate in every shape. One centre stays where it is whatever is added, a vertical pair of marks at 87% and 58% coverage, because its failure is location rather than width.
The true relation’s interval sits on the 95% line at the narrowest width each shape allows, and the distance from it to the quadratic’s aware interval, measured along the width axis, is the cost of not knowing the relation. A flattening relation puts every interval higher on the width axis because more of the groups’ spread is left unexplained, which is also why every method looks good there: a table whose between-group spread is large forgives a misplaced centre, and one whose spread is small does not.
A group from the population’s own tail found that pooling does worse than a group’s own mean for every group far enough from its centre; the intervals here show the same boundary from the coverage side. A small group under a steepening relation is far from any centre that ignores the steepening, and its interval, pulled towards that centre, covers the truth only when the interval is also wide enough to reach back to where the group actually sits.
What is still missing
The centre-aware intervals fall short of 95% by about a point even when the relation is exactly what the model fits — 93.9% for the line and 94.2% for the quadratic with a straight relation. The missing point is the between-group spread’s own uncertainty: every interval here plugs in , and a small group’s interval depends on it through both its width and its weight. What the plug-in forgets and the interval that integrates measured that omission for eight groups, where it cost far more — the plug-in covering 78.8% against the integrated interval’s 95.2% — and a hundred groups estimate the spread well enough that it costs one point here rather than sixteen.
So there are three sources of uncertainty in a small group’s interval, and a league table’s usual interval carries one of them. The group’s own noise is in every interval. The centre’s uncertainty, which this essay adds, matters at the small end and matters most when the relation’s shape there is not known. The spread’s uncertainty matters when the groups are few.
What a league table should print for its small groups
An interval that includes the fitted centre’s uncertainty. It is one extra term, times the prediction variance at the group’s size, and it restores most of the coverage the plug-in interval loses at the small end.
A centre flexible enough at the small end. A straight line in log size is right when the relation is straight and confidently wrong when it bends below the median. A quadratic with its uncertainty carried covered the small groups 94.2%, 94.3% and 93.7% under the three shapes tried; the line, carried the same way, 93.9%, 93.7% and 88.4%.
Never one centre, when size matters. One centre’s interval covers the groups of eight or fewer 35.2% of the time when the relation steepens, and no variance term repairs it.
Posterior means as a set, not as a ranking. Estimates that are too alike found that pooled estimates spread less than the truth, so a set of honest intervals will overlap more than a reader expects; that is the table being candid about how little separates its small groups.
And the widths, beside the intervals. A small group’s honest interval is wide because the table knows little about it; a league table of a hundred found small groups crowding its extremes on their own means, and an honest interval is the statement of why they should not. Borrowing towards a line established why a group should be pooled towards what its size predicts; the width of its interval is how much the table trusts that prediction.
What was counted
Counted, over six hundred tables of a hundred groups: with a straight relation, the small groups’ coverage of 92.8% and 93.9% for the line plugged in and with its centre’s uncertainty, and 92.1% and 94.2% for the quadratic; with a steepening relation, 83.1% and 88.4% for the line, 85.5% and 93.7% for the quadratic, and 58.0% for one centre; for the groups of eight or fewer under the steepening relation, 77.4% for the plugged-in line, 94.6% for the aware quadratic and 35.2% for one centre. Every width quoted is averaged over the small groups of all six hundred tables, and every coverage counts each small group of each table once.
Not claimed: that a quadratic is flexible enough for every bend. It follows a relation that changes slope once, which is what the steepening and flattening shapes do; a relation with a threshold, or one that turns back, would need more, and its prediction variance at the ends would be larger still. Not claimed either that the moment estimate of the spread is the best available; a restricted-likelihood estimate would move the intervals slightly and would not remove the missing point.
Still open: a ranking with its uncertainty
A league table’s intervals are for effects, and its readers want ranks. The earlier essays found a small group’s rank dominated by its noise and pulled back by pooling; with a fitted centre, a small group’s rank also inherits the centre’s uncertainty and bias, and an interval for a rank — which rank positions the group could plausibly hold — is a different object from an interval for its effect, because ranks depend on every group’s estimate at once.
Whether rank intervals built from a fitted size relation, with the centre’s uncertainty carried, cover the small groups’ true ranks, how wide they are for a group of four patients in a table of a hundred, and whether a bent relation moves the small groups’ ranks more than their effects, are measurements this family of results points at and has not made.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- Eight groups, one population — both name hierarchical model, partial pooling, shrinkage
- One population, or two — both name hierarchical model, partial pooling, shrinkage
- Pooling a proportion — both name hierarchical model, partial pooling, shrinkage
- The fewest groups that can borrow — both name hierarchical model, partial pooling, shrinkage
- The slope that borrows — both name hierarchical model, partial pooling, shrinkage
- When borrowing goes wrong — both name hierarchical model, partial pooling, shrinkage
Named objects
A flat tag is an object no other essay names yet.
Credible intervalExtrapolationHierarchical modelLeague tableParameter uncertaintyPartial poolingPlug in estimateShrinkage