A group from the population's own tail
Worth reading first: The weight that decides.
The weight that decides derived the shrinkage weight three ways and found them in agreement, and what the plug-in forgets priced estimating the population spread that the weight needs. Both judged partial pooling the way it is designed to be judged: by the total squared error across all the groups, which it minimises.
A total can be minimised by making most terms much smaller and a few larger. When borrowing goes wrong found one way that happens — a group that was never from the population, placed six widths out, estimated six times worse than by its own mean. That group was an intruder. This essay asks about the groups that belong: ordinary members of a normal population who happen to sit in its tail.
One group’s error, not the total
Take the simplest case, with the population’s centre and spread known and every group measured with the same standard error. The posterior mean shrinks a group’s observed mean towards the centre by the weight . For a group whose true effect is from the centre, its expected squared error is
against for its own mean. The first term is the bias the shrinkage introduces, which grows with the square of how far the group really is from the centre. The second is the variance, which shrinkage always cuts.
Averaged over a normal population the total is — half the unpooled error when the standard error equals the spread. Group by group, the bias term eventually wins, and it wins exactly when
With the standard error equal to the spread, and the break-even is = 1.732 population widths. Every group beyond that distance is estimated worse by pooling than by its own mean, and 8.33% of a normal population is beyond it. None of those groups is unusual. A group at 1.8 widths is at the 96th percentile of an ordinary bell curve.
The hero figure draws the ratio. At the centre, pooling’s error is a quarter of the group’s own; at the break-even it is exactly one; at three widths it is 2.5; at six it is 9.25, and it has no ceiling.
Where the standard error matters
The break-even distance and the share of groups beyond it both depend on how noisy each group’s own mean is.
| each group’s standard error | shrinkage weight | pooling loses beyond | share of groups there | at least one of eight |
|---|---|---|---|---|
| half a width | 0.20 | 1.500 widths | 13.36% | 68.25% |
| one width | 0.50 | 1.732 widths | 8.33% | 50.12% |
| two widths | 0.80 | 2.449 widths | 1.43% | 10.89% |
The direction is the opposite of what a reader might guess. When the groups are measured well, pooling helps less and more groups pay for it: a small weight means a small total gain, and the bias it introduces is enough to overtake a small variance over a large share of the population. When the groups are measured badly, pooling helps a great deal and few groups pay — but the ones that do pay heavily, because a weight of 0.8 pulls a group at 2.5 widths almost all the way to the centre.
The last column is the practical one. With the standard error equal to the spread, a set of eight groups contains at least one member of the population that pooling serves worse than its own mean half the time, before any group is an intruder and before anything has been estimated.
Groups of different sizes
Real groups are not measured equally, and the break-even distance is different for each. With the unit-level spread three times the population’s and groups of three, ten and forty units:
| group size | shrinkage weight | loses beyond | share of the population there | extra error at 2.5 widths | ratio at 2.5 widths |
|---|---|---|---|---|---|
| 3 | 0.750 | 2.236 widths | 2.53% | 0.703 | 1.234 |
| 10 | 0.474 | 1.703 widths | 8.86% | 0.752 | 1.835 |
| 40 | 0.184 | 1.492 widths | 13.58% | 0.136 | 1.603 |
Two things change together as a group gets larger. Its break-even distance falls — towards = 1.414 widths as the group’s own error vanishes, so even a perfectly measured group loses to its own mean if it sits beyond 1.414 widths — and the share of the population beyond it rises accordingly. But the absolute cost falls much faster, because a well-measured group is barely moved. At 2.5 widths the group of forty pays an extra 0.136 in squared error, a fifth of what the group of ten pays.
So the groups most likely to be on the wrong side of the break-even are the large ones, and the groups that pay most when they are there are the middling ones — large enough to have a modest weight, small enough that the modest weight is a long shift. The smallest groups are pulled hardest but, measured so badly, have most to gain from any pull at all.
With the spread estimated, as an analyst has it
The formula assumes the centre and the spread are known. In practice both are estimated from the same groups, and the estimate of the spread is itself unreliable at eight groups. So the measurement that matters is on data an analyst would actually have: eight groups drawn from the population, the spread estimated from the spread of their means, and the group whose true effect is furthest from the centre singled out after the fact.
Over twenty thousand such sets, with each group’s standard error equal to the population’s spread:
- across all eight groups pooling’s squared error is 0.647 of the unpooled error — a clear gain, and short of the known-spread 0.5 by the price the plug-in pays for estimating the spread;
- for the group truly furthest from the centre it is 1.218 of the unpooled error;
- and that group is worse off than with its own mean in 61.6% of the sets.
The most extreme member of any set of eight is the one the analysis is most likely to be asked about — the best-performing school, the worst hospital, the region where the effect was largest — and it is the one pooling is most likely to serve badly. The total gain is real, and it is concentrated in groups nobody was asking about; the loss is concentrated in the one everybody was.
The extreme group an analyst can see
There is a catch in the measurement above, and it matters more than any number in it. The group “truly furthest from the centre” was picked using the true effects, which no analyst has. An analyst picks the group whose observed mean is furthest out, and those are different groups.
For the group whose observed mean is furthest from the rest, pooling’s squared error is 0.417 of the unpooled error, and it is worse off than with its own mean in only 26.8% of sets. That is a larger gain than pooling achieves across the whole set.
The reason is the one regression to the mean describes. A group observed furthest out is, more often than not, a group that is somewhat far out and was measured high. Its own mean overstates it, and pulling it towards the centre is exactly the right correction; the posterior mean is, after all, the best estimate given the observation. The group that loses is the one that really is far out, and among eight groups it is only sometimes the one that looks furthest out.
So the two findings are compatible and both true. Selected by what can be seen, the extreme group gains most from pooling. Selected by what is true, it loses. The loss falls on a group the analysis cannot point to: genuinely exceptional members of the population, whose observed means are exceptional too but not always the most exceptional, and who are pulled towards the centre along with the lucky ones. Pooling is right about the typical group that looks extreme and wrong about the atypical group that is.
That is why the loss is hard to see in any single analysis. Every check an analyst can run conditions on the observations, and conditional on the observations pooling is doing its job. The cost appears only in a calculation that conditions on the truth, which is the calculation a decision about one real group implicitly needs.
A cap on how far a group may move
The loss comes from one source: a group far from the centre is pulled a long way, and the pull is proportional to how far it is. The repair that follows is to limit the pull. Efron and Morris called it limited translation: shrink as usual, but never move any group by more than standard errors.
With the cap, a group’s expected error can never exceed , because the estimate is never further from the group’s own mean than standard errors. The cost is that the groups near the centre, which would have moved less than the cap anyway, are unaffected, and the groups in the middle distance lose some of their benefit.
| cap on the shift | total error | worst group’s error |
|---|---|---|
| none | 0.500 | 9.25 by six widths, unbounded |
| 2 standard errors | 0.500 | 4.96 |
| 1 standard error | 0.528 | 2.00 |
| half a standard error | 0.640 | 1.25 |
| zero | 1.000 | 1.00 |
A cap of one standard error keeps 94% of pooling’s total gain and holds every group’s expected error under twice what its own mean would give. A cap of two standard errors gives back essentially nothing — the total is 0.5004 — and still bounds the worst case at five times, where the uncapped estimator has no bound at all. The shape of the trade is steep at the start and flat at the end: the first units of protection for the tail are almost free, because the groups they protect are rare and the shift they prevent is large.
What the arrows in the usual picture hide
The standard picture of partial pooling draws each group’s own mean, an arrow towards the centre, and the pooled estimate at the arrow’s head.
The picture shows the arrows’ lengths, which depend on each group’s size and not on where it sits. It cannot show which arrows went too far, because that depends on the truth, which the picture marks but a reader of a real analysis never has. In the drawn example most arrows carry their group towards its truth. The ones that carry a group past it, or away from it, are the ones whose true effect was genuinely far out — and a reader of the same figure without the truth marks would see arrows of exactly the same kind.
A report that wants to be honest about the tail can do two things the picture does not. It can state each group’s shift beside its estimate, so that a reader of one group knows how much of the reported value came from the other groups rather than from its own data. And it can state the cap, if one was used, as the most that any group’s estimate was moved — which turns an invisible bias into a stated bound, in the units the reader is already using.
With the spread estimated, the cap costs nothing
The trade in the table above is computed with the population’s spread known. With it estimated from the same eight groups, the cap does something the known-spread arithmetic cannot show.
Rerunning the twenty thousand sets of eight with the shift capped at one standard error:
| uncapped | capped at one standard error | |
|---|---|---|
| all eight groups | 0.647 | 0.626 |
| the group truly furthest out | 1.218 | 1.077 |
| the group observed furthest out | 0.417 | 0.365 |
Every row improves. The cap lowers the error of the whole set, of the truly extreme group and of the observed extreme group at once. It is not trading total accuracy for protection of the tail; it is removing a failure of the estimated weight.
The failure is the one a spread that estimates to zero describes. On a substantial share of sets of eight the estimated spread comes out at or near zero, the weight goes to one, and every group is pulled all the way to the grand mean — a complete pooling that the data did not support. Those are the sets in which the shifts are largest and most wrong, and the cap is precisely a limit on the largest shifts. With the spread known the cap only ever binds on genuinely far groups, and costs a little; with it estimated, it mostly binds on the sets where the estimate of the spread collapsed, and pays for itself.
So the practical recommendation is stronger than the theory that motivates it. Where the population’s spread must be estimated from a handful of groups, a cap on each group’s shift is not a concession to the tail at the expense of the total. On this measurement it is simply a better estimator.
Why the total is the wrong promise for some readers
None of this contradicts the result that pooling dominates. The dominance is a statement about the sum of squared errors over all groups, and the sum is lower. It says nothing about any particular group, and the Stein result that underwrites it is explicit about that: the gain is collective.
Whether the collective promise is the right one depends on who reads the estimates. A funder allocating across all schools pays the total and should take the pooled estimates as they are. A family choosing one school, a regulator deciding about one hospital, a clinician reading one trial’s subgroup, each experiences one term of the sum — and if that term belongs to a group in the population’s tail, the pooled estimate is the worse one for them, systematically, in the direction of making the unusual group look ordinary.
That is exactly the direction that matters most for decisions about outliers. A genuinely excellent school is estimated as merely good; a genuinely dangerous hospital as merely poor. The cap is the adjustment that keeps most of the collective gain while refusing to spend more than a stated amount of any single group’s accuracy on it, and the amount is a choice that can be stated in the report.
What the formulas establish, and what they do not
With the standard error equal to the spread, pooling starts losing at exactly population widths, and the quadrature that computes the capped estimator puts the uncapped risk at exactly the unpooled error there — two routes to one crossing.
Capping the shift at one standard error bounds the worst group’s risk at twice the unpooled error, while without the cap the worst risk passes five times by six population widths and keeps growing; the cap gives back little of the total — 0.528 against 0.500.
Pooling lowers the error across eight groups with the spread estimated and raises the most extreme group’s, 0.647 and 1.218 of the unpooled error.
Selected by its observed mean, the most extreme of eight gains most from pooling — 0.417 of its own error — and selected by its true effect it loses. The two selections are reported together because either alone would mislead: the first is the one an analyst can make and the second is the one a decision about a real group depends on.
What does not survive is partial pooling read as better for every group. A group at 2.5 population widths, with the standard error equal to the spread, has expected error 1.81 times its own mean’s.
Not claimed: that the cap of one standard error is the right cap. It is one point on a trade whose shape is tabulated, and the right point depends on how much a reader cares about the worst group relative to the total. Nor is the extreme-of-eight figure a statement about intruders — every group in that simulation was drawn from the population — or about groups of unequal size, where the weights differ and the break-even distance differs with them.
Still open: the estimates read together
Everything here is about one group’s estimate at a time. A report of a hierarchical analysis usually also uses the estimates together: the spread of the school effects, the number of hospitals above a threshold, a histogram of the regional rates. The posterior means that minimise each group’s error are not the estimates that get those collective questions right, because shrinking every group towards the centre also shrinks the spread of the set.
How far too narrow the set of posterior means is, how badly that miscounts the groups beyond a threshold, and what it costs to have estimates that are right as a set, are the questions of estimates that are too alike.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The fewest groups that can borrow — both name empirical bayes, hierarchical model, james–stein, mean squared error, partial pooling, shrinkage
- One population, or two — both name empirical bayes, hierarchical model, partial pooling, shrinkage
- Borrowing towards a line — both name hierarchical model, partial pooling, shrinkage
- Pooling a proportion — both name hierarchical model, partial pooling, shrinkage
- The interval that integrates — both name empirical bayes, hierarchical model, shrinkage
- The slope that borrows — both name hierarchical model, partial pooling, shrinkage
Named objects
A flat tag is an object no other essay names yet.
Empirical BayesHierarchical modelJames–SteinLimited translationMean squared errorPartial poolingShrinkageWorst case