Groups that borrow

When borrowing goes wrong

Partial pooling wins on the total and can lose badly on one group. Placed six population widths out, the group that was never from the population is estimated six times worse than by its own mean — and nothing in the output says so.

Worth reading first: Eight groups, one population · The weight that decides.

Every measurement in this field so far has been an argument for pooling: 0.88 against 2.23 and 1.15, every group estimated better, the weight derived three ways. All of it rests on one assumption, and this essay is about what happens when the assumption is false.

One group 6 population widths from the restSquared error for each group under partial pooling, with each group's own mean beside it. Seven of the eight are estimated better by pooling. The eighth, which was never from the population, is estimated 6.1 times worse — 13.9 against 2.3.group 11.09own mean 2.26group 21.14own mean 2.08group 31.16own mean 2.37group 41.18own mean 2.16group 51.13own mean 2.19group 61.17own mean 2.23group 71.06own mean 2.24the odd group13.86own mean 2.271,500 datasets, τ = 1, 4 observations per groupthe total is not the experience of the group
Fig. 1 Eight groups, one of which was never from the population the other seven came from. Seven are estimated better by pooling. The eighth is estimated six times worse.

The assumption, stated properly

The groups must be exchangeable: before seeing any data, there is nothing to distinguish one from another. Not identical — the whole subject is that they differ — but unlabelled, in the sense that relabelling the groups would not change what is believed about them.

That is a strong condition and it is almost never checked, partly because it is stated in a vocabulary that sounds technical. In practice it means something a reader can test against a dataset: there is no group that anyone would have singled out in advance.

A set of hospitals where one is the national referral centre is not exchangeable. A set of schools where one is selective is not exchangeable. A set of measurement sites where one uses a different instrument is not exchangeable. In each case the odd group is odd for a reason that was knowable before its data existed, and a model that treats all eight symmetrically is being told something false.

What it costs, counted

The failure is measured by placing one group’s true value a stated number of population widths from the centre and leaving everything else alone.

At six widths out, that group’s squared error under partial pooling is 13.9 against 2.3 for simply reporting its own mean — a factor of 6.1. The other seven groups are estimated better than their own means, exactly as before.

So the model has not broken. It has done precisely what it was asked: it observed a group whose mean was far from the others, judged from the population it was given that such a distance was mostly noise, and pulled it back. Every step was correct given the assumption, and the assumption was the problem.

One group 4 population widths from the rest. Squared error for each group under partial pooling, with each group's own mean beside it. Seven of the eight are estimated better by pooling. The eighth, which was never from the population, is estimated 2.9 times worse — 6.5 against 2.3.
Fig. 2 The same measurement at four population widths. The damage is already large, and every group’s error is still reported as a single unremarkable table.

Almost all of the 13.9 is bias

The two numbers split, and the split says what kind of failure this is.

A shrunk estimate’s squared error is the squared bias it takes on plus the variance it keeps: (Bd)2+(1B)2se2(Bd)^2 + (1-B)^2\,\text{se}^2, with B the weight put on the population and d the group’s distance from the centre. The own-mean figure of 2.3 is se2\text{se}^2 itself, so the standard error is about 1.5.

Solving 36B2+(1B)2×2.25=13.936B^2 + (1-B)^2 \times 2.25 = 13.9 at six widths gives B0.61B \approx 0.61, and with it the two pieces:

  • squared bias 13.56
  • variance 0.34

Ninety-eight per cent of the error is bias. The pooling did reduce the variance, and by a great deal — from 2.25 to 0.34, a factor of 6.6 — and it bought that reduction by taking on a bias more than seven times the variance it removed.

That is worth naming because it is the opposite of how a shrinkage estimator’s failures are usually imagined. Nothing here is noisy. The estimate for the odd group is more stable than its own mean, and it is stably in the wrong place.

The estimated spread has already softened the blow, and it is not enough

The implied weight of 0.61 is itself informative, because a known population spread would have given a different one.

At the settings these figures use — a within-group standard error near 1.5 and a population spread of one — the weight a known τ produces is 2.25/(2.25+1)=0.692.25/(2.25 + 1) = 0.69. The measured behaviour corresponds to 0.61, which is the weight an estimated spread of about 1.19 would give.

So the outlier has raised the estimated population spread by roughly a fifth, and that increase feeds back into its own shrinkage and softens it. The model is partially defending itself, which is the right thing for it to be doing and is a further reason the failure is hard to see: the diagnostic a reader might reach for — an implausibly large τ̂ — moves by twenty per cent, which on eight groups is well inside the noise of estimating τ at all.

Nineteen per cent of protection against a factor of six of damage is the honest summary. The mechanism that would notice the problem is the same mechanism the problem corrupts.

The failure lands on the interesting group

This is the part that makes the failure worth an essay rather than a caveat.

The group that violates exchangeability is, by construction, the group that is unlike the others. That is also, in almost every applied setting, the group somebody wanted to know about. Nobody commissions a study of eight hospitals to find out about the six that are average.

So the error is not distributed randomly across the table with an unlucky row. It is concentrated exactly on the row that motivated the analysis, and its direction is always the same: towards the middle, away from whatever made the group notable.

An analysis that pools will therefore tend to report that the outlying site is less outlying than the raw data says, the exceptional school less exceptional, the anomalous batch less anomalous. Where the group is genuinely exchangeable, that correction is right and is the whole value of the method. Where it is not, the correction is the error, and the two cases produce identical output.

The total goes wrong too

A defence is available for the case above: the total error across all eight groups might still be lower, and hierarchical estimates are usually justified on the total.

At six widths out it is not. The mean squared error per group under partial pooling is 2.72 against 2.22 for taking every group at its word — the shrinkage estimator loses outright, on its own criterion, because one group’s contribution dominates the sum.

That is a useful boundary to have measured. The advantage of pooling is not unconditional and does not survive a sufficiently non-exchangeable set. Somewhere between “all groups from one population” and “one group six widths out” the ranking reverses, and it reverses on the same criterion that established the advantage in the first place.

One group 8 population widths from the rest. Squared error for each group under partial pooling, with each group's own mean beside it. Seven of the eight are estimated better by pooling. The eighth, which was never from the population, is estimated 10.7 times worse — 24.2 against 2.3.
Fig. 3 Eight widths out, where the odd group’s own error dominates every summary the model produces.

How far out is far enough to be damaged

The failure is not a threshold effect, and sweeping the distance shows where the damage starts rather than leaving it as a warning about outliers in general.

At the population’s own scale — one width out, which is an entirely ordinary group — pooling estimates that group at 0.93 against 2.27 for its own mean. It is helped, substantially.

At two widths, the two are level: 2.05 against 2.27, near enough a tie. That is the crossing point, and it is worth noticing how unremarkable a group two population widths from the centre is. In a normal population about one group in twenty sits there legitimately, so a set of eight groups will contain one about a third of the time — and for that group, pooling and not pooling are worth the same.

At three widths pooling is 1.7 times worse for that group, at four 2.9 times, at six 6.1 times. The damage grows as the square of the distance, because the estimate is being pulled a fixed fraction of an increasing gap.

One group 1 population widths from the rest. Squared error for each group under partial pooling, with each group's own mean beside it. Seven of the eight are estimated better by pooling. The eighth, which was never from the population, is estimated 0.4 times worse — 0.9 against 2.3.
Fig. 4 One width out, where the group is unremarkable and pooling helps it as much as it helps the others. The whole failure is what happens to this picture as the distance grows.

The practical reading is that exchangeability does not have to fail badly to matter. A group that is different for a knowable reason is usually different by a lot more than two population widths, which is where the arithmetic has already turned against pooling it.

The estimated spread absorbs the outlier

There is a second failure underneath the first, and it is the reason the damage is not confined to the odd group.

τ is estimated from the spread between the group means. A group placed far from the others inflates that spread, so τ̂ comes out larger than the population it was supposed to describe — and a larger τ̂ means less pooling for every other group.

The seven well-behaved groups therefore borrow less than they should, because of a group they have nothing to do with. That effect is smaller than the direct damage to the odd group, and it is more insidious, because it degrades estimates that appear to have nothing wrong with them.

τ̂ across 2,000 datasets of 12 groups, true τ = 3. The population spread is not supplied to a hierarchical model — it is estimated from how far apart the group means are, after subtracting the noise that would separate them anyway. It averages 2.92 here against a true 3, and comes out exactly zero on 0% of datasets.
Fig. 5 The estimated population spread when the groups genuinely are far apart. The estimator cannot tell that case from one group that does not belong, because both produce widely separated means.

The estimator has no way to distinguish “a wide population” from “a narrow population plus one interloper”. Both produce a large observed spread, and the model’s response to both is the same.

Too few groups is the same failure wearing different clothes

The other way this field’s machinery becomes an assertion is by running it on too few groups.

τ̂ is computed from as many numbers as there are groups. With three or four, the spread between them is estimated so poorly that the weight is close to arbitrary — and, because of the downward bias described in the previous essay, arbitrary in a direction that pools too hard.

A model with three groups will often report τ̂ = 0 and produce three identical estimates. That is formally correct and it is not information; it says only that three numbers were not enough to establish that three groups differ. Presented as a finding — the sites do not differ — it is a statement about the study’s size dressed as a statement about the world, and it is the same error as reading a wide confidence interval as evidence of no effect.

The practical floor is around five to eight groups before τ̂ is worth much, and below it the honest report says which weight was used and that the data did not choose it.

What to do instead

Four responses, in increasing order of effort, and the first is the one most often skipped.

Say that the group is different. If the referral centre is known in advance to be different, take it out of the pooled set and report it separately. Exchangeability is an assumption about knowledge, so knowledge that breaks it belongs in the model rather than in a footnote.

Model the reason. Where the odd group differs for a stated reason — case mix, instrument, selectivity — that reason is a covariate. Pooling groups after adjusting for it restores exchangeability conditional on the covariate, which is the form the assumption usually takes in real work.

Use a heavier-tailed population. A normal population makes a distant group very unlikely, which is why it is pulled so hard. A t population with few degrees of freedom expects the occasional distant group, so it shrinks the well-behaved groups normally and leaves the distant one nearly alone. This is the robust-hierarchical route and it costs little.

Check the residual spread against the model. After fitting, the group estimates should be consistent with the population that was fitted. One group whose own mean sits far outside the fitted population is visible in exactly the way a high-leverage point is visible in a regression: a diagnostic exists, it is cheap, and the failure is silent without it.

What each estimator costs, 4 observations per group. Each point is 1,200 datasets of 8 groups. At τ = 0 the groups are identical and complete pooling is best at 0.29 against 2.23; at τ = 6 they are unrelated and it is worst at 31.5 against 2.23. Partial pooling is at or below both at every point.
Fig. 6 The comparison the whole field rests on, drawn from an exchangeable population. Every claim in it is conditional on the assumption this essay is about.

The diagnostic, and why it is not standard output

The check that would catch this is cheap and is not printed by default, which is worth stating precisely enough to be actionable.

After a hierarchical model is fitted, each group has an estimate and the population has a fitted spread τ̂. The question is whether each group’s own mean is consistent with having been drawn from that fitted population — that is, whether (group mean − population mean) is plausibly within a combination of τ̂ and that group’s own standard error.

A group four or more of those combined widths out is announcing that it does not belong. That is one subtraction and one division per group, and it needs no refitting.

What makes it unusual as a diagnostic is that the model’s own fit gives no hint. Residuals from a hierarchical model are computed against the shrunk estimates, and the shrunk estimate for the odd group has been pulled towards the middle — so the residual is small, the fit looks acceptable, and the diagnostic that would show the problem is the one comparing the unpooled mean to the fitted population.

Eight groups, τ = 2 against a within-group spread of 3. Each row is a group. The hollow circle is the group's own mean, the filled one is the estimate after pooling, and the small mark is the truth the data was generated from. The group of 3 moves 43% of the way to the population mean of -0.01; the group of 40 moves 5%.
Fig. 7 Eight groups at a wider population spread. A reader can see which groups moved and by how much; what no figure of the estimates alone shows is whether a group should have been in this population at all.

This is the same trap as a residual plot that looks fine because the outlier dragged the line to it: the quantity that would reveal the problem has already been adjusted by the problem.

The shape of the mistake, which is not specific to pooling

This failure has the same structure as several others on this site, and naming the structure is what makes it transferable.

A method that borrows strength across units is exactly a method that is wrong about the unit that does not belong. Regression borrows across observations, and one high-leverage point reverses a slope. A meta-analysis borrows across studies, and one study run differently distorts the pooled estimate. Kaplan–Meier borrows across subjects, and censoring that depends on the outcome biases the curve.

In every case the strength being borrowed is real and the method is better than the alternative, under an assumption about the units being comparable. In every case the failure is quiet, produces a well-formed answer, and lands hardest on the unit that motivated the analysis.

The general defence is the same too: state the assumption as a claim about the data rather than as boilerplate, and check it. “The groups are exchangeable” is a testable-sounding sentence that is almost never tested, and the reason is that the model does not require it to be — it will fit, converge, and print estimates either way.

Two cases that look identical and are not

It is worth putting the two situations side by side, because they produce the same table and call for opposite responses.

A wide population. Eight groups genuinely spread over a large range. τ̂ is large, B is small, everything stays near its own mean, and the model is doing the right thing — the estimates barely move because almost nothing separating the groups is noise.

A narrow population plus an interloper. Seven groups close together and one far away. τ̂ is inflated by the eighth, B is small for all of them, and the seven that should have borrowed heavily from each other do not, while the eighth is dragged towards a centre it never belonged to.

The fitted variance components are similar in the two cases. The group estimates are similar. What differs is the shape of the spread between the group means: wide-but-even in the first, tight with one distant point in the second, and that distinction is visible in a plot of the eight means and invisible in the two numbers a model summary reports.

So the recommendation is unusually concrete for something this general: plot the group means before reading the variance components. Eight points on a line take a moment, they answer a question no summary statistic answers, and the alternative is a model that cannot tell a population from a population plus a stranger.

What this field established

Four essays, one idea, and a boundary.

The idea is that the estimate for a group is a weighted average of the group’s own mean and the population’s, that the weight is se²/(se² + τ²), and that this is not a compromise between two positions but the general case of which both obvious answers are limits. Counted against a known truth, it costs 0.88 where they cost 2.23 and 1.15.

The population spread is not asserted. It is read off the distance between the group means with the measurement noise subtracted, it is worth σ²/τ² observations per group — nine, in these figures — and it comes out at exactly zero on more than half of all datasets where the groups are genuinely identical.

And the boundary is here: the whole of it holds when the groups are exchangeable, fails by a factor of six on a group that is not, and gives no indication of which case it is in. The methods that borrow strength are the methods that are wrong about whatever does not belong, and there is no version of them for which that is not true.

The reason to state the boundary as loudly as the result is that this field’s arguments are unusually persuasive. The estimator is better on three separate derivations, it beats both alternatives by a wide margin on a counted criterion, and it needs nothing supplied to it. That combination is exactly what makes the one assumption underneath it easy to stop mentioning — and the assumption is not a technical condition on a proof, it is a claim about the groups in the dataset, which somebody has to make and defend every time.

The four essays are therefore best read as one instruction. Pool, and say which groups were pooled and why they belong together. The first half is worth a factor of two or three in error; the second half is what stops the first from being worth a factor of six in the wrong direction.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

ExchangeabilityHierarchical modelModel misspecificationOutliersPartial poolingShrinkage