Groups that borrow

Where the borrowing goes

Pooling cuts the total squared error across eight groups by 56%. Two of the eight take 61% of that reduction, the four best-measured groups share 11% between them, and the largest group gets 1.5% of what the smallest does. The headline is a fact about the groups nobody was asking about.

Worth reading first: Eight groups, one population.

A pooled analysis of eight groups reduces total squared error from 10.089 to 4.448 — a reduction of 56% — and that sentence is how the result gets reported. It is true, it is the number the essay that first pooled eight groups counted, and it describes an experience nobody has.

Nobody runs eight groups. Somebody runs one of them.

What each group gains from being pooled, τ = 1Eight groups whose sizes span a factor of 13.3. The smallest gains 2.076 of squared error, which is 36.8% of the total reduction; the largest gains 0.030, which is 0.5%. 2 of the eight account for half of everything pooling buys.n = 32.0762.97 → 0.89n = 41.3582.16 → 0.80n = 51.0861.81 → 0.73n = 80.5101.13 → 0.62n = 100.3500.89 → 0.54n = 160.1560.55 → 0.40n = 250.0760.36 → 0.28n = 400.0300.23 → 0.204,000 datasets, τ = 1, σ = 32 groups hold half the gain
Fig. 1 What each of the eight gains, in squared error, with the groups ordered by how much data each has. The smallest gains 2.076 and the largest gains 0.030 — a factor of sixty-nine between the two ends of a study everybody would describe as one study.

The concentration, as shares

The amounts are hard to hold, so the same measurement as shares of the total reduction:

observations in the group weight on the population own error pooled error share of the gain
3 75% 2.967 0.892 36.8%
4 69% 2.157 0.800 24.1%
5 64% 1.815 0.729 19.3%
8 53% 1.125 0.615 9.0%
10 47% 0.888 0.538 6.2%
16 36% 0.552 0.396 2.8%
25 26% 0.357 0.281 1.3%
40 18% 0.228 0.197 0.5%

Two groups of the eight account for 60.9% of everything pooling buys. The four best-measured groups share 10.8% between them. The largest group receives 1.5% of what the smallest receives.

Each group's share of the total gain, τ = 1. Eight groups whose sizes span a factor of 13.3. The smallest gains 2.076 of squared error, which is 36.8% of the total reduction; the largest gains 0.030, which is 0.5%. 2 of the eight account for half of everything pooling buys.
Fig. 2 The same eight numbers as shares, with each group’s gain relative to its own error printed beside it. The group of three has 70% of its squared error removed; the group of forty has 13%.

Why it has to be this way

The concentration is not an accident of these eight sizes. It follows from the weight in two steps, and both are worth doing because they say something slightly different.

A group with more data has less error to lose. The group’s own mean has squared error σ2/n\sigma^{2}/n, so the group of forty starts at 0.228 and the group of three starts at 2.967 — a factor of thirteen before anything is pooled. Even a procedure that removed the same proportion from every group would deliver thirteen times as much to the smallest.

And a group with more data moves less. The shrinkage weight is se2/(se2+τ2)se^{2}/(se^{2}+\tau^{2}), which is 75% for the group of three and 18% for the group of forty. So the proportions are not equal either: the smallest group has 70% of its error removed and the largest 13%.

The two effects multiply. Thirteen times the error to start with, and five times the share of it removed, gives the factor of sixty-nine.

The weight on the population, σ = 3. Each curve is one population spread τ. A group's estimate moves B = se²/(se² + τ²) of the way to the population mean, where se = σ/√n is what the group's own mean does not know. At τ = 1 a group of 9 observations sits halfway.
Fig. 3 The second of the two effects on its own. Nothing about a group’s value decides how far it moves — only how well it is measured — which is the property that makes partial pooling a statement about information. It is also the property that decides who the gain goes to.

This is worth separating from a criticism, because it is not one. The gain is concentrated in the poorly-measured groups because that is where the error was. A procedure that spread its benefit evenly across groups of very different sizes would be doing something indefensible: moving well-measured groups as far as badly-measured ones, which is exactly the thing the weight exists to prevent.

The complaint is not about where the gain goes. It is about the sentence that reports the total.

The factor of sixty-nine is not the whole of it

One more multiplication belongs in the account, and leaving it out would understate the concentration rather than overstate it.

The eight groups here differ in size by a factor of 13.3, which is a large spread and a realistic one: clinics recruit at different rates, batches come out at different sizes, schools have different numbers of pupils. Nothing in a study forces the range to be narrow, and the studies where pooling is most often reached for — a handful of sites, unequal enrolment — are exactly the ones where it is wide.

Squared error before pooling goes as 1/n1/n, so the ratio of the ends is 13.3 before anything else. The weight goes as se2/(se2+τ2)se^{2}/(se^{2}+\tau^{2}), which compresses towards one as sese grows and towards zero as it shrinks, so the ratio of the fractions removed is a further 5.3. Their product is 70, against the counted 69 — the small gap is the simulation’s own error on eight ratios — and both factors widen when the study’s sizes are more unequal.

Push the range to a factor of a hundred and the two effects give a ratio in the hundreds. Close it to a factor of two and they nearly vanish. The concentration is therefore a fact about the design of the study, transmitted through an analysis that did nothing wrong — which is why the repair in the last section of this essay is a design repair and not an analysis one.

Two readers, one analysis

Consider the eight groups as eight clinics in a multi-site study, and two people reading the report.

The methodologist sees total squared error fall 56% and correctly concludes that the hierarchical model is the right analysis. The total is the right quantity for that question, and the answer does not depend on which clinic anyone is at.

The clinic with forty patients sees its own estimate move by 18% of the way to the middle, its own squared error fall from 0.228 to 0.197, and gains 0.5% of the reduction the report is about. It has been given a slightly better estimate and a much less legible one — its published number is now partly somebody else’s data — in exchange for something it would need four significant figures to notice.

Both readings are correct. The second is not an argument against the model; it is an argument about what the report should say, and the missing sentence is who the improvement is for.

Eight groups, τ = 1 against a within-group spread of 3. Each row is a group. The hollow circle is the group's own mean, the filled one is the estimate after pooling, and the small mark is the truth the data was generated from. The group of 3 moves 75% of the way to the population mean of 0.10; the group of 40 moves 18%.
Fig. 4 The same eight groups as the movement each undergoes. The arrows are the gain: long for the groups at the top of the figure, barely visible for the groups at the bottom, and the total the report quotes is mostly the top two arrows.

The concentration tightens as the groups separate

The obvious next question is whether this is a feature of one setting, and it is not — it gets more extreme in the direction that matters.

At a population spread of four the total gain falls to 0.904, a tenth of what it is at a spread of half, because groups that are genuinely far apart have little to borrow. But the share taken by the smallest group rises, from 33.0% to 43.5%, and the share taken by the largest falls from 1.0% to 0.2%.

What each group gains from being pooled, τ = 4. Eight groups whose sizes span a factor of 13.3. The smallest gains 0.393 of squared error, which is 43.5% of the total reduction; the largest gains 0.001, which is 0.2%. 2 of the eight account for half of everything pooling buys.
Fig. 5 The groups four population widths apart. The total reduction is 0.904 rather than 5.641, the group of forty gains 0.001, and the group of three still takes 43.5% of what there is. Pooling has become nearly pointless for six of the eight and is still doing something for two.

So the two regimes a reader might hope for do not exist. Where the groups are alike, pooling helps a great deal and helps the small groups most. Where they are different, pooling helps very little and still helps the small groups most. There is no setting on the slider where the benefit is evenly distributed, and the share taken by the smallest two groups runs from 56.0% to 68.5% across the whole range.

The same budget, spread evenly

The concentration invites a design question, and the answer is not the one the concentration suggests.

The eight groups here hold 111 observations between them, distributed from three to forty. Spread evenly — fourteen each — the same budget gives a total squared error before pooling of 5.141 rather than 10.089, because squared error is convex in 1/n1/n and the group of three is carrying most of the total on its own.

Pooling then removes 1.796 rather than 5.641, and removes it evenly: every group’s share of the gain is between 11.7% and 13.3%, and every group has between 33% and 38% of its own error removed. Four of the eight account for half, which is as even as eight groups can be.

So the equal design is better in total — 3.345 against 4.448 after pooling — and pooling does much less for it. Both halves of that are worth carrying.

Equalising the groups is worth more than pooling them. Going from the unequal design to the equal one saves 4.948 of squared error before any pooling at all, which is nearly as much as pooling saves on the unequal design and is available to anybody who controls the allocation.

And it removes most of what pooling had to offer. A study whose groups are already the same size has less for a hierarchical model to do, because the thing being exploited — that some groups are badly measured and can be told about by the others — has been designed away.

That is the honest shape of the trade and it inverts the usual order of operations. Partial pooling is a repair applied after the data exists; the allocation is a decision taken before it, and how the units are split between the groups is where most of this particular gain actually lives.

What this does to a comparison between groups

The concentration has a consequence for the question multi-site studies are usually asked, and it is not a consequence about error at all.

Pooling moves the small groups a long way and the large groups hardly at all. So a comparison between a small group and a large one is a comparison between a number that has been pulled most of the way to the middle and a number that has barely moved — and the difference between them is systematically understated, by an amount that depends on which two groups are being compared.

Between the group of three and the group of forty, the weights are 75% and 18%. If the two truths differ by dd, the expected difference between the pooled estimates is 0.25d0.25d against 0.82d0.82d — the comparison is attenuated to about 30% of the real gap. Between the group of twenty-five and the group of forty it is attenuated to 88%.

The attenuation is different for every pair. There is no single correction and no single caveat; the ranking of the groups is distorted in a pattern that depends on their sizes and not on their values. That is the mechanism behind the warning every hierarchical modelling text gives about comparing shrunk estimates, stated as a number rather than as advice.

One group 4 population widths from the rest. Squared error for each group under partial pooling, with each group's own mean beside it. Seven of the eight are estimated better by pooling. The eighth, which was never from the population, is estimated 2.9 times worse — 6.5 against 2.3.
Fig. 6 The sharpest version, which the essay on borrowing that goes wrong counts in full: a group genuinely four population widths from the rest is estimated 6.1 times worse by pooling than by its own mean. The total is barely affected, because that group is one of eight.

The fifty-six per cent, taken apart

The reduction the report quotes is a ratio of two sums, and both sums are dominated by the same group.

Before pooling, the group of three contributes 2.967 of the total 10.089 — 29% of the study’s whole squared error sits in the group with 3% of its observations. The three smallest groups together hold 69% of it.

After pooling, the eight errors are much more alike: 0.892, 0.800, 0.729, 0.615, 0.538, 0.396, 0.281, 0.197. The largest is four and a half times the smallest, against thirteen times before.

So what pooling does to the total is mostly what it does to the three worst-measured groups, and what it does to the shape of the eight errors is to compress them. Both are real and neither is “every estimate improved by 56%”.

The compression is worth a line of its own, because it is the effect a reader of the eight numbers would actually notice. Pooled estimates are more alike than unpooled ones in their accuracy as well as in their values, which is a genuine benefit for anything downstream that treats the eight as interchangeable — and a genuine hazard for anything that reads the eight as eight independent readings, since they are no longer independent and no longer equally informative about their own groups.

What the total is good for, and what it is not

Three statements, and they have to be kept apart.

As a comparison of estimators, the total is the right quantity. Partial pooling against the group means against complete pooling is a question about procedures, procedures are judged by risk, and risk is an average over the thing the procedure is applied to. Nothing in this essay changes the ranking the first pooling of eight groups established.

As a description of what the analysis did, the total is misleading by construction. “Error fell 56%” invites the reading that every estimate improved by something like a half, and the truth is that one estimate improved by 70% and three improved by less than 30%. The distribution is not a detail around the average; the average is a number no group is near.

As an argument to a group about why its estimate was moved, the total is irrelevant. A clinic whose number changed deserves the reason its own number changed, and that reason is a weight the clinic can be told: this estimate was moved 18% of the way to the middle because this clinic’s standard error is one number and the spread between clinics is another. The weight is per-group, it is interpretable, and it is already computed.

That last point is the practical one. The quantity to report beside a pooled estimate is its own weight, not the study’s total improvement, and the weight is the one number in the whole apparatus that belongs to the group rather than to the study.

The cross-field version

The pattern is not about group sizes. It is about the weight depending on something the group has more or less of, and a field over there is a case where the something is not a count at all.

Six groups of 10, each fitting its own slope, then borrowing. Each faint line is one group's own least-squares slope through its own centre; each solid line is that slope after pooling towards the population slope of 0.79. Every group has the same 10 observations. The group whose x values span 0.4 has a slope standard error of 1.86 and moves 91% of the way in; the group spanning 2.0 has a standard error of 0.37 and moves 28%.
Fig. 7 Ten groups with ten observations each, pooling a slope rather than a mean. Every group has the same amount of data and they borrow anything from 28% to 91%, decided entirely by where within each group the ten observations sit. The gain is concentrated exactly as it is here, and the axis it is concentrated along is invisible from the group sizes.

There, pooling a slope makes the point that sample size is a proxy for the thing that matters rather than the thing itself. The same measurement applies: the groups whose xx values are bunched have the largest standard errors, borrow the most, and take most of the gain — and a reader of that study cannot even sort the groups by size to find out who benefited.

So the general statement is one line, and it covers both. The gain from pooling is distributed in proportion to how badly each unit was measured, squared, and how badly each unit was measured is usually not in the report.

The number a group should be given

Everything above points at one practical change, and it costs nothing to make.

A pooled analysis already computes, for each group, the weight B=se2/(se2+τ^2)B = se^{2}/(se^{2}+\hat\tau^{2}). It is the fraction of the way that group’s estimate was moved towards the middle, it is a number between zero and one, and it is reported by essentially no software because it is an intermediate quantity rather than an output.

It is also the only number in the analysis that belongs to the group. The total improvement belongs to the study; the population spread belongs to the study; the group’s own estimate has been changed by the study. The weight is the answer to why did my number move, and by how much, and it answers it without any reference to the other groups’ values.

Printing it beside each estimate would make three things visible at once that are currently not. Which groups the analysis is doing work on. How much of each published number is that group’s own data. And — since the weight is a monotone function of the standard error — which groups are the ones the study measured badly, which is a statement about the design rather than about the analysis and is usually the more actionable of the two.

What is claimed here, and what is not

The claim is the distribution of the benefit of partial pooling across eight groups of unequal size: that two of the eight take 60.9% of it, that the four best-measured share 10.8%, that the ratio between the ends is sixty-nine, that the concentration tightens rather than loosens as the population spread grows, and that the two mechanisms producing it — less error to lose, and less movement — multiply rather than add.

Every number is four thousand datasets with the population spread supplied rather than estimated, so nothing here is contaminated by the collapse of τ̂; the concentration is a property of the weight and not of the difficulty estimating it.

What stays out: the attenuation of between-group comparisons, which is computed here from the weights in closed form and not counted, and which deserves the coverage measurement this site would normally demand before quoting it; and loss functions other than squared error, under which the distribution would differ and the direction would not. The equal-allocation comparison is counted here at one budget and one population spread; whether the ordering survives across the range is a measurement the allocation field is better placed to make, and it is the one number in this essay quoted from a single configuration.

Still open: the population the groups came from

Every figure here is drawn from a population that is a single normal, which is what the model assumes. The weight is right, the gain goes where the arithmetic says, and the estimates are pulled towards a centre that is genuinely the middle of the population.

What has never been asked is what happens when the population is the right size and the wrong shape — when the group effects come from two clusters rather than one bell, with the same total spread. The moment estimator recovers the same τ̂, so every weight in the analysis is unchanged and nothing in the output moves. What changes is where the estimates are pulled to: 46% of them land in a region holding 6.6% of the truths. That is what happens when the population is two populations.

The check, and the refusal

Three claims are gated. That the gain falls with every observation a group already has, which is the mechanism stated as an ordering rather than as a rate. That a third of the groups or fewer account for half of the total, which is the concentration as a threshold that a more even distribution would fail. And that the best-measured group gains under a tenth of what the worst does, which is the factor of sixty-nine stated loosely enough to survive the slider and tightly enough to have content.

The refusal is the reading the headline invites: the gain is evenly spread, rejected. One group is required to take more than two and a half times an even share. If that ever passed, the essay’s whole argument would be about a difference too small to matter — and the check would be measuring a study whose groups were all the same size, which is the configuration every textbook picture uses and the reason this field draws unequal ones instead.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Empirical BayesHierarchical modelMean squared errorPartial poolingSample sizeShrinkageStandard errorStudy designVariance components