Groups that borrow

The fewest groups that can borrow

At three groups the estimator that shrinks towards its own data's mean returns the group means untouched, on every dataset, because its constant is J − 3. At two it expands instead of shrinking. And the number of groups at which partial pooling starts to be worth doing is five, or two, or never — it depends on how far apart the groups are.

Worth reading first: Eight groups, one population.

Eight groups, one population settled what borrowing is worth at eight groups: 0.88 of squared error against 2.23 for treating them as unrelated and 1.15 for treating them as one. Every figure in this field since has held the number of groups at eight.

Eight is not what most studies have. Three sites, four batches, five clinics — and the quantity every one of those studies has to estimate before it can borrow anything is a spread computed from three, four or five numbers.

What each estimator costs, against how many groups there areEach point is 4,000 datasets, with the population spread estimated from the data rather than supplied. Partial pooling first beats BOTH of the estimators it sits between at 5 groups; below that, complete pooling — which estimates nothing at all — is the better answer. Its own cost falls from 1.751 at 2 groups to 0.795 at 40.012311.5822.322.5833.584.325.32groupsmean squared error per groupborrowing starts paying at 5each group on its ownone number for allpartial poolingJames–Steinτ = 1, σ = 3, 4 observations a group, τ estimatedthe baseline is the better of the two ends
Fig. 1 The same three estimators against the number of groups rather than against the population spread, with τ estimated from the data as an analyst would have to. Partial pooling first beats both of the estimators it sits between at five groups. Below that, complete pooling — which estimates nothing at all — is the better answer.

The baseline that matters is not the obvious one

The figure’s headline depends entirely on which comparison is being made, and the usual one is too easy.

Partial pooling beats the groups’ own means at two groups: 1.7513 against 2.2641. That looks like a clean win and it is very nearly meaningless, because at two groups it wins by pooling almost completely — and complete pooling, at the same setting, costs 1.6111. The estimator that estimates a population spread from two numbers is beaten by the estimator that assumes the spread is zero.

The same holds at three groups (1.4906 against 1.4209) and at four (1.3562 against 1.3301). Five is where it turns: 1.2384 against complete pooling’s 1.2564, and from there the gap widens without reversing.

A method that is better than the worse of two baselines has not been shown to be worth using. The honest comparison is against the better of them, and against the better of them the answer at this setting is that borrowing needs five groups.

Five, or two, or never

The number is not a property of the method. Move the slider and it moves with it.

At a population spread of two — groups twice as far apart, everything else the same — the crossing is at two groups, because complete pooling has become a bad answer and partial pooling only has to beat the group means. At a spread of four it is four. At a spread of half, partial pooling does not beat both baselines at any size on the sweep, up to forty groups: complete pooling stays ahead at 0.2999 against 0.3235, because the groups genuinely are nearly identical and assuming so is right.

So the question how many groups does partial pooling need has no answer on its own. What it has is an answer given how far apart the groups are, which is the quantity nobody knows before the study — and which, at few groups, is exactly the quantity that cannot be estimated.

The four settings, with the crossing each produces:

population spread complete pooling at 40 groups partial pooling at 40 crossing
σ/6 0.2999 0.3235 none on the sweep
σ/3 1.0291 0.7954 5 groups
2σ/3 3.9459 1.4934 2 groups
4σ/3 15.6131 1.9831 4 groups

The last row is not a typo and it is the one that shows the crossing is not monotone in the spread. At a spread of four, complete pooling is hopeless — fifteen times the cost of partial pooling at forty groups — and yet partial pooling still loses at two and three groups, this time to the group means rather than to complete pooling. With the groups that far apart, a τ̂ that collapses to zero is a catastrophe rather than a mild over-pooling, and at two or three groups it collapses often enough to outweigh everything borrowing buys on the datasets where it does not.

So there are two ways to be below the crossing and they are on opposite sides of the subject: too few groups to tell a small spread from none, and too few groups to be sure a large spread will not be mistaken for none. The second is the more dangerous, because it is the case where an analyst is most confident that pooling is appropriate.

That is worth stating as the circle it is. Borrowing is worth doing when τ is large relative to the standard errors. Whether τ is large relative to the standard errors is what τ̂ is for. And τ̂ from three groups is the estimator with an atom at zero, which at three groups reports zero on well over half of datasets.

What each estimator costs, against how many groups there are. Each point is 4,000 datasets, with the population spread estimated from the data rather than supplied. Partial pooling first beats BOTH of the estimators it sits between at 2 groups; below that, complete pooling — which estimates nothing at all — is the better answer. Its own cost falls from 2.069 at 2 groups to 1.493 at 40.
Fig. 2 The same sweep with the groups twice as far apart. Complete pooling is now a poor answer at every size, so partial pooling only has to beat the group means and the crossing moves to two groups. Nothing about the estimator changed.
What each estimator costs, against how many groups there are. Each point is 4,000 datasets, with the population spread estimated from the data rather than supplied. Partial pooling first beats BOTH of the estimators it sits between at no size on this sweep; below that, complete pooling — which estimates nothing at all — is the better answer. Its own cost falls from 1.585 at 2 groups to 0.323 at 40.
Fig. 3 And with the groups genuinely alike. Complete pooling costs 0.2999 at forty groups against partial pooling’s 0.3235, and the two curves do not cross anywhere on the sweep. Where the population spread is small, assuming it is zero is better than estimating it — at every number of groups this figure reaches.

A degree of freedom, spent on the middle

Underneath the simulated crossing there is an exact result that needs no datasets at all, and it belongs to the other estimator in this field.

James–Stein shrinks each group’s mean towards a centre by a factor computed from how spread out the means are:

c=1(J3)σ2Sc = 1 - \frac{(J-3)\,\sigma^{2}}{S}

when the centre is the mean of the same observations, and with J2J-2 in place of J3J-3 when the centre is a point fixed in advance. The difference is one, and it is a degree of freedom: locating the centre from the data costs a group.

The shrinkage factor, and the group it spends locating the centre. The factor is one minus (J − 3) times the within-group variance over the sum of squares when the centre is the data's own mean, and one minus (J − 2) times the same ratio when the centre is given. At three groups the first constant is zero, so the factor is exactly 1 and nothing moves; at two it is negative and the estimator expands on 100% of datasets. Each point of the estimated-centre curve sits close to the given-centre curve one group to its left.
Fig. 4 The factor itself, averaged over three thousand datasets at each size. One means no borrowing and zero means complete pooling. At three groups the estimated-centre factor is exactly 1.0000 — the constant is zero, so the correction is zero, on every dataset and not on average.

At three groups J3=0J-3 = 0, so c=1c = 1 identically, and the estimator returns the group means unchanged. That is not an approximation or a small-sample degradation. It is an algebraic identity, checked here to 101210^{-12} on forty datasets, and it says that an estimator shrinking towards its own data’s mean cannot shrink at all with three groups.

The two curves say the same thing a second way. At four groups the estimated-centre factor averages 0.5950; at three groups the given-centre factor averages 0.5949. The estimated-centre estimator at JJ groups behaves almost exactly like the given-centre one at J1J-1, which is the degree of freedom made visible. The agreement is close rather than exact — at five against four it is 0.5015 against 0.5037, about half a per cent apart — because the two sums of squares are taken about different points.

τ̂ across 2,000 datasets of 3 groups, true τ = 1. The population spread is not supplied to a hierarchical model — it is estimated from how far apart the group means are, after subtracting the noise that would separate them anyway. It averages 0.80 here against a true 1, and comes out exactly zero on 45% of datasets.
Fig. 5 Why the other estimator has its own difficulty at three groups, from the other direction: the spread estimated from three group means comes out exactly zero on 50.8% of datasets that have a real spread of one in them. At four groups it is 45.0%, at five 40.7%, at eight 32.7%.

Those two facts are the same fact seen twice. An estimator shrinking towards its own mean has a constant that vanishes at three groups; an estimator estimating a population spread from three group means gets zero half the time. Both are saying that three numbers do not determine how much three numbers vary, and both say it by returning a value at the edge of their own range rather than by failing.

Two groups, where the clamp guards the wrong end

At two groups the constant is 1-1, and a negative constant makes cc larger than one.

The positive part — the max(0,)\max(0, \cdot) that every modern statement of the estimator carries — exists to stop cc going below zero, which would shrink each group past the centre and out the other side. Nothing in the formula stops cc going above one, and above one it expands: each group’s deviation from the mean is multiplied up rather than damped.

It happens on 100% of two-group datasets, and the average factor is 413,931, because the sum of squares in the denominator is a single squared difference that is sometimes very close to zero. The resulting mean squared error is 349,375 against the group means’ 2.2641.

This shape keeps recurring, and it is worth naming rather than filing as a curiosity. The clamp is a guard written against the failure somebody had in mind, the other failure was outside the range the formula was derived for, and the output is a perfectly well-formed number several orders of magnitude wrong. Nothing throws. Nothing warns. The same structure produced a tick function returning an empty array and an interval on a prior that covers nothing.

The practical form is one line: the estimator is not defined below four groups, and the arithmetic does not say so.

Why partial pooling does not have the same cliff

The empirical-Bayes estimator in every other figure here — precision-weighted centre, weight se2/(se2+τ^2)se^{2}/(se^{2}+\hat\tau^{2}) — has no such constant and no such cliff. At two groups it returns 1.7513 rather than 349,375.

The reason is the clamp, in the other direction. τ^2\hat\tau^{2} is max(0,sbetween2se2)\max(0, s^{2}_{\text{between}} - \overline{se^{2}}), and at two or three groups the subtraction lands below zero most of the time, so τ^=0\hat\tau = 0, so the weight is one, so the estimate is the grand mean. Partial pooling’s failure mode at few groups is complete pooling, which is a sensible answer, and James–Stein’s is an unbounded expansion, which is not.

That is a real difference between two estimators usually described as the same estimator from two directions, and it shows up only where the number of groups is small. It is also why the crossing in the first figure is at five for partial pooling and at twelve for James–Stein: the second spends its first several groups recovering from being badly behaved rather than starting from a reasonable baseline.

What each estimator costs, 4 observations per group. Each point is 1,200 datasets of 8 groups. At τ = 0 the groups are identical and complete pooling is best at 0.29 against 2.23; at τ = 6 they are unrelated and it is worst at 31.5 against 2.23. Partial pooling is at or below both at every point.
Fig. 6 The field’s original comparison for contrast, at eight groups throughout: against the population spread rather than against the number of groups, all four estimators are well behaved and partial pooling is never the worse of the two ends. Everything in this essay is invisible on that axis.
The weight on the population, σ = 3. Each curve is one population spread τ. A group's estimate moves B = se²/(se² + τ²) of the way to the population mean, where se = σ/√n is what the group's own mean does not know. At τ = 1 a group of 9 observations sits halfway.
Fig. 7 The weight the whole field runs on, for reference: how far a group moves, against how much data it has. Nothing in this curve mentions the number of groups, which is the reason none of this essay’s difficulty is visible from it — the weight is a statement about one group and every problem here is a statement about how many there are.

That is worth pausing on, because the weight is how partial pooling is usually explained and it is complete as far as it goes. Given τ, a group of four moves 69% of the way to the middle and a group of forty moves 18%, and the number of groups never enters. Every difficulty in this essay lives in the given τ, and the picture that explains the method is a picture in which that quantity has already been supplied.

What a study with four groups should do

The arithmetic above supports three statements and refuses a fourth, and separating them is most of what a reader needs.

Below four groups, do not use an estimator that shrinks towards its own mean. That is exact. The James–Stein form returns the data unchanged at three and misbehaves at two, and neither is a small-sample caveat.

At four or five groups, compare against complete pooling rather than against the group means. The counted numbers say partial pooling loses that comparison at four and wins it at five when the spread is σ/3, and the crossing moves with the spread — so the comparison has to be made rather than looked up.

Where the spread genuinely is small, complete pooling is not a straw man. At τ = σ/6 it wins at every size up to forty groups. An analysis that reports one number for every group because the evidence for any difference is weak has not failed; it has arrived at the answer the data supports, which is what the collapse of τ̂ to zero is saying when it happens.

What the arithmetic refuses is the fourth statement, the one that would be most useful: use partial pooling when there are at least N groups. There is no such N. The crossing is a function of the ratio of the population spread to the standard errors, that ratio is the thing being estimated, and at the sizes where the rule would matter it is estimated from three or four numbers.

How often "there is no spread between the groups" is reported about data that has one. Every dataset here was generated with a real population spread of 1. The moment estimator is the difference between the observed spread and what noise alone would produce, clamped at zero, and the difference comes out negative often: at eight groups it reports exactly zero on 32.6% of datasets, which is an instruction to pool completely and give all eight groups the same estimate. The rate falls to 4.2% at 48 groups.
Fig. 8 The same difficulty stated as a rate rather than as a risk: how often the estimated population spread comes out exactly zero, against how many groups there are. It takes forty-eight groups to get the rate below five per cent. Every number in this essay is downstream of that curve.

The one number that does generalise

There is a quantity that does not move with the spread, and it is the one to carry away.

The cost of estimating τ rather than being told it is the gap between the partial-pooling curve and the known-centre curve, and at four groups it is 1.3562 against 1.3022 — about 4%. At eight it is 1.0679 against 0.9764, about 9%; at forty, 0.7954 against 0.7584, about 5%.

That gap is small and nearly flat, which is the reassuring half of this essay. Estimating the population spread is not expensive. What is expensive is that the estimate of it is, at few groups, usually zero — and a weight of one is not a small error in a weight, it is a different analysis.

So the difficulty at four groups is not an accumulation of noise that more careful arithmetic would reduce. It is that the estimator has two regimes, and at four groups it spends most of its time in the one where it has stopped borrowing and started insisting.

Two ways of running out of groups

Putting the two failures side by side says what each is and why they are not the same.

The constant runs out. J3J-3 is a count of degrees of freedom and at three groups there are none left: two have gone on the spread of the means and one on the centre. This is exact, it is the same on every dataset, and it produces an estimator that does nothing at three and something absurd at two. No amount of data inside the groups repairs it, because the constant counts groups.

The estimate collapses. τ^\hat\tau is a difference of two positive quantities and the difference is negative more than half the time at three groups. This is a rate rather than an identity — it varies by dataset — and it produces an estimator that pools completely, which is a defensible answer arrived at for an indefensible reason. More data inside the groups does repair it, by shrinking the se2\overline{se^{2}} being subtracted.

The second is the one that governs practice, because the empirical-Bayes form is what gets fitted. And it has a consequence worth stating for anyone reading a report from a four-group study: the finding that the groups did not differ is, about half the time, a fact about there having been four of them. The estimator was asked a question its input cannot answer and returned the edge of its range, and the report of that is a row of identical estimates with intervals around them.

The third thing this sits beside, which neither of the two is, is the outlier case. There the model is wrong about one group and every quantity in the fit is computed correctly. Here the model is right and the quantities cannot be computed. Both produce a plausible-looking output and they are different faults with different repairs, which is why this field’s ledger keeps them apart.

What is claimed here, and what is not

The claim is how partial pooling and James–Stein behave as the number of groups falls: that the estimated-centre shrinkage factor is exactly one at three groups and above one at two, both as identities rather than as measurements; that the number of groups at which partial pooling beats both of its baselines is five at one population spread, two at another and does not exist at a third; and that the cost of estimating the spread rather than being told it is a few per cent while the cost of the estimate collapsing is a change of estimator.

The simulated numbers are four thousand datasets a point with the spread estimated from the data. The two algebraic results are exact and are gated at 101210^{-12}.

What stays out: restricted maximum likelihood, which has a different small-sample bias and the same atom at zero and would need its own counting; the fully Bayesian answer, which has no atom and is the field next door; unequal group sizes, which are the default in every other figure in this field and are held equal here so that the number of groups is the only thing changing; and the case of two groups treated as a comparison rather than as a hierarchy, which is a t-test and a different question.

What this does not say about eight groups

Nothing here weakens the earlier results, and it is worth being explicit because a finding about small samples can read as a retraction.

At eight groups partial pooling costs 1.0679 against the group means’ 2.2502 and complete pooling’s 1.1645, with the spread estimated from the data. That is the number the essay that first pooled eight groups reported, it is a win against both baselines, and it holds. At twelve, twenty and forty groups the margin grows: 0.9622, 0.8724, 0.7954 against complete pooling’s 1.1120, 1.0620 and 1.0291.

What changes here is the shape of the claim rather than its sign. Partial pooling beats both extremes is true at eight groups and at a population spread of σ/3, and it is a statement about a region rather than a theorem. The region is large and it does not include four groups, and there is no way to tell from inside a four-group study which side of it the study is on.

Still open: who the borrowing is for

Every comparison here is a total — squared error summed over all the groups and divided by how many there are. That is the right quantity for ranking estimators and it is nobody’s experience, because nobody runs all eight groups.

When borrowing goes wrong took one version of that apart: the group that is genuinely not from the population is estimated worse. What has never been asked is the ordinary case, where every group is from the population and the total improves by a third. Where does that third come from? The weight is a function of how well a group is measured and of nothing else, so the answer is available from the arithmetic before any counting — and it is more lopsided than the arithmetic suggests. That is where the borrowing actually goes.

The check, and the refusal

Two claims are gated at machine precision because they are identities rather than measurements: at three groups the estimated-centre estimator returns every group mean unchanged to 101210^{-12}, checked on forty datasets, and its risk at three groups therefore equals no-pooling’s risk exactly. Beside them, the given-centre version is required to move at three groups, because an identity demonstrated only for the estimator it is about is an identity that might be an error in the arithmetic behind both.

The refusal is the claim a reader would most like to be true: borrowing always helps, rejected. At three groups the estimator that shrinks towards its own mean is required to be no better than not borrowing at all, and at four groups with the population spread set wide it is required to buy less than a tenth of the error. A check that found borrowing profitable at every size and every spread would be a check measuring its own baseline, which is what a refusal that never rejects always turns out to be.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Degrees of freedomEmpirical BayesHierarchical modelJames–SteinMean squared errorMethod of momentsPartial poolingSample sizeShrinkageVariance components