The weight that decides
Worth reading first: Eight groups, one population.
The estimate that beats both obvious answers is a weighted average, and the weight is one line:
B = se² / (se² + τ²)
This essay is about that line. It arrives from three arguments that have almost nothing in common, which is the strongest evidence available that it is not an artefact of any of them.
Read as a ratio
The numerator is what the group’s own mean does not know: se², the variance of an average of that group’s observations. The denominator is everything that is not known — the group’s noise plus the real spread between groups.
So B is the share of the apparent difference between this group and the middle that is attributable to noise, and the estimate moves that share of the way towards the middle. Nothing else in the expression: not the group’s value, not its rank, not whether it is above or below average.
Two limits are worth taking, because they are the two obvious answers from the previous essay reappearing as special cases.
When τ → 0 the groups are identical, B → 1, and the estimate is the population mean for every group: complete pooling. When τ → ∞ the groups have nothing to do with each other, B → 0, and the estimate is the group’s own mean: no pooling. Partial pooling is not a third method sitting between two others. It is the general case, and both obvious answers are what it does at the ends of a range.
The landmark on the curve is where se = τ, which puts B at exactly a half. At a within-group spread of 3 and a population spread of 1 that is a group of nine observations, and the figure marks it because it is the only point on the curve a reader can locate without arithmetic.
The half-point, as a group size
Every curve crossing a half where the group’s standard error equals the population spread is a statement about n, and writing it that way makes the weight a single hyperbola.
With ,
So the whole curve has one scale in it. A group of exactly observations is half its own and half the population’s; at and that is nine observations — the same nine as the population is worth nine observations to each group, which is the same arithmetic said from the other side.
And the hyperbola is slow. A group needs observations to be 90% pooled and to be 90% its own, so the middle four fifths of the weight spans an eighty-one-fold range of group size.
That is worth carrying into any table of groups. Unless the group sizes vary by two orders of magnitude, they will all sit in the middle of the curve together, and the weights will look far more alike than the sizes do.
The second derivation: it is a posterior mean
Suppose the group’s true value has a normal distribution centred at the population mean with spread τ, and the group’s observed mean is that value plus normal noise of size se. That is the same model stated in Bayesian language: the population distribution is a prior, and the group’s own mean is the likelihood.
The posterior for the group’s value is then normal, and its mean is
B × (prior mean) + (1 − B) × (observed mean) with B = se²/(se² + τ²)
which is the same expression, arrived at with no mention of squared error, estimators, or how well anything performs across repeated datasets.
This site does not accept an algebraic identity as evidence that an implementation is right, so the identity is checked the way everything else here is checked: by computing the posterior mean a second way, with no shared arithmetic. The second route multiplies the prior density by the likelihood on a grid of twenty thousand points and integrates numerically. The two agree to about fifteen decimal places across every case tested, which is the check’s real content — a formula and a numerical integral that disagree would mean one of them was not the model.
The consequence for a reader is that “shrinkage” and “using a prior” are the same operation, and the choice between the vocabularies is a choice of which part to argue about. The Bayesian account makes the assumption explicit and invites the objection that the prior was invented. The hierarchical account makes the same assumption and answers the objection in advance: the prior is not invented, it is estimated from the other groups.
The third derivation, which mentions no population at all
The third argument is the strangest and it is worth stating carefully, because the usual telling makes it sound like a paradox rather than a fact.
Take eight quantities. They are fixed numbers, not draws from anything. There is no population, no prior, no hierarchy, and no sense in which the eight are related — one may be a mortality rate, another a rainfall total, another the batting average of a stranger. Each is observed once, with independent normal noise of known size.
The obvious estimate of each quantity is its own observation. It is unbiased, it is the maximum likelihood estimate, and for any single one of the eight it cannot be improved on.
Stein showed in 1956 that for the eight together it can. Shrinking every observation towards their common average, by a factor computed from how spread out the observations are, has a lower total squared error than the observations themselves — always, for every configuration of the eight true values, whatever they are.
The shrinkage factor is
1 − (k − 3)σ² / Σ(observed − average)²
clamped below at zero, and its shape is exactly the shape of B: when the observations are bunched, most of what separates them is noise and the factor is small, so everything moves to the middle; when they are far apart, it is near one and they stay.
Measured here on eight fixed, equally spaced truths that were drawn from no population whatever, the total squared error is 1.91 against 2.23 for taking each observation at its word — a 15% reduction obtained by pooling information between quantities that have nothing to do with each other.
What the strangeness actually is
The result is usually presented as an affront to intuition, and the affront is real but is not where it is usually located.
It is not that the estimate of a rainfall total is improved by knowing a batting average. It is not. Each individual estimate is, in general, made worse — the shrunk estimate of any particular one of the eight has higher expected squared error than that one’s own observation, for some true values.
What is improved is the total, and the total is a choice of criterion. Adding the eight squared errors together is a decision to treat them as one problem, and once that decision is made, the eight observations are jointly informative about how much of their spread is noise — which is the only thing the estimator uses them for.
So the honest summary is: the strangeness is in the criterion, not in the arithmetic. A reader who genuinely cares about one of the eight quantities and not the others has not been shown that shrinking helps, and the counting in the fourth essay of this field is about exactly that reader.
The “− 3” in the factor is where the criterion shows its hand: the result requires at least four quantities, and at three or fewer no such improvement exists. Nothing about a fourth measurement makes the first three easier to estimate individually. It makes the collection estimable as a collection.
Small groups move, large groups do not
Returning from the general result to the practical one, the property that decides what a hierarchical estimate looks like in a table is that the movement is by size, not by value.
At the population spread in these figures, the group of three moves 75% of the way to the middle and the group of forty moves 18%. Two groups reporting the same extreme value move by completely different amounts if they are of different sizes, and two groups of the same size move by the same fraction whether they are at the top of the table or the bottom.
This has a consequence that surprises people who expect shrinkage to be a form of conservatism: when the groups genuinely differ a lot, partial pooling barely changes anything. It is not a smoothing preference. It is a measurement of how much of the visible variation is real, and where the answer is “most of it”, the estimates stay where they were.
Unequal groups are the interesting case
Every figure in this field uses groups of unequal size — three to forty — and the reason is that equal groups hide the mechanism entirely.
With equal groups every se is the same, so every B is the same, and every estimate moves the same fraction of the way to the middle. The picture is a uniform contraction, which looks like a stylistic choice: someone decided the estimates should be less spread out.
With unequal groups the contraction is differential, and it becomes visible that the estimate is doing arithmetic rather than expressing a preference. The small group’s mean is mostly noise and it is treated as such; the large group’s mean is mostly signal and it is left alone. That is why the opening figure is drawn at sizes spanning a factor of thirteen, and why a demonstration of shrinkage on equal groups is a demonstration of nothing much.
The same weight in other clothes
Once B is read as a variance ratio rather than as a pooling rule, it turns up in places that are not usually filed under this heading, and recognising it is worth more than the individual results.
The Kalman gain is B. A filter combining a prediction with a new measurement weights them by their variances in exactly this form, with the prediction’s variance playing the part of τ² and the measurement’s noise playing se². A tracking filter that trusts a noisy sensor less is doing what a hierarchical model does to a three-observation hospital.
Ridge regression is B applied to coefficients. Shrinking a fitted coefficient towards zero by a factor that depends on its own standard error is the same weighted average with the population mean fixed at zero, and the ridge penalty is a way of writing τ.
A reliability coefficient is 1 − B. The share of observed variance that is real, in a psychometric test or a rating scale, is the complement of the share that is noise, and the standard correction for attenuation is this weight applied to a correlation.
None of these is an analogy. They are the same arithmetic reached from different starting points, and a reader who has understood one has understood the others — which is the practical reason the expression is worth memorising rather than looking up.
What the estimate does to the two ends of a table
The weight is a statement about each group separately, but its effect on a table is concentrated at the ends, and that is where its consequences are argued about.
The largest movements belong to the smallest groups, and the smallest groups are the ones most likely to be at the top and the bottom of an unpooled ranking. So partial pooling does most of its work precisely on the rows a reader looks at, and it will always look as though the method has singled out the interesting cases for correction.
It has not. It applies the same rule to every row and the rule happens to bite hardest where the evidence is thinnest. A group of forty at the top of the table stays near the top; a group of three at the top moves most of the way to the middle; and the difference between those two outcomes is a statement about how much was known, not about how much is believed.
The weight when the noise is not known
One simplification runs through everything above and deserves a paragraph rather than a footnote: se has been treated as known.
In practice the within-group spread σ is estimated too, usually pooled across the groups, and the estimate of it is far better than the estimate of τ because it has all the observations behind it rather than one number per group. That is why this field’s attention goes to τ: with eight groups, τ̂ rests on eight numbers, and with eight hundred observations spread across them, σ̂ rests on eight hundred.
The practical consequence is an asymmetry in what can go wrong. Getting σ slightly wrong moves every B slightly. Getting τ wrong — and with few groups it is easy to — changes the whole character of the answer, from “these groups are all the same” to “these groups differ”. The next essay is about that estimate.
When the weight should not be used at all
Three conditions have been assumed throughout and each of them can fail, so it is worth saying what the failure looks like rather than leaving the assumptions implicit.
The groups must be exchangeable. Not identical — the whole point is that they differ — but unlabelled: knowing nothing beyond the data, any group could have been any other. A set of groups that includes one measured by a different instrument, or one from a different country, is not exchangeable, and the model will treat a real difference as noise to be removed.
The population must be roughly the shape assumed. A normal population with one heavy outlier is not a normal population, and τ̂ inflates to accommodate the outlier, which weakens the pooling for every well-behaved group. The estimate is not robust in the technical sense, and a single wild group changes what happens to the other seven.
The noise must be independent of the value. Where a group’s variability grows with its level — counts, rates near zero, anything on a log scale — se is not a constant per group and the weights are wrong in a systematic direction. The usual repair is to model the transformed quantity, which is a change of model rather than a correction to it.
None of the three is exotic and all three are checkable in the data at hand. What they have in common is that each turns the weight from a measurement into an assumption, which is the state the whole field exists to get out of.
Three arguments, one estimator
The three derivations answer three different objections, which is why it is worth having all of them.
Why this weight and not another? Because it minimises squared error, which the first essay counts directly: 0.88 against 2.23 and 1.15.
Why is it legitimate to use other groups’ data on this group? Because under the model it is the posterior mean — the estimate the model’s own probability calculus produces — and that is checked against numerical integration to fifteen digits rather than trusted.
What if the groups are not really a population? Then Stein’s argument still applies, needs no population, and still reduces the total error, which is measured here at 1.91 against 2.23 on truths drawn from nothing at all.
Between them they cover the ground that a single derivation leaves exposed, and none of them is a statement about what is reasonable. Every one is a quantity that has been counted against a truth generated from a stated rule.
What is left to establish
The weight requires τ, and τ has so far been treated as available. The next essay is about where it actually comes from: a moment estimator built out of the spread between the group means, biased downward by about six per cent at twelve groups, and collapsing to exactly zero on more than half of all datasets when the groups are genuinely identical.
That is also where this field meets the Bayesian one most directly. A prior that the data supplies is not the prior anybody objects to, and the cost of supplying it — about a fifth of the estimate’s advantage, measured — is the price of not having to assert one.
What links here
Computed from the collection, not written here: the essays that point at this one.
Reads more easily once this is understood
Essays that name this one as worth reading first.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A prior on the spread — both name posterior, posterior mean, prior
- One population, or two — both name exchangeability, partial pooling, shrinkage
- The fewest groups that can borrow — both name james–stein, partial pooling, shrinkage
- An interval for something else — both name posterior, posterior mean
- Borrowing towards a line — both name partial pooling, shrinkage
- The base rate was always Bayes — both name posterior, prior
Named objects
A flat tag is an object no other essay names yet.
ExchangeabilityJames–SteinPartial poolingPosteriorPosterior meanPriorShrinkage