One population, or two
Worth reading first: Eight groups, one population.
When borrowing goes wrong is about one group that is not from the population. Everything else in that essay’s setup is correct: the other seven groups are genuine draws from a normal, the model is right about them, and the failure is localised.
This is the case where the model is wrong about all of them at once, and no single group is unusual.
Matched on the only thing the fit reads
The construction matters, so it is worth stating precisely.
Half the groups have truths near and half near , with a little spread of their own around each cluster. The two parameters are chosen so that the total variance of the group truths is exactly — the same a single normal population would have. Ninety per cent of that variance is between the clusters and ten per cent inside them, which leaves the middle of the gap genuinely empty: 6.6% of the truths fall in the middle half of it.
The moment estimator reads exactly one thing about the population: the spread between the observed group means, minus the average measurement error. That quantity is a variance. Variance does not know about shape.
So from the mixture and from a normal population of the same spread are the same number, every shrinkage weight is identical, every group moves the same fraction of the way it would have, and nothing in the fitted output is different.
What the estimates do instead
The weights being right does not make the estimates right, because the weight decides how far each group moves and the population’s shape decides where it should have moved to.
Each group is pulled a fraction of the way towards the grand mean. For a normal population the grand mean is the middle of the distribution and pulling towards it is pulling towards where the truths mostly are. For two clusters the grand mean is the middle of the gap, and pulling towards it is pulling away from both places the truths actually are.
Forty-six per cent against six point six. The analysis has taken a bimodal population and produced an output whose densest region is the one place the population is not.
At a spread of two, 50.5% of groups are moved further from their own cluster’s centre than their own mean was, which is the same statement about individual groups rather than about the output as a whole: for half the groups, borrowing moved the estimate in the wrong direction.
Exchangeable is not the same as normal
The assumption partial pooling is usually stated under is exchangeability: the groups are interchangeable before the data is seen, so any ordering of them is as likely as any other. That is a weaker and more defensible assumption than normality, and it is often offered as the reason the normal population is not really an assumption at all.
The two clusters here are exchangeable. Nothing labels a group as belonging to one cluster or the other, the groups arrive in no order, and permuting them changes nothing about the analysis. A Bayesian’s justification for the model is intact.
What exchangeability delivers is that the group effects are draws from some common distribution. What it does not deliver is which one, and the analysis has to commit — to a normal, because that is what makes the shrinkage a weighted average and the arithmetic closed-form. So exchangeability licenses the form of the model and the normality is a separate assumption sitting inside it, doing work, and it is the one that fails here.
That distinction is worth holding because it is usually collapsed. A defence of the hierarchical model from exchangeability is not a defence of the normal population, and every figure in this essay is a case where the first holds and the second does not.
What it costs, which is less than it looks
The squared error is where this becomes an argument about how much to care, and the honest answer is “less than the pictures suggest”.
Against the two-cluster population, partial pooling costs 1.1576 per group. Against a normal population of the same spread it costs 1.0655. That is 8.6% worse — real, reproducible, and nothing like the factor the gap-landing figures imply.
It is also still much better than the alternatives on the same data: the groups’ own means cost 2.2534 and complete pooling costs 1.2694. Partial pooling against a badly misspecified population still beats both of the estimators it sits between.
The penalty also shrinks as the clusters separate: 8.6% at a spread of one, 8.1% at two, 4.7% at three, 3.1% at four. That is the opposite of what the pictures suggest, and the reason is that a large spread means a small weight — the groups are barely pooled at all, so the wrong centre is barely used.
So squared error, which is the criterion this field has ranked everything on, says the misspecification is a minor inefficiency.
Why squared error is the wrong criterion here
That conclusion is correct and it is not the one to act on, and separating the two is the whole of what follows.
Squared error is a sum over groups of a squared distance. It is dominated by the groups that are furthest from their estimate, which are the poorly-measured ones, and it is almost unaffected by whether an estimate is in a place the population occupies. A set of estimates that all sat exactly at the grand mean would have a large squared error and would at least not be making a claim about structure. A set that is pulled from two clusters into a gap has a slightly larger squared error than it should and is making a false claim about structure, and the second thing is not in the criterion at all.
The questions a two-level study is usually run to answer are questions squared error does not score:
Are the groups one population or several? The output says one, because it was told to say one and its estimates are more unimodal than the data was.
Which groups are alike? The pooled estimates in the gap are groups from both clusters, moved towards each other. Two groups reported as similar may be from opposite clusters.
Is any group unusual? The most extreme group in each cluster has been pulled towards the middle hardest in absolute terms. Unusualness has been compressed out.
Every one of those is a statement about the shape of the set of estimates, and the shape is the thing the assumption supplied rather than the thing the data determined.
Half the groups are moved the wrong way
The gap-landing count is a statement about the output as a whole. The same failure read one group at a time is sharper.
Pooling is justified to a group as follows: the group’s own mean is noisy, the other groups say something about where it probably sits, and moving it towards them moves it towards the truth. The checkable version of that promise is whether the pooled estimate is closer to the group’s own cluster’s centre than the group’s own mean was.
At a spread of one it is not, for 38.7% of groups. At a spread of two, 50.5%. At three, 53.0%, and at four, 53.2%.
So on a two-cluster population, once the clusters are separated at all, more than half the groups are moved further from where they came from. The promise is broken for the majority, and the total squared error still improves, because the groups it is broken for are the well-measured ones whose errors are small and the groups it is kept for are the badly-measured ones whose errors dominate the sum.
That is the same concentration the accounting of where the borrowing goes measured, arriving here as a reason the total cannot be trusted to report this failure. A criterion dominated by two groups out of eight will not notice a procedure that misbehaves for five of them.
The contrast with the single outlier
Setting this beside the field’s earlier failure says what is new about it.
The outlier case is a large error in a known place. The group that was hurt is the one that looked unusual, which is the group anybody would check, and its own mean is available as a sanity check that visibly disagrees with its pooled estimate.
The mixture case is a small error everywhere, in no identifiable place. No group’s pooled estimate disagrees dramatically with its own mean, nothing looks like an outlier, and the total error is barely affected. What is wrong is a property of the eight numbers together, and each of the eight is individually unremarkable.
That shape has a name by now. It is the same class as a variance estimated as exactly zero and an interval on a prior that covers nothing: the output is well-formed, every quantity in it is computed correctly, and the defect is in what the output is a picture of.
What does see it
The fitted model cannot detect the shape because the model has no parameter for shape. Three things can, and none of them is inside the standard fit.
The observed group means themselves. They are the data, they are blurred by noise but not transformed, and at a spread of one they land in the gap 20.1% of the time against the truths’ 6.6% and the estimates’ 46.3%. The raw means are the least misleading picture of the population’s shape available, and a pooled analysis replaces them.
A group-level predictor, where one exists.
If the two clusters correspond to something recorded — rural and urban sites, two manufacturing lines, two protocol versions — then a single indicator in the group-level model turns the mixture into two normal populations, each correctly specified, and the whole difficulty disappears. The cure for a bimodal population is usually a column somebody already has.
The raw means deserve a second sentence, because the comparison is unusually favourable to them. Twenty point one per cent against six point six is an overstatement of the gap’s population by a factor of three; the pooled estimates’ 46.3% is an overstatement by a factor of seven. The noise-blurred data is wrong in the same direction and less than half as far, and it is the object a pooled analysis exists to improve on. Wherever the question is what shape is this population, the improvement runs backwards.
And a posterior predictive check, which is the general answer when no such column exists: simulate group effects from the fitted model, compare their spread shape against the observed means, and see whether the fitted unimodal population could have produced what was seen. That is a real diagnostic with real power and it is not counted here, because counting it properly means choosing a discrepancy statistic and measuring its size and power, which is an essay rather than a paragraph.
Where this leaves the central claim
Partial pooling has now been put four times against something it was not built for, and the results are consistent enough to state as one sentence.
Against the population it assumes, it beats both of the estimators it sits between, decisively. Against too few groups it degrades gracefully into complete pooling, which is a defensible answer. Against one group that does not belong it ruins that group’s estimate and improves the rest. Against a population of the wrong shape it costs 8.6% of squared error and produces an output that misdescribes the population.
Every one of those failures is invisible in the fitted output, and none of them is a failure of the arithmetic. The weight is computed correctly in all four cases. What changes is whether the thing the weight is a weight on — a common population, with a centre worth moving towards — is there.
That is why this field’s repairs are all outside the fit: more groups, a group-level predictor, a prior with a wider tail, a look at the raw means. None of them is a better estimator. They are ways of checking the sentence the estimator takes as given, and the estimator has no way of checking it because it is the sentence that makes the estimator an estimator.
What is claimed here, and what is not
The claim is what partial pooling does to a population with the right variance and the wrong shape: that the estimated population spread differs from a normal population’s by 5% of its own between-dataset spread and is therefore unreadable from one study; that squared error worsens by 8.6% at a spread of one and by less as the clusters separate; and that the estimates land in the empty middle of the gap 46.3% of the time against the truths’ 6.6%.
Everything is three thousand datasets per point with the spread estimated from the data, which is what an analyst has. The two populations share a seed sequence and differ only in how the group truths are drawn.
What stays out: a posterior predictive check, named above and not measured; mixture models fitted as mixtures, which are the correct analysis and a different subject with their own label-switching and identifiability problems; three or more clusters, where nothing in the argument changes and the arithmetic is longer; and heavy-tailed populations, which are the other common way for a normal assumption on group effects to be wrong and which fail in the opposite direction — there the outliers are real and the model over-shrinks them, which is the single-group case arriving many times over.
The 6.6% figure is a property of the construction rather than a measurement, and the choice to put a tenth of the variance inside the clusters is a choice: at a quarter the gap holds 19% of the truths and every contrast in this essay is milder. The construction is stated so the number can be read as what it is.
Still open: what the estimates would have to show
Every diagnostic named above is external to the fit. The question left unanswered here is whether the fitted output itself carries a usable signal of the shape it has assumed away — whether the eight pooled estimates, or the eight residuals between each estimate and its group’s own mean, are distributed differently under a mixture than under a normal.
They must be, since the estimates are a deterministic function of the data. Whether the difference is large enough to see at eight groups, after shrinkage has compressed exactly the structure a test would look for, is a question about the power of a test nobody here has written — and the answer, if it is the one the compression suggests, would mean the fit destroys the evidence of its own misspecification.
The check, and the refusal
Two claims are gated, and the second is the one the essay rests on. That the population spread recovered from two clusters differs from a normal population’s by less than a quarter of what the estimate moves between datasets — stated relative to the estimator’s own spread rather than as an absolute tolerance, because a difference is only invisible compared with something. And that the estimates land in the gap at more than twice the rate the truths do, at every spread on the sweep.
The refusal is the reassurance a reader would reach for: the fitted model sees the clusters, rejected. The two τ̂ values are required to agree to within a quarter of the estimate’s own between-dataset spread. If that check ever failed — if the fit did carry a large signal of the shape — the essay would be describing a detectable problem rather than an invisible one, and the whole argument would have to be rewritten around a diagnostic that already exists.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A group from the population's own tail — both name empirical bayes, hierarchical model, partial pooling, shrinkage
- A level with two units — both name hierarchical model, method of moments, random-effects, variance components
- The prior the data estimates — both name empirical bayes, method of moments, partial pooling, variance components
- The slope that borrows — both name hierarchical model, partial pooling, random-effects, shrinkage
- What a two-unit study should report — both name hierarchical model, partial pooling, random-effects, variance components
- What the plug-in forgets — both name empirical bayes, hierarchical model, partial pooling, shrinkage
Named objects
A flat tag is an object no other essay names yet.
Empirical BayesExchangeabilityHierarchical modelMethod of momentsMixtureModel misspecificationPartial poolingRandom-effectsShrinkageVariance components