The shape the estimates keep
Worth reading first: Eight groups, one population.
One population, or two gave a hierarchical model twelve groups whose effects came from two clusters rather than from one bell, with the total spread between them exactly the spread the model assumes. The fit recovered that spread, used the shrinkage weight a normal population would have earned, and pulled nearly half its estimates into the empty middle between the clusters, while reporting nothing unusual. It ended on a question about the fit’s own output. The estimates are a deterministic function of the data, so they must differ under a mixture from under a normal population; whether the difference is large enough to see at eight groups, after shrinkage has compressed exactly the structure a test would look for, was the open part — with the suspicion that the fit destroys the evidence of its own misspecification.
The suspicion has a precise answer, and for the commonest case it is that the fit destroys nothing.
An affine map keeps a shape
With every group the same size, every group’s mean has the same standard error, and partial pooling moves every mean the same fraction of the way towards the same centre. The pooled estimate of group j is the grand mean plus times the group’s deviation from it, with one weight B for every group. That is an affine map — a shift and a uniform compression — and a uniform compression changes the scale of a set of numbers and nothing else. Any statistic of the estimates that does not depend on their scale or location is exactly the same number computed on the raw group means.
The figure is one dataset. Sixteen groups of four observations, effects from two clusters with a total spread of two, noise of three. The true effects have a kurtosis of 1.203 — far below a normal population’s 3, which is what two clusters look like to a fourth moment. The observed means, blurred by the noise, have a kurtosis of 2.348; the noise has pushed them most of the way back towards a bell. The fitted spread is 1.701, so every estimate is pulled 43.7% of the way to the centre, and the pooled estimates have a kurtosis of 2.348 — the observed means’ to the last digit, because they are the observed means, scaled.
So the compression the earlier essay worried about happens, and it is invisible to every shape test. What made the twelve estimates look like one population was not the pooling. It was the noise in the group means, which the pooling inherits and does not add to.
Where pooling does erase the shape
There is one place the map is not merely a compression. When the fitted spread between groups is zero, the weight is one, every estimate is the grand mean, and the estimates have no spread and no shape at all. A test on them cannot reject anything.
At the earlier essay’s spread of one, with groups of four observations and noise of three, that happens often. On eight groups the fit pools completely on 29.9% of two-cluster datasets and 31.5% of normal ones; on sixteen, 18.3% and 19.7%; on sixty-four, 2.1% and 2.6%. Those are the datasets where the pooled estimates genuinely carry no information about the population’s shape, and they are not a rare corner: at eight groups they are a third of all studies.
But they are datasets where the raw means carry almost none either. A fitted spread of zero means the observed means are no more dispersed than their noise alone would make them, so whatever shape the true effects have is buried under noise several times its size. The pooled estimates throw away a signal the raw means hold only faintly — and on these datasets, the raw-means version of the test reports two clusters 6.6% of the time, about as often as it would by chance.
The evidence was never there
That is the real finding, and it is about power rather than pooling.
The test reads the sample kurtosis of the group values and rejects when it is lower than a normal population of the same spread would produce 95% of the time. At eight groups and a spread of one, it rejects a two-cluster population 6.5% of the time — a test of size 5% with essentially no power. At a spread of two it reaches 12.2%, and at four, where the clusters are separated by several times the noise, 45.3%. Sixteen groups lift those to 6.0%, 19.4% and 69.4%; sixty-four groups to 8.1%, 45.5% and 99.5%.
So at the size the earlier essay worked at — eight groups, a spread comparable to the noise — no test of shape applied to the fit’s output or to the data could have seen the clusters, and the fit is not the reason. The noise in a group mean of four observations, a standard error of 1.5, is nearly as large as the distance between the clusters’ centres, 1.9, and eight such means are too few to estimate a fourth moment. The model’s failure to report its own misfit is a failure the data shares.
Why the noise, and not the pooling, hides the clusters
How much of the population’s shape survives into the group means can be written down, because noise adds to a fourth moment in a simple way. A group mean is its true effect plus an independent normal error, and for a sum of independent parts the fourth cumulant adds while a normal part contributes none. So the excess kurtosis of the means — their kurtosis less a normal’s 3 — is the true effects’ excess kurtosis times , the square of the share of the means’ variance that is real.
For this two-cluster population the true effects’ kurtosis is 1.38, an excess of −1.62: as far below a bell as a symmetric population with a little spread inside each cluster can be. With groups of four observations and noise three, each mean’s standard error is 1.5. At a spread of one, the real share of the means’ variance is under a third, its square under a tenth, and the means’ kurtosis is 2.85 — a departure from the bell of 0.15. At a spread of two the means’ kurtosis is 2.34, and at four, 1.75.
Against those signals stands the sampling error of a kurtosis estimated from J values, whose standard deviation for a normal population is about : 1.7 at eight groups, 0.6 at sixty-four. A departure of 0.15 measured with an error of 1.7 cannot be seen, and one of 0.66 measured with an error of 0.6 is seen about half the time — which is where the counted powers sit, two routes that share nothing arriving at the same verdict. Pooling never enters the arithmetic, because a uniform compression multiplies the deviations from the centre by one factor and the kurtosis divides it out.
The formula also says what would make a shape visible at eight groups. Larger groups shrink the standard error and raise the real share of the variance; at sixteen observations per group the standard error is 0.75, the share at a spread of one is 0.64, and the means keep two fifths of the true effects’ departure instead of a tenth. The design decision that makes a hierarchical model’s population checkable is the same one the fewest groups that can borrow priced for its estimates: how much each group is measured, not how many groups there are.
The argument does not depend on the departure being two clusters. A population of group effects with heavy tails — a few groups far from the rest, the case when borrowing goes wrong met one group at a time — has a positive excess kurtosis, and the same dilution applies with the same factor: the noise divides the departure by the same square of the real share, whichever sign it has. And the same invariance holds: with equal group sizes the pooled estimates have the means’ kurtosis exactly, so a check for heavy tails is no weaker for being run on the estimates. What differs between the two departures is only what a reader of the estimates is shown — clusters crowded into a gap, or outlying groups pulled in too far — and where the borrowing goes is the account of that, group by group.
The check an analyst can run
The test above is calibrated against the true null — a normal population with the actual spread — which no analyst has. The version an analyst can run uses the fitted model instead: simulate group means from the fitted normal population, with the fitted spread and each group’s own standard error, two hundred times, and see how often the simulated means are as low in kurtosis as the observed ones. That is a posterior predictive check with a plug-in spread, and the dots in the figure are its power.
It loses almost nothing for not knowing the truth. Its size, measured on normal populations, stays between 4.5% and 5.3% at every setting, and its power tracks the calibrated test’s closely: 6.2% at eight groups and a spread of one, 43.6% at a spread of four, 91.1% at thirty-two groups and four, 45.9% at sixty-four and two. The fitted spread is estimated well enough for the purpose because the shape test asks about the fourth moment relative to the second, and a reference distribution built from an estimated second moment is still the right shape.
Unequal groups
The affine argument needs every group to be shrunk by the same weight, and groups of different sizes are not. A group of two observations has a standard error of 2.1 and a group of sixteen 0.75, so the first is pulled much further towards the centre than the second. That is no longer a uniform compression, and the estimates’ shape is no longer the means’ shape — and here, the change helps.
With sixty-four groups of mixed sizes and a spread of two, the test on the observed means has 45.5% power, and on the pooled estimates 75.0%. Means standardised by their own predictive spread reach 68.3%, and the fitted-model check, which uses that standardisation, 65.7%. At eight groups the gaps are small: 16.3%, 18.9%, 17.6% and 16.3%.
The reason is that the observed means carry noise unevenly. The small groups’ means are scattered by noise larger than the gap between the clusters, and they fill the gap with values that look like a single broad population. Shrinkage pulls exactly those noisy means hardest towards the centre and leaves the precisely measured large groups where they are, so the pooled estimates are dominated by the groups whose means reflect their clusters. The fit that was suspected of hiding the clusters makes them easier to see, by trusting each group in proportion to what it actually knows.
What this changes in the earlier account
One population, or two said the raw means are the least misleading picture of a population’s shape, and for the question it asked — how many values land in the empty middle between the clusters — that remains true: the pooled estimates crowd into the gap and the raw means do not. For the question asked here — whether a test can tell two clusters from one bell — the answer is different. With equal group sizes the raw means and the estimates are the same evidence; with unequal ones, the estimates are better evidence.
The two answers are about different things. Counting values in a region measures where individual estimates sit, and pooling moves them; a shape test measures a scale-free property of the whole set, and pooling either preserves it exactly or improves it. A reader looking at a forest plot of pooled estimates is reading the first kind of thing, and is misled; a check computed from the same estimates reads the second kind, and is not.
The two can be put side by side in one dataset. In the earlier essay’s setting nearly half the pooled estimates sit in the middle half of the gap between the clusters, where a fifth of the raw means and a fifteenth of the true effects sit, so the plot shows one population more convincingly than the data do. The kurtosis of the same estimates is the raw means’ kurtosis exactly. The picture moves the estimates into the gap by compressing everything towards the centre, and the compression that fills the gap is the same compression a scale-free statistic divides out. What the plot exaggerates, the check ignores, and what the check measures, the plot never showed — which is why the honest report of a hierarchical fit carries both the estimates and the check, and why neither stands in for the other.
What a hierarchical analysis should report about its population
A shape check, with its power stated. A posterior predictive check of the group effects’ kurtosis is cheap, holds its size, and at sixty-four groups of moderate spread it finds a two-cluster population most of the time. At eight groups a clean check is not evidence of one population, and the report should say how much power it had.
Whether the fit pooled completely. A fitted spread of zero leaves estimates with no shape at all; on eight groups at a modest spread that is a third of studies, and on those the question of shape cannot be asked.
The group sizes. With unequal groups the pooled estimates are the better basis for any check on the population, and a check run on raw means throws away the weighting that makes it powerful — the same weighting the weight that decides derived for each estimate, doing a second job.
Counted, on what
Four thousand datasets at every setting, the two-cluster population of the earlier essay — ninety per cent of the spread between the clusters, ten per cent within each — and a normal population with the same spread, groups of four observations with noise three, or groups of 2, 4, 8 and 16 in turn. The spread between groups is estimated by moments and floored at zero, as in eight groups, one population. The test’s 5% point is the fifth percentile of its statistic over the four thousand normal datasets at the same setting; a dataset whose values have no spread is counted as not rejecting. The fitted-model check uses two hundred simulated sets per dataset. With four thousand datasets a power near 50% carries a standard error of about 0.8 points, and the gaps between the equal-size versions of the test are within that. Only one shape statistic is used; a test designed for bimodality rather than for low kurtosis might do better on two clusters and worse on other departures, and is not measured.
Still open: a test that knows what it is looking for
The kurtosis test is a general check on the fourth moment, and two clusters are a specific departure: two modes, a gap between them, and a known number of components. A test built for that — the likelihood ratio between one normal population and a mixture of two, fitted to the group means with their standard errors — should be more powerful, and it brings back the problems when borrowing goes wrong met in their simplest form: a mixture’s likelihood ratio has no standard reference distribution, and its fitted components can collapse onto single groups.
Whether a mixture-likelihood test finds two clusters among eight groups where the kurtosis test cannot, what its size is when calibrated by simulation from the fitted normal, and whether a positive result can be acted on — by fitting the mixture, or by looking for the group-level column that would explain it, which borrowing towards a line showed is usually the better repair — is the measurement this leaves.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A group from the population's own tail — both name empirical bayes, hierarchical model, partial pooling, shrinkage
- A history that stepped — both name hierarchical model, model misspecification, partial pooling, statistical power
- Estimates that are too alike — both name empirical bayes, hierarchical model, partial pooling, shrinkage
- The relation the table has to estimate — both name hierarchical model, model misspecification, partial pooling, shrinkage
- The small groups one centre protects — both name empirical bayes, hierarchical model, partial pooling, shrinkage
- What the plug-in forgets — both name empirical bayes, hierarchical model, partial pooling, shrinkage
Named objects
A flat tag is an object no other essay names yet.
Between group varianceEmpirical BayesHierarchical modelKurtosisMixture distributionModel diagnosticsModel misspecificationMonte CarloPartial poolingPosterior predictive checkShrinkageStatistical power