The spread, and its own uncertainty

When the spread estimates to zero

The usual estimate of a population spread is a difference of two positive quantities, clamped at zero. On a third of eight-group datasets with a real spread in them the difference comes out negative, the estimate is exactly zero, and every group is pooled completely on data that said no such thing.

Worth reading first: Eight groups, one population · The weight that decides.

The estimate of the population spread that partial pooling uses is a subtraction.

The observed group means are spread out for two reasons: the groups really do differ, and each group’s mean is measured with error. Their observed variance is therefore the real spread plus the average measurement variance, so the real spread is the observed variance minus the average measurement variance. Estimate both, subtract, take the square root.

There is nothing wrong with that argument. It is unbiased on the variance scale, it needs no distributional assumption beyond the two variances existing, and it is the reason the prior the data estimates can claim that τ is recovered from the data rather than supplied.

The difficulty is arithmetic. A difference of two positive quantities can be negative, a variance cannot, and something has to be done about that.

How often "there is no spread between the groups" is reported about data that has oneEvery dataset here was generated with a real population spread of 1. The moment estimator is the difference between the observed spread and what noise alone would produce, clamped at zero, and the difference comes out negative often: at eight groups it reports exactly zero on 32.6% of datasets, which is an instruction to pool completely and give all eight groups the same estimate. The rate falls to 4.2% at 48 groups.00.2000.4004681216243248number of groupsshare of datasets on which τ̂ is exactly zeroeight groups: 32.6%τ̂ = √max(0, between − within), truth τ = 132.6% of eight-group datasets report no spread at all
Fig. 1 Every dataset behind this curve was generated with a real population spread of 1. The curve is how often the estimate of that spread comes out exactly zero. At eight groups it is 32.6% of datasets; at four it is 44.5%; it takes forty-eight groups to get it down to 4.2%.

Zero is not a small number here

What is done about it is the obvious thing: the estimate is clamped at zero, because reporting a negative variance would be arithmetically purer and useless.

The clamp is correct and it is not the problem. The problem is what an estimate of exactly zero means downstream. Substituted into the shrinkage weight,

B = se² / (se² + 0²) = 1

for every group, whatever its standard error. A weight of 1 puts all the weight on the population mean and none on the group’s own data. Every group receives the same estimate. The analysis has not returned a small amount of borrowing; it has returned complete pooling, which is one of the two answers partial pooling exists to avoid.

One dataset on which the moment estimate collapsed. The eight groups' own means as hollow circles, the integrated posterior's interval as a bar, and the plug-in estimate as a filled mark. The moment estimate of τ came out exactly zero on this dataset, so every plug-in estimate is the same number, 0.14. The integrated posterior, on the same observations, still gives the groups different values, because its interval for τ runs from 0.02 to 1.82 and it never has to pick one.
Fig. 2 One dataset on which the estimate collapsed. Eight groups with standard errors from 0.5 to 2.1, and their own means scattered across a range of several. The filled marks on the left are the plug-in estimates: all eight are the same number. The bars are what the integrated posterior gives on exactly the same observations, and it still separates them, because its interval for τ runs well away from zero.

That figure is the whole essay. Two analyses of one dataset, differing only in whether a single number was estimated and substituted or carried and integrated over, and one of them has reported that eight groups are indistinguishable while the other has not.

Why it happens so often

An estimator that fires on a third of datasets is not misbehaving on unusual data. It is doing what its sampling distribution says it does, and the sampling distribution has an atom at zero — a lump of probability at a single point, not a density near it.

τ̂ across 2,000 datasets of 8 groups, true τ = 1. The population spread is not supplied to a hierarchical model — it is estimated from how far apart the group means are, after subtracting the noise that would separate them anyway. It averages 0.87 here against a true 1, and comes out exactly zero on 25% of datasets.
Fig. 3 The estimate’s whole distribution over two thousand datasets, at eight groups. The bar at zero is not a bin like the others: it is the probability of exactly zero, and it is the largest thing in the picture. Everything to the right of it is spread across a factor of several.

The mechanism is short. With J groups, the observed between-group variance is roughly τ² + (average se²) times a chi-square on J − 1 degrees of freedom over J − 1. At eight groups that multiplier has a standard deviation of about 0.53 — it is a very noisy quantity. Subtract a fixed average se² from something that noisy and the answer lands below zero whenever the multiplier lands below (average se²)/(τ² + average se²).

The rate has a closed form, where the design is simple enough

That last paragraph is an argument, and this site does not accept arguments about arithmetic without a second route. There is one available, and it needs no simulation at all.

If every group is measured to the same precision — the same standard error se throughout — then the observed between-group variance is exactly (τ² + se²) times a chi-square on J − 1 degrees of freedom divided by J − 1, and the average within-group variance it is compared against is the constant se². The estimate collapses precisely when

χ²(J − 1) < (J − 1) · se² / (τ² + se²)

which is a distribution function evaluated once. Nothing about the data enters except the ratio of the two spreads and the number of groups.

Against twenty thousand simulated studies at each setting, the two routes give:

groups τ se counted closed form
8 1 1 16.59% 16.48%
8 0.5 1 41.05% 41.28%
16 1 1 5.69% 5.77%
4 2 1 10.53% 10.36%

Every pair agrees to well inside the simulation’s own standard error — the largest disagreement is 0.8 of one. The closed form also settles the boundary case the figures cannot reach: at a population spread of exactly zero, where the groups genuinely are identical, the collapse probability at eight groups is 57.1%. That is the estimator behaving correctly. Where there is no spread, saying so most of the time is what it should do.

The unequal standard errors the figures use have no such formula — the between-group variance is then a weighted sum of chi-squares with different scales — which is why they count. Where both routes apply they agree, and where only one applies it is the counted one that is trusted.

Two further things follow, and the figures show both.

It is a small-J problem, not a defect of the estimator. The rate falls monotonically with the number of groups because the chi-square concentrates. At forty-eight groups it is 4.2%.

It is a small-τ problem too. Move the slider on the first figure and the whole curve drops: at a population spread of 2 the eight-group rate is 6.1%, and at 3 it is under 2%. Where the groups are genuinely far apart the subtraction is not close to the boundary.

How often "there is no spread between the groups" is reported about data that has one. Every dataset here was generated with a real population spread of 0.5. The moment estimator is the difference between the observed spread and what noise alone would produce, clamped at zero, and the difference comes out negative often: at eight groups it reports exactly zero on 51.2% of datasets, which is an instruction to pool completely and give all eight groups the same estimate. The rate falls to 32.9% at 48 groups.
Fig. 4 And where the groups are genuinely rather alike, the collapse is the typical outcome rather than a minority one: 51.2% at eight groups, and still 32.9% at forty-eight. Data generated with a real spread of 0.5, against standard errors running from 0.5 to 2.1.

Both dimensions point the same way and it is the same direction as everything else in this field. The estimator collapses exactly where partial pooling is most worth having: a handful of groups that are somewhat different, where neither obvious answer is right.

What the collapse costs, and what it does not

It would be easy to overstate this. Complete pooling is not a catastrophe; it is one of the two sensible answers, and on data where the groups really are nearly identical it is the better of the two. An estimator that says “pool completely” when the evidence for any difference is weak is not behaving stupidly.

So the point estimate survives the collapse in reasonable shape, and eight groups, one population already counted the squared errors that say so. What does not survive is everything else.

Every interval becomes the wrong width. With B = 1 the variance of a group’s estimate is the variance of the population mean and nothing else, which is much too small. This is the sharpest version of the 79% coverage the previous essay counted: the datasets where τ̂ collapses are the datasets where the plug-in interval is worst, and they are a third of all of them.

Every comparison between groups becomes impossible. All eight estimates are equal, so any question of the form “is group 3 higher than group 6” has the answer “they are identical”, which is a statement the data did not make.

And the report says nothing unusual happened. No warning is produced. The output is eight estimates, all the same, with intervals. An analyst who has not seen this before will read it as a finding — the groups turned out not to differ — when it is a property of the estimator at this sample size.

That last point is what makes this belong on this site rather than in a footnote. The failure is not that a number is wrong; it is that a number is absent, and the site has now met that shape three times. A tick function returning an empty array drew no gridlines. A class no stylesheet defined drew no lines. Here, an estimator returning its boundary reports no spread. In every case the output is well-formed, every check passes, and what is missing is missing quietly — because a check can ask whether a value is right and cannot easily ask whether a value should have been something other than the edge of its own range.

The general repair is the one the fleet has arrived at from three directions: look for the case where the machinery returns its own boundary, and ask how often that happens. Here the answer is a third of the time.

The rate is a function of the weight it is trying to compute

The closed form contains one quantity from the data and it is not a new one. Rewrite the threshold:

se² / (τ² + se²) = B

which is the shrinkage weight itself, at the common standard error. So the condition for the estimate to collapse is that a chi-square on J − 1 degrees of freedom falls below (J − 1) times the weight the analysis exists to compute. The collapse probability is a function of the answer.

The table’s rows are then one curve rather than four settings. At τ = se the weight is 0.5 and eight groups collapse 16.5% of the time; at τ = se/2 the weight is 0.8 and they collapse 41.3%; at τ = 2·se the weight is 0.2 and four groups collapse 10.4%. Every one of those is the same chi-square tail read at a different multiple of its degrees of freedom.

Written that way the essay’s central complaint stops being a coincidence and becomes an identity. The estimator fails most often exactly where the borrowing it is computing would have been largest — because a large weight means τ is small relative to the standard errors, and a small τ is what puts the subtraction near the boundary. There is no setting in which a study needs to borrow heavily and can also expect to be told how much to borrow.

It also gives a usable rule before any data exists. An analyst who can guess the ratio of the between-group spread to a typical standard error can read off the collapse probability from a chi-square table alone — no simulation, no pilot. Expecting to borrow three-quarters of each group’s estimate at eight groups means expecting the estimator to return zero on something near two datasets in five, and that is a design fact rather than an analysis surprise.

At a spread of exactly zero, more groups never help

The boundary case the closed form settles has a consequence the counted curves cannot show, and it is the sharpest thing in this essay.

At τ = 0 the weight is exactly 1, so the condition becomes χ²(J − 1) < J − 1: a chi-square falling below its own mean. At eight groups that is the 57.1% quoted above. At sixteen it is about 55%, at forty-eight about 53%, and as the number of groups grows it approaches 50% and stops. A chi-square’s median sits about two-thirds below its mean at every degrees of freedom, and that gap shrinks relative to the spread rather than disappearing.

So the reassurance that this is a small-J problem holds for every genuine spread and fails at zero. Where the groups really are identical, the estimator says so about half the time however many groups there are, and reports some positive spread the other half. Forty-eight groups do not settle it and four hundred and eighty would not either.

That is not a defect to be repaired, and it is worth saying why. A study of identical groups has nothing to estimate: τ sits on the boundary of its own parameter space, the sampling distribution of an estimator at a boundary is half an atom and half a density, and no amount of data moves a point off an edge. What it does mean is that “the estimate came out zero” carries no information about whether the groups are alike once the study is large — a coin flip carries none — while at eight groups it carries a little. The statistic that looked most informative on the smallest study is the one that stops meaning anything on the largest.

The same atom, one level up

This is not a quirk of one estimator. Any variance component estimated as a difference of mean squares has the atom, and a study with more than one level of grouping has more than one such component.

A three-level design — observations inside groups inside clusters — estimates its top-level variance as the difference between the between-cluster mean square and the between-group mean square, divided by the number of observations in a cluster. It is the same subtraction with the same clamp, and it collapses for the same reason. Measured on eight clusters of three groups of four, with a genuine cluster-level spread of 0.5 present in the data, the top component is reported as exactly zero on 18.9% of datasets.

The consequence is worse one level up than it is here, because the collapse of a top-level component does not merely pool the clusters. It removes the level from the model entirely: with the cluster variance at zero, the design effect is one, the clustering has no consequence for any standard error, and the analysis proceeds as though the study had been a simple random sample. What two levels at once counts is what that costs.

The pattern is worth naming because it will recur wherever this machinery goes. A variance estimated by subtraction has a probability of being reported as absent, and that probability is largest exactly when the number of units at that level is smallest — which is the top of every hierarchy, because a hierarchy narrows as it rises.

What the integral does instead

The posterior for τ has no atom at zero. It cannot: a continuous prior against a likelihood that is finite at zero gives a density, and a density puts zero probability on any single point.

What eight groups say about τ, when the truth is 0.25. The posterior density for the population spread after eight groups whose standard errors run from 0.5 to 2.1. The shaded band is the central 95% interval, from 0.36 to 3.17; the posterior median is 1.26 and the mean 1.39. The vertical mark at 0.00 is the moment estimate that empirical Bayes substitutes and then treats as known.
Fig. 5 The posterior for τ on data whose moment estimate is exactly zero. Its median is 1.26 and its central interval runs from 0.36 to 3.17. The data has not said the groups are identical; it has said it cannot distinguish a spread of 0.4 from one of 3, which is a different statement and the true one.

That is not the integral being cleverer. It is the integral being asked a different question. The moment estimator answers what single value of τ best explains the spread between these means, and on a third of datasets the honest answer to that question is “none of them is any better than zero”. The integral answers what does the data say about τ, and the answer is a range, which the data can always supply.

The consequence for the estimates is what the second figure showed: the groups stay separated, because the posterior averages over values of τ that separate them, weighted by how well each explains the data. No value of τ has been chosen, so no value of τ can be chosen badly.

The check this needed

The rate quoted here is a count and the count is asserted rather than reported, which on this site means it has to be written so that it can fail.

Two earlier versions of that assertion were wrong in opposite directions and both are worth recording, because they are the same mistake. The first claimed that the estimate is low on average — that the clamp biases it downwards. That is true at a population spread of 2 and false at 0.25, where the clamp pushes the mean estimate up to 0.43 against a truth of 0.25, because an estimator that cannot go below zero and can go arbitrarily high has its mean dragged upward. The second claimed the distribution is right-skewed, which is true at 0.25 and false at 2, where the median and mean have converged.

Both were facts about a region of the slider, asserted as though they were facts about the argument. The assertion that survives is the one the figure is actually about: at eight groups the estimate collapses to exactly zero on some share of datasets that all have a real spread in them, and the share is reported. That holds at every setting the slider reaches, which is what an assertion in a figure is allowed to claim — anything that needs two settings to state belongs in the site’s own gate, where the comparison across settings can be made.

What to do about it

Three responses are available and only one of them is the one usually reached for.

Use more groups. Correct, and rarely possible. The number of groups is a property of the study.

Use a better point estimator. Marginal maximum likelihood has the same atom — it hits the boundary whenever the moment estimate would go negative — so this does not help as much as it sounds. Restricted maximum likelihood improves the bias and keeps the atom.

Stop taking a point estimate. The integral has no atom, needs no clamp, and produces intervals that cover. It costs a one-dimensional quadrature, which on eight groups is a few hundred evaluations of a normal density.

The third is the one this field argues for, and the argument is not that Bayesian machinery is preferable. It is that the quantity being estimated is badly determined by eight numbers, and every method that insists on returning one value for it will spend a third of its time returning the most extreme value available.

There is a fourth response worth mentioning because it is the one most often taken in practice, and it is not a response at all: report the collapse as a result. No significant between-group variation was detected. The chi-square tail above says exactly what that sentence is worth at eight groups with a real spread of 1 — it will be written about one dataset in six even when the groups do differ by a whole standard error, and about two datasets in five when they differ by half of one. It is a statement about how many groups the study had.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

DiscretenessEmpirical BayesHierarchical modelMethod of momentsPartial poolingPosteriorShrinkageVariance components