The prior the data estimates
Worth reading first: Eight groups, one population · What a prior is worth.
The weight that decides how much a group borrows is B = se²/(se² + τ²), and τ is the real spread between the groups. An analyst does not have it. What this essay is about is that nobody needs to supply it, because it is visible in the data — and that this changes what kind of object a prior is.
Two things are mixed together, and one of them is subtractable
The group means differ from each other for two reasons: the groups are genuinely different, and each group’s mean is measured with error. The observed spread contains both, and adding them is the easy direction —
observed spread ≈ real spread + measurement noise
The measurement noise is not an unknown. Each group’s own standard error is computable from its own data: σ/√n, with σ estimated from the within-group variation, which has every observation behind it. So the average noise is known, and subtracting it leaves the real spread:
τ̂² = max(0, spread between the group means − average se²)
That is the whole derivation. It is the method of moments, it is nearly unbiased for τ², and it is what the term variance components refers to: the total variation decomposed into a part between groups and a part within them, with the between-groups part being the one the pooling weight needs.
The max is doing something honest
The clamp at zero is not tidiness. When the groups really are identical, the observed spread is on average exactly the noise, so the difference is negative about half the time — and a negative variance estimate is not an embarrassment to be hidden, it is the data reporting that the groups are no further apart than sampling error alone would put them.
Reported as zero, that is an instruction: pool completely. Measured across two thousand datasets of twelve identical groups, the estimate comes out exactly zero on 53% of them.
An estimator that could never say zero would be an estimator that could never say the groups are the same, and it would put a little bit of separate estimation into every dataset regardless of evidence. The clamp is what allows the answer “there is nothing here”.
What it costs to estimate rather than know
Every advantage claimed for partial pooling in the first essay of this field was measured with τ supplied. That is the idealised case and it is worth separating from the real one.
With τ known, the mean squared error per group is 0.88. With τ estimated from the same eight groups being pooled, it is 1.05. Both against 2.23 for taking each group at its word and 1.15 for one number for everybody.
So the fee for not knowing τ is about a fifth of the advantage, and what is bought with it is that nothing was asserted. That trade is the whole argument for empirical Bayes, and it is a number rather than a preference.
The two figures differ in exactly one respect and the difference is visible at the left-hand end. Where the groups are genuinely identical, an estimator that has to discover that fact does slightly worse than one that was told it. Everywhere else the two are close, and both are far below the obvious answers.
The fee depends entirely on what it is a fee against
“About a fifth” is the fee measured against the error itself — 1.05 against 0.88 is nineteen per cent more error — and that is the mildest of the three denominators available. The other two are worth writing down, because they are what somebody choosing between the three procedures is actually deciding between.
Against complete pooling. One number for everybody costs 1.15. Knowing τ buys 0.27 of that; being made to estimate it buys 0.10. So estimating the population spread gives up 63% of what knowing it was worth, against the better of the two procedures that need no population at all.
Against no pooling. Taking each group at its word costs 2.23. Knowing τ buys 1.35; estimating it buys 1.18, which is 87% of it.
Both are correct and neither is the headline on its own. The honest summary is that the fee is small beside the whole gain and large beside the margin over the nearest competitor — which is what a fee usually looks like when a procedure is being compared with the closest thing to it rather than with the worst thing on the list.
The same three numbers, in observations
The section above prices the population in observations, and the same conversion prices every row of the table. A group’s error under no pooling is σ²/n, so an error of E is the accuracy of σ²/E observations however it was obtained.
At σ = 3 the four rows read:
- taking each group at its word: 2.23, which is 4.0 observations — the four each group actually has
- complete pooling: 1.15, which is 7.8
- partial pooling with τ estimated: 1.05, which is 8.6
- partial pooling with τ known: 0.88, which is 10.2
So the population is worth 6.2 extra observations to each group when its spread is known, and 4.6 when it has to be discovered — against the nine that σ²/τ² promises, the difference being what a grand mean estimated from eight groups costs rather than one supplied.
That is the whole field in one column. Four observations behave like ten if the analyst is told how alike the groups are, like nine if the data has to say so, and like eight if the question is not asked at all.
The prior is worth a stated number of observations
The Bayesian field on this site treats a prior as a component with a measurable cost, and states that cost in observations: a Beta(a, b) prior is worth a + b of them, and the posterior mean is the data’s proportion pulled towards the prior’s in exactly that proportion.
The same statement is available here and it is the cleanest way to read B. The population acts on every group as
σ²/τ² extra observations
At the within-group spread and population spread used throughout these figures — σ = 3, τ = 1 — that is nine observations. A group with three of its own is therefore three parts population to one part itself, which is the 75% weight; a group with forty is forty parts to nine, which is 18%.
That is a genuinely useful sentence for reading a fitted model. “The population is worth nine observations to each group” is checkable, arguable and immediately interpretable, in a way that “τ̂ = 1.02” is not.
The estimate is biased, and in which direction
τ̂ is not an unbiased estimate of τ, and the direction is worth knowing because it is the direction that makes the model look better behaved than it is.
Across two thousand datasets of twelve groups generated at τ = 1.5, the estimate averages 1.41 — about six per cent low. Two mechanisms push it there. The moment estimator is close to unbiased for τ², and the square root of a noisy estimate of a square is biased downward, by an amount that grows with the noise. The clamp at zero pushes the other way, and at moderate τ the square root wins.
The consequence is a mild systematic tendency to pool slightly harder than the truth warrants. A group’s estimate is pulled a little further towards the middle than a model with the true τ would pull it, and the differences between groups are reported as slightly smaller than they are.
For eight or twelve groups that effect is small next to the gain, which is why the estimator is used. For three or four groups it is not small, and that is one of the reasons the next essay treats a small number of groups as a failure mode rather than as a mild inconvenience.
What six per cent of τ is worth at the weight
A bias in τ matters only through B, and B is not very sensitive to it, which is why the estimator survives being biased.
At these settings the standard error of a group mean is σ/√n = 1.5, so se² = 2.25. Generating at τ = 1.5 makes the true weight
an even split between the group and the population. The average estimate of 1.41 gives τ̂² = 1.99 and a weight of 0.531.
Three points of over-pooling for six per cent of bias in τ. In the observation register of the section above, the population is credited with 4.5 observations where it has earned 4.0 — half an observation too many, applied to every group.
That is why the direction of the bias is worth naming and its size is not worth correcting. The effect on each group’s estimate is a shift of three per cent of the distance to the grand mean, which is inside the noise of the estimate being shifted, and any correction would have to be estimated from the same eight numbers that produced the bias.
The reason it stops being negligible with three or four groups is that both halves get worse together: τ̂ is noisier, so the square root’s downward bias is larger, and the weight is more sensitive because there are fewer groups holding the grand mean in place. The failure is not that the bias grows; it is that it stops being small relative to everything else.
What τ is not
Two spreads are in play and they are routinely confused, including in published summaries, so it is worth stating the difference in the terms a table presents them in.
σ is the variation between observations inside a group. It is what a group’s own standard deviation estimates, it has every observation in the study behind it, and it says nothing whatever about whether groups differ.
τ is the variation between the groups’ true values. It has as many pieces of information behind it as there are groups — eight numbers, not eight hundred — and it is the only one of the two that answers the question a group-level table is asked.
The confusion is easy to make because both are called “the standard deviation” in different sentences of the same report, and it always runs in the same direction: σ is large and precisely estimated, τ is smaller and roughly estimated, and quoting the first where the second belongs makes group differences look both larger and better established than they are.
A concrete case from these figures: σ = 3 and τ = 1. The observations within a group scatter three times as widely as the groups’ true values do. Anybody reading the raw spread of the data as evidence about how different the groups are would be out by a factor of three, in the direction that manufactures differences.
The estimator real software uses
The moment estimator above is used throughout this field because it is one line, has a closed form, and its bias can be stated exactly. It is not what a mixed-model package fits.
Standard software maximises a likelihood — usually the restricted likelihood, REML, which is maximum likelihood applied to a set of contrasts chosen so the variance estimate is not dragged downward by having fitted the means first. That correction is the same one that puts an n − 1 in a sample variance, generalised.
Three things are worth carrying from that difference.
The answers are close but not identical, and they differ most where there are fewest groups — which is where every estimate of τ is worst anyway. Two packages reporting different variance components on the same data are usually reporting different estimators rather than disagreeing.
Both can return zero. A REML fit reporting a variance component of exactly zero is the same event as the moment estimator’s clamp, it happens for the same reason, and it is routinely treated as a convergence failure to be worked around by refitting with a different optimiser. It is not a convergence failure. It is an answer.
Neither is unbiased for τ. Unbiasedness for a variance does not survive a square root, so every route to a population spread understates it slightly, and every hierarchical model therefore pools marginally harder than the truth warrants.
The objection about using the data twice
The standard objection to empirical Bayes is that the prior is estimated from the same data that is then updated by it, so the data is used twice and the resulting intervals are too narrow.
The objection is correct and its size is knowable. A hierarchical estimate that treats τ̂ as though it were τ ignores the uncertainty in τ̂, so the reported uncertainty in each group’s estimate is smaller than the true one — most severely when there are few groups, because that is when τ̂ is worst.
There are two responses and they are honest ones rather than dismissals.
Report what it costs. The gap between the idealised and estimated cases is the measurement above: 0.88 against 1.05, which is what ignoring the uncertainty in τ̂ buys and what discovering τ costs.
Or put a distribution on τ as well, which is what a full Bayesian treatment does: rather than estimating the population spread and conditioning on it, give it a prior of its own and integrate over it. The estimates barely move; the intervals widen, correctly, and most where there are fewest groups.
What neither response does is return to asserting a prior. The choice is not between an estimated prior and no prior — a model with no prior is a model with τ = ∞, which is no pooling, which is a choice about the population made with no evidence at all.
What this changes about the word “prior”
The Bayesian field’s essays argue that a prior is a component with a measurable effect rather than a philosophical position. This field can say something stronger, and it is the reason the two fields belong beside each other.
Where there are groups, the prior is data. It is not a belief, not a summary of previous literature, and not a regularisation constant chosen by cross-validation. It is the observed distribution of the other groups, with their measurement error subtracted, and it can be reported, checked and argued about like any other estimate.
That does not make every prior an estimate. A single-group problem has no population to read one off, and there the objections about where the prior came from apply in full. But an enormous share of the places where priors are actually used in practice — schools, hospitals, batches, sites, sensors, regions, repeated experiments — are group problems, and in all of them the question “where did the prior come from” has an answer that points at a column of the dataset.
Reading one dataset, in order
The arithmetic is short enough to follow end to end on the eight groups these figures use, and doing so makes the parts separable in a way the formula does not.
One. Compute each group’s own mean and each group’s own standard error. The sizes run from three to forty, so the standard errors run from 3/√3 = 1.73 down to 3/√40 = 0.47 — a factor of nearly four between the best- and worst-measured groups, before anything is estimated.
Two. Compute the spread of those eight means around their average. That number contains the real differences and the measurement error together, and on its own it is an overestimate of the real spread every time.
Three. Subtract the average of the squared standard errors. What remains, floored at zero, is τ̂². Here it lands near 1, and it would land near zero if the eight groups had been generated identically.
Four. For each group, form B = se²/(se² + τ̂²) and move that fraction of the way from the group’s own mean to the population’s. The group of three moves 75% of the way; the group of forty moves 18%.
Nothing in that sequence requires an opinion, and every intermediate quantity is reportable. That is the practical content of the claim that the prior is estimated: each step is a number a reader can be shown, and a reader who disagrees with the answer has to disagree with one of them.
What to ask of a fitted model
Three questions, all answerable from standard output, and all of which this essay’s measurements make concrete.
What is τ̂, and what is it worth in observations? σ̂²/τ̂² is the number of observations the population contributes to each group, and it converts a variance component into something a reader can weigh against the group sizes in the table.
How many groups is it estimated from? Eight or more and the estimate is doing real work; three or four and the model is closer to asserting a weight than measuring one.
Did it come out at zero? A τ̂ of exactly zero is not a numerical failure and should not be worked around. It is the model reporting that the groups are indistinguishable, and the correct response is to say so rather than to force a small positive value that produces prettier group-level estimates.
The third question has the same shape as the one the multiple-comparison field asks about a correction that removes every finding: an uncomfortable answer produced by a correct procedure is a result, and the pressure to refit until the output looks like a finding is exactly the pressure that forking paths are made of.
What is left
This field has one essay to go, and it is about the assumption every measurement here rests on: that the groups are exchangeable, that any of them could have been any other.
When that is false, the machinery does not report a problem. It produces an estimate for the group that is genuinely different, computed by pulling it towards a population it never came from, and the error is a factor of six rather than a few per cent. Every other number in this field is an argument for pooling; that one is the boundary of the argument, and it is measured the same way.
What links here
Computed from the collection, not written here: the essays that point at this one.
Reads more easily once this is understood
Essays that name this one as worth reading first.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- One population, or two — both name empirical bayes, method of moments, partial pooling, variance components
- The fewest groups that can borrow — both name empirical bayes, method of moments, partial pooling, variance components
- Where the borrowing goes — both name empirical bayes, partial pooling, variance components
- A group from the population's own tail — both name empirical bayes, partial pooling
- A level with two units — both name method of moments, variance components
- Estimates that are too alike — both name empirical bayes, partial pooling
Named objects
A flat tag is an object no other essay names yet.
Empirical BayesMethod of momentsPartial poolingPriorVariance components