Concept

Method of moments — where it appears

Estimating parameters by matching sample moments to their theoretical expressions and solving. It is simple, it needs no likelihood, and it can return estimates outside the parameter space — a negative variance component is the standard example.

Named by 6 essays across 3 fields — each of them below, with the objects they name alongside it.

τ̂ across 2,000 datasets of 12 groups, true τ = 1.5. The population spread is not supplied to a hierarchical model — it is estimated from how far apart the group means are, after subtracting the noise that would separate them anyway. It averages 1.41 here against a true 1.5, and comes out exactly zero on 5% of datasets.

The prior the data estimates

A hierarchical model needs a population spread, and it does not ask for one. It reads τ off the distance between the group means — biased six per cent low, exactly zero on 53% of datasets where the groups are identical — and the prior stops being a belief.

hierarchical · Prior
How often "there is no spread between the groups" is reported about data that has one. Every dataset here was generated with a real population spread of 1. The moment estimator is the difference between the observed spread and what noise alone would produce, clamped at zero, and the difference comes out negative often: at eight groups it reports exactly zero on 32.6% of datasets, which is an instruction to pool completely and give all eight groups the same estimate. The rate falls to 4.2% at 48 groups.

When the spread estimates to zero

The usual estimate of a population spread is a difference of two positive quantities, clamped at zero. On a third of eight-group datasets with a real spread in them the difference comes out negative, the estimate is exactly zero, and every group is pooled completely on data that said no such thing.

fullbayes · Pooling
What 20 clusters of 20 correlated observations do to a 95% interval. Each study has 400 observations arranged as 20 clusters of 20. The lower points are the counted coverage of the usual interval, which treats them as 400 independent observations; the curve through them is 2Φ(1.96/√deff) − 1 with deff = 1 + 19ρ, computed before any data was drawn. At ρ = 0.81 the interval covers 36% rather than 95%. The upper points treat the cluster as the unit and need no variance components at all.

Two levels at once

A third level of grouping adds no new arithmetic and produces one number — the design effect — that decides how many independent observations a clustered study is worth. It is the same quantity the time-series field computes for autocorrelated data, arrived at from a completely different picture.

multilevel · Levels
What each estimator costs, against how many groups there are. Each point is 4,000 datasets, with the population spread estimated from the data rather than supplied. Partial pooling first beats BOTH of the estimators it sits between at 5 groups; below that, complete pooling — which estimates nothing at all — is the better answer. Its own cost falls from 1.751 at 2 groups to 0.795 at 40.

The fewest groups that can borrow

At three groups the estimator that shrinks towards its own data's mean returns the group means untouched, on every dataset, because its constant is J − 3. At two it expands instead of shrinking. And the number of groups at which partial pooling starts to be worth doing is five, or two, or never — it depends on how far apart the groups are.

hierarchical · Pooling
What a variance estimated from K units is worth. The between-unit mean square is a scaled chi-square on K − 1 degrees of freedom, so the estimator's whole distribution is decided by the number of units. At two units its interquartile range spans a factor of 13.03 and its ten-to-ninety range a factor of 171.3, and it comes out exactly zero on 26.7% of studies. The closed form and 3,000 simulated studies agree to 0.051 at every quantile.

A level with two units

A variance estimated from two units is a scaled chi-square on one degree of freedom. Its interquartile range spans a factor of thirteen, its ten-to-ninety range a factor of a hundred and seventy-one, and it comes out exactly zero on 26.7% of studies — so the design effect it decides runs from 1.00 to 7.01 against a truth of 4.69.

multilevel · Levels
Twelve groups from two clusters, τ = 1. Every group's truth is in one of two clusters, and the population's total spread is exactly τ = 1, so the analysis recovers τ̂ = 1.515 and every shrinkage weight is what it would be for a single normal population. The estimates are pulled towards the grand mean, which is the middle of the gap — a place 0 of the 12 truths are and 6 of the estimates end up.

One population, or two

Group effects from two clusters rather than one bell, with the same total spread. The analysis recovers the same population spread, uses the same weight for every group, and reports nothing unusual — while 46% of its estimates land in a region holding 6.6% of the truths.

hierarchical · Pooling

Named alongside it

The objects these essays reach for when they reach for this one.

Variance componentsHierarchical modelEmpirical BayesPartial poolingSample sizeShrinkageDegrees of freedomDiscretenessEffective sample sizeRandom-effectsStudy designBlocking

All concepts