Hierarchy past one number

A level with two units

A variance estimated from two units is a scaled chi-square on one degree of freedom. Its interquartile range spans a factor of thirteen, its ten-to-ninety range a factor of a hundred and seventy-one, and it comes out exactly zero on 26.7% of studies — so the design effect it decides runs from 1.00 to 7.01 against a truth of 4.69.

Worth reading first: The slope that borrows.

A two-site trial. Two manufacturing lines. Two laboratories, two cohorts, two years. The design has a level, the level has two units, and every hierarchical model fitted to it estimates a variance from them.

That estimate has a sampling distribution, it is available in closed form, and what it says is that nothing downstream of it means anything.

What a variance estimated from K units is worthThe between-unit mean square is a scaled chi-square on K − 1 degrees of freedom, so the estimator's whole distribution is decided by the number of units. At two units its interquartile range spans a factor of 13.03 and its ten-to-ninety range a factor of 171.3, and it comes out exactly zero on 26.7% of studies. The closed form and 3,000 simulated studies agree to 0.051 at every quantile.012311.58234.325.91units at the levelthe variance estimate, against a truth of 1.0the truth, 1.0spans ×171 at two units3,000 studies a point, bands at 10–90 and 25–75the shape is a chi-square, and K − 1 decides it
Fig. 1 The variance estimate against the number of units it is computed from, with the truth at 1.0. At two units its interquartile range spans a factor of 13.03 and its ten-to-ninety range a factor of 171.3, and it comes out exactly zero on 26.7% of studies.

One degree of freedom

The whole result is one line of distribution theory and it needs no simulation.

With KK units of mm observations each, the between-unit mean square is

MSbetween(ω2+σ2m)χK12K1\text{MS}_{\text{between}} \sim \Bigl(\omega^{2} + \frac{\sigma^{2}}{m}\Bigr)\cdot \frac{\chi^{2}_{K-1}}{K-1}

exactly, for normal data. The estimate of the unit variance is that minus σ2/m\sigma^{2}/m, clamped at zero. So the estimator’s entire sampling distribution is a scaled chi-square, and its shape is decided by K1K - 1 and by nothing else.

At K=2K = 2 that is one degree of freedom, and a chi-square on one degree of freedom is a remarkably spread-out object:

units degrees of freedom interquartile range spans ten-to-ninety spans
2 1 ×13.03 ×171.3
3 2 ×4.82 ×21.9
4 3 ×3.39 ×10.7
8 7 ×2.12 ×4.2
20 19 ×1.56 ×2.3
60 59 ×1.28 ×1.6

The middle half of the distribution spans a factor of thirteen. Not the tails — the middle half. An estimate from two units that comes back at 0.4 and one that comes back at 5.2 are both ordinary draws from the same truth.

The interquartile range is the right statistic to quote here and it is worth saying why, because a standard error would be the usual one and would be misleading. The estimator’s distribution is heavily skewed — the mean of a chi-square on one degree of freedom is 1 and its median is 0.455 — so a standard error describes a symmetric spread that the quantity does not have, and quoting one would suggest the estimate is usually within a standard error of the truth. It is usually below the truth, and occasionally far above.

That is the reason to report quantiles rather than a variance wherever a distribution is this skewed, and it is why the table above is quantile ratios rather than standard deviations.

Both routes

A closed form with no count beside it is a claim nobody has tested, so it is tested here.

Three thousand studies at each size, with the estimate computed exactly as an analysis would compute it, and the ten-to-ninety quantiles compared against the chi-square. The two agree to 0.051 at every quantile of every size, sharing only the four parameters.

The agreement matters more here than usual because the closed form is doing the arguing. A simulation could say the spread is large; only the chi-square says why, and the why is what generalises to every variance component estimated by differencing mean squares, at any level, in any design.

Why one degree of freedom is so much worse than two

The table’s first row is far away from its second, and the reason is worth a paragraph because it explains why “two units” is a category rather than a point on a scale.

A chi-square on kk degrees of freedom is a sum of kk squared normals, and summing is what makes a distribution concentrate. At k=2k = 2 the sum of two squares already has a mode away from zero; at k=1k = 1 it does not — the density of χ12\chi^{2}_{1} is infinite at zero and falls monotonically, so the most likely values of the mean square are the smallest ones.

That is the shape behind every number in the table. The median of χ12\chi^{2}_{1} is 0.455 against a mean of 1, so the estimate is more often below the truth than above it and the ones above it are sometimes far above. The interquartile range runs from 0.102 to 1.323 — thirteen times — because the lower quartile is pressed against a density that is rising as it approaches zero.

One degree of freedom is not a small number of degrees of freedom; it is a different kind of distribution, and the qualitative break between the first two rows of the table is that break rather than a continuation of the trend.

What everything downstream inherits

The variance is not an output. It is an input to three quantities a study reports, and each is a deterministic function of it.

What 2 units decide, and how much each is known. the variance itself: 0.000 at the tenth percentile of the variance estimate, 0.390 at the median and 2.900 at the ninetieth, against a true 1.000; the intraclass correlation: 0.000 at the tenth percentile of the variance estimate, 0.213 at the median and 0.668 at the ninetieth, against a true 0.410; the design effect: 1.000 at the tenth percentile of the variance estimate, 2.918 at the median and 7.014 at the ninetieth, against a true 4.689; effective observations: 20.000 at the tenth percentile of the variance estimate, 6.854 at the median and 2.851 at the ninetieth, against a true 4.266. Each is a deterministic function of the same estimate, so the spread is the estimate's spread carried through.
Fig. 2 The same estimate carried through. At the tenth percentile of the variance estimate the intraclass correlation is 0.000, the design effect is 1.000 and the study is worth 20.00 independent observations. At the ninetieth they are 0.668, 7.014 and 2.851. The truth is 0.410, 4.689 and 4.266.

The design effect runs from one to seven, and one is not a small design effect. A design effect of one says the level has no consequence at all — the observations are independent, the clustering can be ignored, the study is worth its full twenty observations. A design effect of seven says it is worth under three.

So the same study, analysed the same way, reports somewhere between “the clustering does not matter” and “the study is a seventh of its nominal size”, depending on which draw from a chi-square on one degree of freedom it happened to get.

And the study cannot tell which it got. There is no internal diagnostic: the estimate is a number, the number is used, and the analysis proceeds.

What 3 units decide, and how much each is known. the variance itself: 0.000 at the tenth percentile of the variance estimate, 0.638 at the median and 2.444 at the ninetieth, against a true 1.000; the intraclass correlation: 0.000 at the tenth percentile of the variance estimate, 0.307 at the median and 0.629 at the ninetieth, against a true 0.410; the design effect: 1.000 at the tenth percentile of the variance estimate, 3.763 at the median and 6.663 at the ninetieth, against a true 4.689; effective observations: 30.000 at the tenth percentile of the variance estimate, 7.972 at the median and 4.502 at the ninetieth, against a true 6.399. Each is a deterministic function of the same estimate, so the spread is the estimate's spread carried through.
Fig. 3 Three units rather than two. The design effect runs from 1.00 to 6.66 — still reaching one at the low end, because the estimate still collapses to zero on a share of studies — and the median has moved most of the way towards the truth. One extra unit is the largest single improvement available anywhere in this account.

Zero, again

The most common single outcome at two units is the boundary.

The estimate is a difference of two positive quantities and comes out negative — hence zero — on 26.7% of studies with a genuine variance of 1.0 in them. That is the atom the two-level field met at eight groups, one level up and much larger: at eight units the rate is a fifth of a per cent.

What a zero does here is worse than what it does there. In a two-level model a collapsed τ^\hat\tau pools the groups completely, which is a defensible answer. A collapsed top-level variance in a nested design removes the level: the intraclass correlation is zero, the design effect is one, and the analysis proceeds as though the study had been a simple random sample.

So on a quarter of two-site studies the output says the sites are interchangeable. No warning is produced, and the statement is a property of there having been two of them.

What 20 clusters of 20 correlated observations do to a 95% interval. Each study has 400 observations arranged as 20 clusters of 20. The lower points are the counted coverage of the usual interval, which treats them as 400 independent observations; the curve through them is 2Φ(1.96/√deff) − 1 with deff = 1 + 19ρ, computed before any data was drawn. At ρ = 0.81 the interval covers 36% rather than 95%. The upper points treat the cluster as the unit and need no variance components at all.
Fig. 4 What the design effect decides, from two levels at once: counted coverage of an interval that ignores the clustering, against the intraclass correlation, with the closed form through it. Everything on that curve is read at a value of ρ\rho, and the question here is how well that value is known when the level has two units.

What the interval does

The design effect’s range is abstract; what it does to the interval a study reports is not.

The overall mean’s interval is scaled by deff\sqrt{\text{deff}}, so a design effect running from 1.00 to 7.01 is an interval whose width runs from its naive value to 2.65 times it. The same twenty observations, the same two sites, the same model — and the reported precision varies by a factor of two and a half according to which draw from a one-degree-of-freedom chi-square the study happened to make.

The direction of the error is the part that matters. On the studies where the estimate collapsed, the interval is the naive one and covers as badly as ignoring the clustering altogether does — which, at an intraclass correlation of 0.41 and ten observations a site, is under half the time. On the studies where the estimate came in high the interval is more than twice as wide as it needs to be and covers far more than it claims.

Neither of those is a small error and they point in opposite directions, so they do not average out into a study that is roughly right. A collection of two-site studies contains some that are far too confident and some that are far too cautious, and nothing in any of them says which it is.

What the level is for, and whether it is doing it

The uncomfortable conclusion is worth stating plainly rather than hedged.

A level with two units cannot estimate its own variance. The estimate exists, it is unbiased for the variance on the variance scale, and its distribution is so wide that no function of it carries information a reader could act on. Fitting the level is not wrong; it is empty, and it is empty in a way the output does not show.

That leaves three honest responses and they are all decisions rather than analyses.

Treat the two units as fixed. Two sites are not a sample from a population of sites; they are two sites. A fixed effect for the site estimates the difference between them — which is a well-determined quantity, one degree of freedom spent on one contrast — and makes no claim about a population. This is what the multilevel field’s own rule prescribes: use a level when the units are a sample and the question is about the population, and a covariate or a fixed effect when the grouping is a small fixed set whose members are the question.

Supply the variance from outside. A prior, a value from a previous study, a regulatory assumption — and say so. This is honest in a way an estimate from two units is not, because the source of the number is stated.

Or integrate over it. The fully Bayesian answer has no atom at zero, produces intervals that cover, and its prior does visible work at two units — which is the correct report, because the data cannot determine the quantity and a method that says so is telling the truth. The field next door measures what that costs at eight groups: 31% more width for sixteen points of coverage. At two units the prior is doing nearly all the work and the width reflects it.

What 8 units decide, and how much each is known. the variance itself: 0.316 at the tenth percentile of the variance estimate, 0.884 at the median and 1.829 at the ninetieth, against a true 1.000; the intraclass correlation: 0.180 at the tenth percentile of the variance estimate, 0.380 at the median and 0.560 at the ninetieth, against a true 0.410; the design effect: 2.621 at the tenth percentile of the variance estimate, 4.423 at the median and 6.036 at the ninetieth, against a true 4.689; effective observations: 30.519 at the tenth percentile of the variance estimate, 18.086 at the median and 13.254 at the ninetieth, against a true 17.063. Each is a deterministic function of the same estimate, so the spread is the estimate's spread carried through.
Fig. 5 The same picture at eight units, for scale. The design effect now runs from 2.62 to 6.04 against a truth of 4.69 — still a factor of two and a quantity a reader can use, which is the difference between an estimate and a number.

The same shape at every level of every hierarchy

The result generalises further than the two-site trial, and the generalisation is one sentence.

A variance component is estimated from the units at its own level, and a hierarchy narrows as it rises: a three-level study with sixty observations in twelve groups in three clusters estimates its observation variance from sixty things, its group variance from twelve and its cluster variance from three. The degrees of freedom for the three are 48, 9 and 2.

So the top-level variance is always the worst-determined quantity in the model, always by a wide margin, and always the one that decides the design effect and therefore every interval. That is exactly what the two-level field found from the other end when it measured a top-level component reported as exactly zero on 18.9% of three-level datasets with a real spread in them.

The pattern is worth naming because it recurs wherever variance components are fitted. A variance estimated by subtraction has a probability of being reported as absent, and that probability is largest exactly where the number of units is smallest — which is the top of every hierarchy, because a hierarchy narrows as it rises.

Where the line is

The table above gives the crossing, and it is earlier than the arithmetic suggests.

Between two units and four the interquartile range falls from ×13.03 to ×3.39 — most of the improvement in the whole table happens in the first two units added. Between eight and sixty it falls from ×2.12 to ×1.28, which is an improvement nobody would restructure a study for.

The useful reading is that the first few units are worth enormously more than the rest, which is the opposite of the usual sample-size intuition and follows from the degrees of freedom being K1K-1 rather than KK. Going from two units to four triples the degrees of freedom; going from twenty to forty doubles them.

That is a design statement and it is actionable. A study with two sites and a large budget is better spent on a third and a fourth site than on more subjects at the two, by a margin the table quantifies — and the argument is the same one two levels at once makes about clusters, arriving at the level above and with a sharper slope.

What a variance estimated from K units is worth. The between-unit mean square is a scaled chi-square on K − 1 degrees of freedom, so the estimator's whole distribution is decided by the number of units. At two units its interquartile range spans a factor of 13.03 and its ten-to-ninety range a factor of 171.3, and it comes out exactly zero on 12.1% of studies. The closed form and 3,000 simulated studies agree to 0.045 at every quantile.
Fig. 6 And with six times as many observations in each unit. The variance estimate’s spread at two units is unchanged — the chi-square has one degree of freedom whatever mm is — and all that moves is the subtraction, which is smaller. The estimate collapses to zero on 12.1% of studies rather than 26.7%, and its median rises from 0.390 to 0.453 — a real improvement in the subtraction and nothing at all in the spread. More data inside the units does not help, which is the sharpest version of the whole finding.

What a reader of a two-site study should ask

The findings translate into three questions a reader can put to a report, and none of them needs the data.

How many units are at the top level? It is almost never stated prominently — a paper says “a multi-centre trial” and the number of centres is in a table — and it is the single number that decides whether the top-level variance means anything.

Was the top-level variance reported as zero? At two units it will have been on a quarter of such studies, and a report of zero between-site variance from two sites is not a finding. It is the estimator at its boundary, and the correct reading is that the study had two sites.

And is the level a sample or a set? Two sites chosen because they were the two available are not a sample from a population of sites, and a random-effects model for them is making a claim about a population the study never drew from. The fixed-effect analysis answers the question the study can answer; the random-effects analysis answers a question with one degree of freedom of evidence behind it.

That third question is the one that decides, and it is not statistical. It is about what the two units were, and the answer is in the methods section rather than in the model.

What is claimed here, and what is not

The claim is what a variance component estimated from few units is worth: that its sampling distribution is a scaled chi-square on K1K-1 degrees of freedom, so at two units the interquartile range spans a factor of 13.03 and the ten-to-ninety range 171.3; that the closed form and three thousand simulated studies agree to 0.051 at every quantile of every size; that the estimate is exactly zero on 26.7% of two-unit studies with a genuine variance in them; and that the design effect it decides runs from 1.00 to 7.01 between the tenth and ninetieth percentiles against a truth of 4.69.

The chi-square result is exact for normal data. Every rate and quantile beside it is three to four thousand studies.

What stays out: restricted maximum likelihood, which has the same one degree of freedom and a different small-sample behaviour at the boundary; the Satterthwaite and Kenward–Roger corrections, which are the standard repair for the degrees of freedom of a test in this situation and do nothing about the variance estimate itself; non-normal data, where the chi-square is an approximation and the spread is generally larger rather than smaller; and unbalanced designs, where the mean squares are weighted sums of chi-squares and the closed form is a Satterthwaite approximation rather than an identity.

Still open: what a two-unit study should report

What is established above is that a variance from two units is uninformative, and three responses are named without measuring any of them. The measurement that would settle the choice is a comparison on the quantity studies actually report — the interval for the overall mean — across the three: a fixed effect for the unit, a supplied variance, and an integrated one.

Each makes a different claim about what the two units represent, so the comparison is not purely technical: an interval from a fixed-effect analysis is a statement about these two sites and an interval from a random-effects one is a statement about sites in general, and they are answers to different questions that happen to be printed in the same place.

The check, and the refusal

Three claims are gated. That the closed form and the count agree at every quantile of every size, within the simulation’s own error — with the tolerance stated relative to each quantile’s size, because a quantile in the tail of a chi-square on one degree of freedom is estimated with a spread that grows with it, and an absolute tolerance passed at ten observations a unit and failed at four. That the interquartile range at two units spans more than a factor of ten. And that the spread falls monotonically as units are added, which would catch a quantile computed on the wrong degrees of freedom.

The refusal is the reading that would make this a footnote: the estimate is merely noisy at two units, rejected. The ten-to-ninety span at two units is required to be more than twenty times what it is at sixty. It is a hundred and seven times. A quantity whose middle half spans a factor of thirteen is not a noisy estimate of anything; it is a draw from a distribution whose shape the data barely constrains, and the refusal is what separates the two readings.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Degrees of freedomDiscretenessEffective sample sizeHierarchical modelMethod of momentsRandom-effectsSample sizeStudy designVariance components