A level with two units
Worth reading first: The slope that borrows.
A two-site trial. Two manufacturing lines. Two laboratories, two cohorts, two years. The design has a level, the level has two units, and every hierarchical model fitted to it estimates a variance from them.
That estimate has a sampling distribution, it is available in closed form, and what it says is that nothing downstream of it means anything.
One degree of freedom
The whole result is one line of distribution theory and it needs no simulation.
With units of observations each, the between-unit mean square is
exactly, for normal data. The estimate of the unit variance is that minus , clamped at zero. So the estimator’s entire sampling distribution is a scaled chi-square, and its shape is decided by and by nothing else.
At that is one degree of freedom, and a chi-square on one degree of freedom is a remarkably spread-out object:
| units | degrees of freedom | interquartile range spans | ten-to-ninety spans |
|---|---|---|---|
| 2 | 1 | ×13.03 | ×171.3 |
| 3 | 2 | ×4.82 | ×21.9 |
| 4 | 3 | ×3.39 | ×10.7 |
| 8 | 7 | ×2.12 | ×4.2 |
| 20 | 19 | ×1.56 | ×2.3 |
| 60 | 59 | ×1.28 | ×1.6 |
The middle half of the distribution spans a factor of thirteen. Not the tails — the middle half. An estimate from two units that comes back at 0.4 and one that comes back at 5.2 are both ordinary draws from the same truth.
The interquartile range is the right statistic to quote here and it is worth saying why, because a standard error would be the usual one and would be misleading. The estimator’s distribution is heavily skewed — the mean of a chi-square on one degree of freedom is 1 and its median is 0.455 — so a standard error describes a symmetric spread that the quantity does not have, and quoting one would suggest the estimate is usually within a standard error of the truth. It is usually below the truth, and occasionally far above.
That is the reason to report quantiles rather than a variance wherever a distribution is this skewed, and it is why the table above is quantile ratios rather than standard deviations.
Both routes
A closed form with no count beside it is a claim nobody has tested, so it is tested here.
Three thousand studies at each size, with the estimate computed exactly as an analysis would compute it, and the ten-to-ninety quantiles compared against the chi-square. The two agree to 0.051 at every quantile of every size, sharing only the four parameters.
The agreement matters more here than usual because the closed form is doing the arguing. A simulation could say the spread is large; only the chi-square says why, and the why is what generalises to every variance component estimated by differencing mean squares, at any level, in any design.
Why one degree of freedom is so much worse than two
The table’s first row is far away from its second, and the reason is worth a paragraph because it explains why “two units” is a category rather than a point on a scale.
A chi-square on degrees of freedom is a sum of squared normals, and summing is what makes a distribution concentrate. At the sum of two squares already has a mode away from zero; at it does not — the density of is infinite at zero and falls monotonically, so the most likely values of the mean square are the smallest ones.
That is the shape behind every number in the table. The median of is 0.455 against a mean of 1, so the estimate is more often below the truth than above it and the ones above it are sometimes far above. The interquartile range runs from 0.102 to 1.323 — thirteen times — because the lower quartile is pressed against a density that is rising as it approaches zero.
One degree of freedom is not a small number of degrees of freedom; it is a different kind of distribution, and the qualitative break between the first two rows of the table is that break rather than a continuation of the trend.
What everything downstream inherits
The variance is not an output. It is an input to three quantities a study reports, and each is a deterministic function of it.
The design effect runs from one to seven, and one is not a small design effect. A design effect of one says the level has no consequence at all — the observations are independent, the clustering can be ignored, the study is worth its full twenty observations. A design effect of seven says it is worth under three.
So the same study, analysed the same way, reports somewhere between “the clustering does not matter” and “the study is a seventh of its nominal size”, depending on which draw from a chi-square on one degree of freedom it happened to get.
And the study cannot tell which it got. There is no internal diagnostic: the estimate is a number, the number is used, and the analysis proceeds.
Zero, again
The most common single outcome at two units is the boundary.
The estimate is a difference of two positive quantities and comes out negative — hence zero — on 26.7% of studies with a genuine variance of 1.0 in them. That is the atom the two-level field met at eight groups, one level up and much larger: at eight units the rate is a fifth of a per cent.
What a zero does here is worse than what it does there. In a two-level model a collapsed pools the groups completely, which is a defensible answer. A collapsed top-level variance in a nested design removes the level: the intraclass correlation is zero, the design effect is one, and the analysis proceeds as though the study had been a simple random sample.
So on a quarter of two-site studies the output says the sites are interchangeable. No warning is produced, and the statement is a property of there having been two of them.
What the interval does
The design effect’s range is abstract; what it does to the interval a study reports is not.
The overall mean’s interval is scaled by , so a design effect running from 1.00 to 7.01 is an interval whose width runs from its naive value to 2.65 times it. The same twenty observations, the same two sites, the same model — and the reported precision varies by a factor of two and a half according to which draw from a one-degree-of-freedom chi-square the study happened to make.
The direction of the error is the part that matters. On the studies where the estimate collapsed, the interval is the naive one and covers as badly as ignoring the clustering altogether does — which, at an intraclass correlation of 0.41 and ten observations a site, is under half the time. On the studies where the estimate came in high the interval is more than twice as wide as it needs to be and covers far more than it claims.
Neither of those is a small error and they point in opposite directions, so they do not average out into a study that is roughly right. A collection of two-site studies contains some that are far too confident and some that are far too cautious, and nothing in any of them says which it is.
What the level is for, and whether it is doing it
The uncomfortable conclusion is worth stating plainly rather than hedged.
A level with two units cannot estimate its own variance. The estimate exists, it is unbiased for the variance on the variance scale, and its distribution is so wide that no function of it carries information a reader could act on. Fitting the level is not wrong; it is empty, and it is empty in a way the output does not show.
That leaves three honest responses and they are all decisions rather than analyses.
Treat the two units as fixed. Two sites are not a sample from a population of sites; they are two sites. A fixed effect for the site estimates the difference between them — which is a well-determined quantity, one degree of freedom spent on one contrast — and makes no claim about a population. This is what the multilevel field’s own rule prescribes: use a level when the units are a sample and the question is about the population, and a covariate or a fixed effect when the grouping is a small fixed set whose members are the question.
Supply the variance from outside. A prior, a value from a previous study, a regulatory assumption — and say so. This is honest in a way an estimate from two units is not, because the source of the number is stated.
Or integrate over it. The fully Bayesian answer has no atom at zero, produces intervals that cover, and its prior does visible work at two units — which is the correct report, because the data cannot determine the quantity and a method that says so is telling the truth. The field next door measures what that costs at eight groups: 31% more width for sixteen points of coverage. At two units the prior is doing nearly all the work and the width reflects it.
The same shape at every level of every hierarchy
The result generalises further than the two-site trial, and the generalisation is one sentence.
A variance component is estimated from the units at its own level, and a hierarchy narrows as it rises: a three-level study with sixty observations in twelve groups in three clusters estimates its observation variance from sixty things, its group variance from twelve and its cluster variance from three. The degrees of freedom for the three are 48, 9 and 2.
So the top-level variance is always the worst-determined quantity in the model, always by a wide margin, and always the one that decides the design effect and therefore every interval. That is exactly what the two-level field found from the other end when it measured a top-level component reported as exactly zero on 18.9% of three-level datasets with a real spread in them.
The pattern is worth naming because it recurs wherever variance components are fitted. A variance estimated by subtraction has a probability of being reported as absent, and that probability is largest exactly where the number of units is smallest — which is the top of every hierarchy, because a hierarchy narrows as it rises.
Where the line is
The table above gives the crossing, and it is earlier than the arithmetic suggests.
Between two units and four the interquartile range falls from ×13.03 to ×3.39 — most of the improvement in the whole table happens in the first two units added. Between eight and sixty it falls from ×2.12 to ×1.28, which is an improvement nobody would restructure a study for.
The useful reading is that the first few units are worth enormously more than the rest, which is the opposite of the usual sample-size intuition and follows from the degrees of freedom being rather than . Going from two units to four triples the degrees of freedom; going from twenty to forty doubles them.
That is a design statement and it is actionable. A study with two sites and a large budget is better spent on a third and a fourth site than on more subjects at the two, by a margin the table quantifies — and the argument is the same one two levels at once makes about clusters, arriving at the level above and with a sharper slope.
What a reader of a two-site study should ask
The findings translate into three questions a reader can put to a report, and none of them needs the data.
How many units are at the top level? It is almost never stated prominently — a paper says “a multi-centre trial” and the number of centres is in a table — and it is the single number that decides whether the top-level variance means anything.
Was the top-level variance reported as zero? At two units it will have been on a quarter of such studies, and a report of zero between-site variance from two sites is not a finding. It is the estimator at its boundary, and the correct reading is that the study had two sites.
And is the level a sample or a set? Two sites chosen because they were the two available are not a sample from a population of sites, and a random-effects model for them is making a claim about a population the study never drew from. The fixed-effect analysis answers the question the study can answer; the random-effects analysis answers a question with one degree of freedom of evidence behind it.
That third question is the one that decides, and it is not statistical. It is about what the two units were, and the answer is in the methods section rather than in the model.
What is claimed here, and what is not
The claim is what a variance component estimated from few units is worth: that its sampling distribution is a scaled chi-square on degrees of freedom, so at two units the interquartile range spans a factor of 13.03 and the ten-to-ninety range 171.3; that the closed form and three thousand simulated studies agree to 0.051 at every quantile of every size; that the estimate is exactly zero on 26.7% of two-unit studies with a genuine variance in them; and that the design effect it decides runs from 1.00 to 7.01 between the tenth and ninetieth percentiles against a truth of 4.69.
The chi-square result is exact for normal data. Every rate and quantile beside it is three to four thousand studies.
What stays out: restricted maximum likelihood, which has the same one degree of freedom and a different small-sample behaviour at the boundary; the Satterthwaite and Kenward–Roger corrections, which are the standard repair for the degrees of freedom of a test in this situation and do nothing about the variance estimate itself; non-normal data, where the chi-square is an approximation and the spread is generally larger rather than smaller; and unbalanced designs, where the mean squares are weighted sums of chi-squares and the closed form is a Satterthwaite approximation rather than an identity.
Still open: what a two-unit study should report
What is established above is that a variance from two units is uninformative, and three responses are named without measuring any of them. The measurement that would settle the choice is a comparison on the quantity studies actually report — the interval for the overall mean — across the three: a fixed effect for the unit, a supplied variance, and an integrated one.
Each makes a different claim about what the two units represent, so the comparison is not purely technical: an interval from a fixed-effect analysis is a statement about these two sites and an interval from a random-effects one is a statement about sites in general, and they are answers to different questions that happen to be printed in the same place.
The check, and the refusal
Three claims are gated. That the closed form and the count agree at every quantile of every size, within the simulation’s own error — with the tolerance stated relative to each quantile’s size, because a quantile in the tail of a chi-square on one degree of freedom is estimated with a spread that grows with it, and an absolute tolerance passed at ten observations a unit and failed at four. That the interquartile range at two units spans more than a factor of ten. And that the spread falls monotonically as units are added, which would catch a quantile computed on the wrong degrees of freedom.
The refusal is the reading that would make this a footnote: the estimate is merely noisy at two units, rejected. The ten-to-ninety span at two units is required to be more than twenty times what it is at sixty. It is a hundred and seven times. A quantity whose middle half spans a factor of thirteen is not a noisy estimate of anything; it is a draw from a distribution whose shape the data barely constrains, and the refusal is what separates the two readings.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The fewest groups that can borrow — both name degrees of freedom, hierarchical model, method of moments, sample size, variance components
- One population, or two — both name hierarchical model, method of moments, random-effects, variance components
- Where the borrowing goes — both name hierarchical model, sample size, study design, variance components
- One number for a table of candidates — both name degrees of freedom, effective sample size, sample size
- A block size that changes — both name degrees of freedom, sample size
- A coverage table with its own error — both name discreteness, sample size
Named objects
A flat tag is an object no other essay names yet.
Degrees of freedomDiscretenessEffective sample sizeHierarchical modelMethod of momentsRandom-effectsSample sizeStudy designVariance components