Two levels at once
Worth reading first: The slope that borrows · The observations that repeat each other.
Patients inside wards inside hospitals. Pupils inside classes inside schools. Measurements inside runs inside batches. Three levels rather than two, and the model gains one line:
y = μ + (cluster effect) + (group effect) + (observation noise)
with three variances instead of two. Nothing about the estimation changes: each variance is still a difference of mean squares, each is still clamped at zero, and each still says how much of the total variation belongs to its level.
What is worth the essay is not the third level. It is the quantity that comes out of it.
The design effect
Two observations from the same cluster are correlated. Call that correlation ρ — the intraclass correlation, the share of the total variance that belongs to the cluster level. The variance of a mean of m such observations is not σ²/m; it is
σ²/m · (1 + (m − 1)ρ)
and the factor in brackets is the design effect. It says by how much the variance of the study’s estimate exceeds what an independent sample of the same size would give. Divide the sample size by it and the result is the number of independent observations the study is actually worth.
At m = 20 and ρ = 0.41 the design effect is 8.79, so four hundred clustered observations carry the information of forty-six independent ones. That is not a correction of a few per cent. It is a factor of nine.
Counted against a closed form
The design effect is a statement about a variance, and a variance is not something a reader can check. Its consequence for an interval is.
An interval built as though the observations were independent is too short by a factor of √deff, so it covers 2Φ(1.96/√deff) − 1 instead of 95%. That expression involves no data at all — only ρ and the cluster size — and it can be written down before the study exists.
Against three thousand simulated studies:
| intraclass correlation | 0.410 |
| design effect at 20 per cluster | 8.79 |
| counted coverage, ignoring the clustering | 49.3% |
| predicted from the design effect | 49.2% |
| counted coverage, cluster as the unit | 94.0% |
A tenth of a percentage point apart, from two calculations sharing nothing: one a tally over three thousand studies, the other a normal distribution function evaluated once.
The last row is the repair and it is embarrassingly simple. Take the mean of each cluster, treat those twenty numbers as the data, and form the ordinary interval from them. No variance components are needed, no design effect has to be estimated, and the twenty cluster means are independent by construction. The information lost by throwing away the within-cluster structure is exactly the information the design effect says was never there.
Why 94.0% rather than 95.0%
The cluster-level interval covers 94.0%, not 95%, and the shortfall is not noise — three thousand studies give a standard error of about 0.4 percentage points, so a full point is more than two of them.
It is the same defect the t that fixes a small sample is about, at the cluster level. The interval above uses 1.96, which is the normal quantile; with twenty clusters the right multiplier is the t quantile on nineteen degrees of freedom, 2.093. Using the normal where a t is needed produces an interval about 7% too short, and 94.0% is what that costs.
That is worth stating because it is the second time in this essay that the effective sample size has turned out to be the number of clusters rather than the number of observations. It sets the width of the interval, and it sets the degrees of freedom too. A study of four hundred patients in twenty hospitals is, for every purpose that matters here, a study of twenty.
The identity with the time-series field
Here is the part worth carrying elsewhere.
The observations that repeat each other computes, for a series with lag-one correlation φ, an effective sample size of
n(1 − φ)/(1 + φ)
and measures what happens to an interval that ignores it: at φ = 0.8 and n = 50, a 95% interval covers 47.0%.
The two quantities look unrelated. One is about time order and an autoregressive process; the other is about nesting and variance components. They are the same quantity.
Both are the variance of a mean divided by the variance a mean of independent observations would have had. For the clustered case that ratio is 1 + (m − 1)ρ, because every pair inside a cluster contributes ρ and there are m(m − 1) ordered pairs. For the series it is the same sum over pairs with the correlation falling off geometrically, which is what produces (1 + φ)/(1 − φ) in the limit. Neither derivation mentions the other’s picture and both are counting the same covariances.
The practical value of noticing is that the intuition transfers in both directions. A reader who understands why fifty autocorrelated observations are worth six understands why four hundred clustered ones are worth forty-six. And a reader who knows that a clustered design is repaired by analysing cluster means knows what the corresponding repair for a series is: analyse block means, with blocks long enough to be nearly independent.
Blocking is the same quantity again, used deliberately
There is a third place this arithmetic already appears on the site, and in that one it is an advantage rather than a cost.
The variance removed before the data shows blocking taking the between-block variance out of a comparison, and computes the gain as σ²/(σ² + β²). That is 1 − ρ, with ρ the intraclass correlation of this essay. The same decomposition, the same two numbers, used twice in opposite directions.
The difference is which comparison is being made. A cluster-level question — what is the overall mean, does the treatment differ between hospitals — has to carry the cluster variance and pays the design effect for it. A within-cluster question — does the treatment differ between two patients in the same ward — has the cluster variance cancel, and gains exactly what the other loses.
That is why the same structure is a disaster for one study and a gift for another, and why “is the data clustered” is not a question with a single consequence. The question is whether the contrast of interest lives inside the clusters or across them. A trial randomising whole hospitals to arms pays the full design effect; a trial randomising patients within each hospital pays none of it and gains the blocking.
The variance components themselves
The three variances are recovered by differencing mean squares, and they are recovered accurately when there is enough at each level. Over four hundred datasets of sixty clusters of four groups of five, against a truth of 1.0, 0.7 and 1.2:
| level | truth | recovered |
|---|---|---|
| between clusters | 1.0 | 0.998 |
| between groups | 0.7 | 0.700 |
| within groups | 1.2 | 1.201 |
Three digits on all three, which is what a moment estimator should do with sixty clusters.
With eight clusters it does something else. The top-level component is reported as exactly zero on 18.9% of datasets generated with a real cluster spread of 0.5 in them — the atom from when the spread estimates to zero, one level up.
The consequence is worse here than it is for a two-level model, and the reason is what a zero at the top does. A cluster variance of zero makes the intraclass correlation zero, which makes the design effect one, which makes the clustering have no consequence for any standard error in the analysis. The study is then reported as though it had been a simple random sample of four hundred, with intervals √8.79 ≈ 3 times too short — and the report contains no indication that a level was dropped.
That is the failure mode this whole phase keeps meeting: the machinery returns the boundary of its own range, the boundary means “this thing is not there”, and nothing distinguishes it from the thing genuinely not being there.
What this means for designing one
The three figures together give the design rule, and it is not the one sample-size arithmetic usually produces.
Adding observations inside a cluster buys very little. The effective sample size is Km/(1 + (m − 1)ρ), which as m grows approaches K/ρ — a hard ceiling set by the number of clusters and the intraclass correlation, no matter how many observations are collected. At ρ = 0.41 and twenty clusters, that ceiling is forty-nine effective observations. The study measured four hundred and could not have exceeded forty-nine by measuring four thousand.
Adding clusters buys everything. K appears linearly and outside the ceiling. Twenty clusters of twenty and forty clusters of ten are the same four hundred observations and not the same study: the second is worth 40·10/(1 + 9·0.41) = 84 effective observations against the first’s 46.
That is the same shape as how many subjects arrives at for power, and it is worth putting the two together.
A power calculation that ignores clustering asks for a sample size that will be missed by the design effect, which is a factor rather than a percentage. Sixty-four per arm gives 80% power at half a standard deviation; delivered as eight clusters of eight at ρ = 0.41, those sixty-four observations are worth eighteen, and the power is nearer 30%.
The repair is standard and is worth stating in the form that makes it hard to get wrong: compute the sample size as though the observations were independent, then multiply by the design effect. Not add a margin, not round up — multiply, by a number that at ρ = 0.41 and clusters of twenty is nearly nine. And since ρ has to be guessed before the study, guess it high: the cost of over-recruiting is linear and the cost of under-recruiting is a trial that cannot answer its question.
What one more observation is worth, in each of the two directions
The design rule above is stated as a contrast between two things that buy very little and a great deal, and the ratio between them has a closed form worth carrying, because it is a single number an experimenter can hold.
The effective sample size is Km/(1 + (m − 1)ρ), so the two derivatives are the two ways of spending an observation. Adding one observation to every cluster spends K observations and raises the effective sample size by K(1 − ρ)/deff². Adding one new cluster of m spends m observations and raises it by m/deff. Divide the per-observation returns and everything cancels except
deff / (1 − ρ)
At m = 20 and ρ = 0.41 that is 8.79/0.59 = 14.9. An observation added inside an existing cluster is very nearly fifteen times more expensive, per unit of information, than the same observation recruited as part of a new one. In the units the study was measured in: a new cluster of twenty buys 2.28 effective observations, and a twenty-first patient in each of the twenty existing wards buys 0.15.
The expression says why the ceiling exists rather than merely that it does. The numerator grows with the cluster size and the denominator does not, so the penalty for growing clusters instead of adding them compounds — it is not a fixed inefficiency to be tolerated but one that gets worse exactly as the study gets larger in the wrong direction. And it collapses correctly at both ends: at ρ = 0 the design effect is one, the ratio is one, and the two ways of spending an observation are identical, which is the independent case arriving as a special case rather than as an exception.
How wrong the guess about ρ can afford to be
ρ has to be guessed before the study, and the design effect is linear in it with a slope of m − 1. That makes the consequence of guessing arithmetic rather than a matter of judgement.
Guess ρ = 0.30 where the truth is 0.41, in clusters of twenty. The design effect used is 1 + 19(0.30) = 6.70 against a true 8.79, so the recruitment multiplier is 24% short — and since the repair is a multiplication rather than a margin, 24% short on the multiplier is 24% short on the whole study. The same eleven-hundredths of error in clusters of five gives 2.20 against 2.64, which is 17% short. The guess did not get better; the clusters got smaller.
That is the second reason to prefer many small clusters over few large ones, and it is independent of the first. Large clusters raise the design effect, which is a cost known in advance and can be recruited against. Large clusters also raise the sensitivity of the design effect to the one input nobody knows, at a relative rate of (m − 1)/deff per unit of ρ — 2.16 at clusters of twenty against 1.52 at clusters of five. A design that is expensive is manageable; a design whose cost is uncertain in proportion to its size is not, which is the same reasoning the design that hedges applies to a guessed parameter in a quite different setting.
The practical form of that is the instruction already given, with a reason attached. Guess ρ high, and prefer the arrangement in which being wrong about it matters least — because the quantity being guessed enters multiplied by m − 1, and m is the one thing in the whole calculation the experimenter sets directly.
Where the third level earns its keep
Nothing so far has needed three levels — the design effect only involves the cluster and the observation. The middle level earns its place when a question is asked about it.
A ward’s estimate in a three-level model is shrunk towards its own hospital’s estimate, and the hospital’s estimate is shrunk towards the overall one. A ward in a hospital whose other wards are all performing well is therefore pulled up, and a ward in a poorly performing hospital is pulled down, which is the correct treatment of the fact that wards in one hospital share staffing, catchment and management.
Compare that with the two-level alternative of ignoring the hospitals and shrinking every ward towards the global mean. That treats two wards in the same hospital as no more related than two wards in different countries. It will over-shrink the wards of a genuinely unusual hospital, all in the same direction, and it will report the hospital’s own effect as nothing at all — because the hospital is not in the model.
The rule for choosing between them is not statistical sophistication. Use a level when the units at that level are a sample from a population of such units and the question is about the population; use a covariate when the grouping is a small fixed set whose members are the question. Twenty hospitals out of a health service are a sample; two treatment arms are not.
What the level actually cost
Nothing about the arithmetic. The third variance component is estimated the same way as the second, the shrinkage weight at each level is the same expression, and a group inside a cluster is shrunk towards its cluster’s estimate exactly as a group was shrunk towards the population’s.
What it cost is the top of the hierarchy. Every level added narrows: sixty clusters of four groups of five has three hundred observations at the bottom, sixty at the middle and one number’s worth of information at the top about how much clusters vary. The estimate of the top-level variance is the worst-determined quantity in the model, it is the one with the atom at zero, and it is the one that decides the design effect and therefore every interval.
The repair is the one this site has arrived at three times from three directions. Do not estimate the top-level variance and substitute it. Integrate over it — or, where that is more machinery than the question deserves, sidestep it entirely by taking the cluster as the unit of analysis, which needs no variance components at all and covers within a percentage point of what it claims.
The second option deserves more respect than it usually gets. Analysing cluster means throws away information in the sense that it ignores everything about the within-cluster structure, and the design effect is the statement that most of that information was not there. What it buys is a procedure with no variance components in it, no atom at zero, no design effect to estimate, and a coverage of 94% that becomes 95% by using a t rather than a normal. For a question about the overall mean, that is a better trade than it looks — and the whole of this field’s difficulty lives in the machinery it avoids.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The check before the standard error — both name correlation, coverage, dependence, sample size
- The fewest groups that can borrow — both name hierarchical model, method of moments, sample size, variance components
- Where the borrowing goes — both name hierarchical model, sample size, study design, variance components
- A block size that changes — both name blocking, coverage, sample size
- A schedule that reads the mean — both name blocking, coverage, sample size
- One number for a table of candidates — both name dependence, effective sample size, sample size
Named objects
A flat tag is an object no other essay names yet.
BlockingCorrelationCoverageDependenceEffective sample sizeHierarchical modelMethod of momentsSample sizeStudy designVariance components