The count that is not the rows
Worth reading first: What a wrong model estimates · The observations that repeat each other.
Three hundred rows, a covariate that is constant inside each cluster, and a within-cluster correlation of 0.1 — a value most analysts would describe as negligible. Put those three hundred rows in five clusters of sixty and the slope’s variance is 6.9000 times what an independent-rows calculation reports. The 95% interval built by counting rows covers 53.42%.
Nothing about the row count changes across the whole of what follows. The clusters move from five of sixty to a hundred of three, and the naive interval’s coverage climbs from 53.42% to 92.83% on the same three hundred rows, the same correlation and the same truth. The number that decides a standard error is the count of independent things, and the row count is not it — which is the point at which the sandwich in this field stops being a variance correction and becomes a statement about what a sample is.
The correction is well known: sum the filling over clusters rather than over rows. What is less often quoted is how it behaves at the cluster counts real studies have. At five clusters the cluster-robust interval covers 74.43% against a normal reference and 85.08% against a on , and it does not reach 94% until a hundred clusters. And the repair everyone quotes for unequal sizes is measurably the wrong formula for a slope, in a way a pair of arrangements makes unmistakable.
One cluster size, and the effect is exact
With the covariate constant inside a cluster, the design effect is not an approximation. Every row in a cluster shares the same covariate value, so the cluster contributes as a single unit whose error is the average of correlated errors, and the arithmetic gives
exactly. At that is 6.9000 at five clusters of sixty, 3.9000 at ten of thirty, 2.4000 at twenty of fifteen, 1.5000 at fifty of six and 1.2000 at a hundred of three. Counted as the variance of six thousand fitted slopes divided by the independent-rows variance, the same five cells read 7.0467, 3.9255, 2.3561, 1.4659 and 1.1973 — two routes, no shared arithmetic, agreeing across a factor of six in the quantity.
The reading that makes the number concrete is its square root. At five clusters 2.6268, so a standard error that counts rows is short by a factor of two and a half before anything else goes wrong. That is the same quantity, computed from a different picture, as the effective sample size in the essay on observations that repeat each other, where a lag-one correlation of 0.8 leaves a fifty-point series worth about six independent observations. Clustering and autocorrelation are one arithmetic wearing two names, and the essay that finds the design effect at a third level of grouping makes the same identification from the multilevel side.
A coverage that knows nothing about clusters
The naive interval’s coverage has a closed form of its own, and it is the cleanest check in this field because it reads nothing about the design at all.
If the interval is with short by a factor of , then the standardised distance to the truth is times too large, and the coverage is
At that is 54.44%, against a counted 53.42% whose binomial standard error is 0.563 points. At the other four cluster counts it reads 67.90%, 79.42%, 89.05% and 92.64% against counted 67.38%, 80.18%, 89.98% and 92.83%. The formula contains no cluster count, no cluster size and no correlation — only the design effect — so agreement across five cells is evidence that the design effect is the whole of what clustering does to this interval, rather than one of several things it does.
That matters for interpreting the failure. The naive interval is not miscentred, its estimate of the slope is fine, and nothing about the fit is wrong. It is a correct interval divided by the wrong number, and the wrong number is a count.
The repair behaves like a sample the size of the cluster count
Summing the filling over clusters is the standard answer, and the sweep prices it rather than recommending it. Against a normal reference the cluster-robust interval covers 74.43% at five clusters, 86.00% at ten, 91.50% at twenty, 93.97% at fifty and 94.20% at a hundred.
So the repair is a repair, and at the cluster counts a lot of applied work actually has — a handful of schools, sites, countries, years — it recovers about two thirds of the shortfall and leaves a 95% interval covering three quarters of the time. The reason is visible in the estimate rather than in the interval: the cluster-robust variance estimate’s average, divided by the variance the slope has, is 0.5699 at five clusters, 0.8025 at ten, 0.9125 at twenty, 0.9868 at fifty and 0.9900 at a hundred. It is a mean of terms, so it inherits the small-sample behaviour of an average of five things, and no amount of rows inside those five clusters changes it. The estimator is exactly as good as a sandwich on observations, which is the honest version of the folk rule about needing forty clusters.
A reference on degrees of freedom is the usual patch and it is worth what it costs. At five clusters that quantile is 2.7764 rather than 1.959964, and the coverage goes from 74.43% to 85.08% — ten points, from a change that costs nothing, which is the same trade the small-sample coverage sweep prices for a reference distribution at twenty rows. It does not close the gap, because the estimate’s own shortfall of 0.5699 is far too large for a quantile to absorb.
The uncorrected cluster sandwich covers 70.00% at five clusters against the finite-sample-corrected version’s 74.43%, which is the same ordering as the row-level corrections in the essay on the four leverage weightings and for the same reason: a sandwich’s filling is short by construction, and a factor slightly above one is the cheapest partial repair.
The same sizes, laid out two ways
The formula quoted for unequal cluster sizes replaces with — the size a randomly chosen observation finds itself in — and reports . That formula is derived for a mean, and this is where it fails.
Take five clusters of sizes 100, 80, 60, 40 and 20, summing to the same three hundred rows. Place the large ones at the ends of the covariate’s range and the design effect is 9.3158. Place the same five sizes with the large ones in the middle and it is 5.4652 — a factor of 1.7046 between two designs that share every quantity the formula reads. Their effective size is 73.333 in both cases and the formula returns 8.2333 for both, over-stating one arrangement by a tenth and under-stating the other by a half.
The intervals follow the truth rather than the formula. The cluster-robust interval covers 60.90% on the edge arrangement and 82.92% on the centre one — 22.02 points apart, on the same rows, the same cluster count, the same sizes and the same correlation.
Why a slope is not a mean
The reason is one line of algebra and it is worth writing out, because it says exactly which quantity the usual formula is missing.
For a mean, each cluster enters the estimate weighted by its size, so the variance inflation is a size-weighted average of sizes and is the right summary. For a slope, each cluster enters weighted by — its size times its squared distance from the covariate mean — because that is what a least-squares slope does with an observation. The design effect is therefore
a -weighted average of each cluster’s own inflation, and cannot see the at all. Putting the large clusters where is largest gives the biggest clusters the biggest weight in that average and the effect goes up; putting them where is smallest does the reverse.
A cluster’s cost is its size times its leverage, which is the same that the field’s account of the leverage corrections computes one observation at a time, arriving here at the level of a group. That is also why the effect of arrangement shrinks as the clusters multiply: at a hundred clusters the edge and centre arrangements read 1.2792 and 1.1431, because sizes of four and two cannot differ by much and the weights are spread across a hundred positions.
The gap is not confined to the five-cluster case, which would make it a curiosity of a small sweep. At ten clusters the two arrangements read 5.1555 and 3.1819 against a formula of 4.4533 for both; at twenty they read 3.0170 and 2.0395 against 2.6567; at fifty, 1.7429 and 1.3601 against 1.5987. The formula sits between the two truths at every cluster count and is wrong about both. The ratio between the two arrangements is 1.70 at five clusters, 1.62 at ten and 1.48 at twenty, and it closes only as the sizes themselves become nearly equal — at a hundred clusters the ladder runs from two to four and there is little left to arrange. So an analyst who applies the size-only repair is not making a small correction in the right direction — the correction’s error is comparable to the correction.
One consequence is worth stating against the intuition it contradicts. Unequal sizes are not uniformly worse than equal ones: the centre arrangement’s design effect at five clusters is 5.4652, below the equal-size 6.9000. Concentrating the rows into clusters that sit near the covariate’s mean concentrates them where the slope reads them least, which is a way of wasting data rather than a way of losing precision. The reflex that unequal cluster sizes are a problem to be corrected is a reflex about means.
The estimate that counts rows reads a seventh of the variance
Coverage mixes several failures, and the naive interval’s can be read as one number. The model-based variance estimate’s average, divided by the variance the slope actually has, is 0.1364 at five clusters, 0.2502 at ten, 0.4206 at twenty, 0.6804 at fifty and 0.8332 at a hundred.
Those are almost exactly the reciprocals of the design effects — against 0.1364, against 0.8332 — which is the sharpest way to say what the row-counting estimate is doing. It is not mildly optimistic and it is not noisy. It is estimating the variance a sample of three hundred independent rows would have had, on a sample that contains five independent things, and it reports one seventh of the truth with the composure of an estimator that has converged. Adding rows inside those five clusters moves it towards a smaller fraction rather than a larger one, because the quantity it is converging on gets further from the truth as the clusters grow.
That is the same failure as the flat model-based line in the coverage sweep under a leaning error variance — an estimator converging on the wrong quantity — with the departure moved from the shape of the error variance to its correlation structure. And it is why a cluster-robust error is not an optional refinement on a study of this shape: the default is not approximately right, it is out by the design effect, and the design effect is a product of a small correlation and a large cluster size.
The corollary for design is worth stating because it is the decision that actually matters. Given three hundred rows to collect, a study choosing between five clusters of sixty and a hundred of three is choosing between an effective sample of about forty-three and one of about two hundred and fifty. That choice is made before any data exists, it is usually made on cost, and nothing in the analysis recovers what it gives away — the essay on a slope pooled across groups reaches the same conclusion from the opposite direction, where how much a group borrows is decided entirely by where its observations sit rather than by how many it has.
What could have produced these numbers without the claim being true
Simulation noise. Six thousand draws at each cell give binomial standard errors between 0.297 and 0.630 points, so the twenty-two-point gap between the two arrangements is about thirty-five standard errors and the forty-point gap between naive coverage at five clusters and at a hundred is larger still. The design effects have a second route: the counted values 9.1005 and 5.3103 against closed forms of 9.3158 and 5.4652, agreeing to within about two per cent, which is what six thousand draws buys for a variance rather than for a proportion.
The closed form could be the sampler restated. It is not: the design effect is an expectation against the stated error structure, computed from the sizes and positions alone, while the counted effect is the empirical variance of six thousand fitted slopes divided by the independent-rows variance. And the naive coverage has a third route, , which reads only the design effect — three quantities, three derivations, one number.
The covariate being constant within a cluster could be doing the work. It is the case the design effect is exact for, and it is chosen for that reason: it is also the case a cluster-robust error is usually reached for, since the treatment, the policy or the site characteristic is a cluster-level variable. A covariate that varies within clusters splits the slope’s information into a within part and a between part, the within part is unaffected by the cluster correlation, and the design effect falls somewhere between one and the value computed here. So the numbers above are the ceiling of the effect rather than a typical value, and saying so is more useful than quoting them as though every clustered regression carried them.
The correlation could be doing the work rather than the cluster size. It cannot be doing much of it: is 0.1 at every cell in this essay, chosen because it is a value that gets described as small and ignored. The design effect is linear in and the largest cluster size is sixty, so is a product of a small number and a large one. A correlation nobody would bother reporting, multiplied by a cluster size everybody reports, is a factor of seven in a variance.
What was not swept
was held at 0.1 throughout, and only the cluster count and the arrangement move. The design effect’s closed form is linear in , so the level of every number here scales predictably and there is nothing to learn from sweeping it. The coverage at few clusters is a different matter: it depends on how far the cluster-robust variance estimate sits from its own target, which is an average over terms and has nothing to do with being small. Whether the 0.5699 at five clusters is a property of the estimator or of this correlation is therefore not settled by anything here, and it is one sweep away.
Nor is the degrees-of-freedom question settled. The on used above is a convention with the same standing as the on in the small-sample coverage sweep: a heuristic that helps, tested rather than derived. It buys ten points at five clusters and leaves the interval nine points short, and the alternatives — a degrees-of-freedom count read off the design, or a reference distribution generated by resampling whole clusters — are the obvious next comparison and are not run.
What the sweep does settle is the thing worth carrying. A standard error is a statement about how many independent things a study contains, and neither the row count nor a formula reading only the cluster sizes can supply that number. The row count is wrong by a factor of seven at five clusters of sixty; the size-only formula is wrong by a factor of 1.13 in one direction and 1.51 in the other on the same set of sizes, which is worse than being wrong by a constant because it cannot be adjusted for. The quantity that is right reads the sizes and where they sit, and it is the same weighting that decides everything else in this field.
What links here
Computed from the collection, not written here: the essays that point at this one.
Reads more easily once this is understood
Essays that name this one as worth reading first.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A block size that changes — both name confidence interval, coverage, degrees of freedom, independence
- A schedule that reads the mean — both name confidence interval, coverage, degrees of freedom, independence
- A width promised for a difference — both name confidence interval, coverage, degrees of freedom, effective sample size
- How many observations a weight leaves — both name design effect, effective sample size, intraclass correlation, kish effective size
- The count or the length — both name confidence interval, coverage, degrees of freedom, dependence
- The degrees of freedom in the sums — both name confidence interval, coverage, degrees of freedom, independence
Named objects
A flat tag is an object no other essay names yet.
Cluster-robust standard errorCluster sizeClustered samplingConfidence intervalCoverageDegrees of freedomDependenceDesign effectEffective sample sizeIndependenceIntraclass correlationKish effective sizeLeverageSampling variationSandwich estimator