A standard error for a model that is wrong

An effective number of clusters

Ten clusters of thirty rows, a covariate spread evenly across them, and a cluster-robust interval for the slope read against a t on G − 2 = 8 covers 91.1%. The slope needed about five degrees of freedom, not eight, because the two outermost clusters carry 49% of its variance and the middle two 0.6%. A count computed from the sizes and positions before any outcome is seen — 4.45 here, 3.32 when the big clusters sit at the edges — with the bias-reduced variance covers between 94.5% and 96.8% in all twelve layouts measured, where the usual interval on G − 2 covers between 76.2% and 94.5%.

Worth reading first: The bread and the filling · The count that is not the rows.

The reference the sandwich is read against separated two causes of a cluster-robust interval’s failure at few clusters. The estimator is too small on average, and the normal reference it is read against is too thin-tailed for a variance built from a handful of terms. Reading the same estimate against a tt on G−2G-2 degrees of freedom recovered much of the miss: at five equal clusters, from 75.05% to 87.95%. That essay held the cluster sizes equal and varied their number. The count that is not the rows had done the opposite, varying the sizes and their arrangement at a comfortable cluster count. Both closed on the case studies actually present, clusters that are few and unequal at once, and on a specific worry. If a few clusters dominate the sum the variance is built from, the number of terms that matter is smaller than GG, and a tt on G−2G-2 is too optimistic exactly where it is most needed.

The worry turns out to be right, and to understate the problem. The number of terms that matter is smaller than G−2G-2 even when every cluster has the same size. What decides it is not the sizes alone but where each cluster sits along the covariate. It can be computed from the design before any outcome is seen, and a tt on that count closes the gap in every layout measured here.

The setting

Three hundred rows are split into six, ten or twenty clusters. The covariate is constant within a cluster, and clusters sit evenly from −1 to 1, which is the arrangement every clustered regression on this field has used. The within-cluster correlation is 0.2 and the slope is 0.6. Sizes are equal, or run in a straight line from smallest to largest with a ratio of 5 or 15, laid out two ways: the largest clusters at the ends of the covariate range or in its middle. Each setting is drawn five thousand times, and every interval below is computed on the same fits.

Two estimators of the slope’s variance are compared. The usual cluster-robust estimator is the squared sum of each cluster’s score with the small correction every package ships. The bias-reduced estimator inflates each cluster’s score by its own leverage, as the HC2 correction does for single rows. Because the covariate is constant within a cluster, a cluster’s block of the hat matrix is its leverage times a matrix of ones, so the inflation is one number per cluster: the score divided by 1−mghg\sqrt{1 - m_g h_g}, with mghgm_g h_g the cluster’s total leverage.

Where the miss comes from

Coverage of the cluster-robust slope interval at six, ten and twenty clusters of unequal size, by reference. 5,000 studies of 300 rows at each setting, within-cluster correlation 0.2. 6 clusters, equal sizes: 89.1%, 92.6%, 96.0%; 6 clusters, 5:1, big at the edges: 80.1%, 88.1%, 95.2%; 6 clusters, 5:1, big at the centre: 93.8%, 95.4%, 96.8%; 6 clusters, 15:1, big at the edges: 76.2%, 86.8%, 94.5%; 10 clusters, equal sizes: 91.1%, 93.0%, 95.8%; 10 clusters, 5:1, big at the edges: 86.3%, 90.3%, 94.7%; 10 clusters, 5:1, big at the centre: 94.5%, 95.7%, 96.4%; 10 clusters, 15:1, big at the edges: 84.9%, 89.6%, 94.8%; 20 clusters, equal sizes: 92.9%, 93.8%, 95.1%; 20 clusters, 5:1, big at the edges: 90.9%, 92.7%, 94.8%; 20 clusters, 5:1, big at the centre: 94.5%, 95.0%, 95.6%; 20 clusters, 15:1, big at the edges: 90.6%, 92.4%, 94.7% — for the usual variance on G − 2, the bias-reduced variance on G − 2, and the bias-reduced variance on its effective count.
Fig. 1 Coverage of the slope interval in twelve settings — six, ten and twenty clusters, equal sizes and three unequal layouts — for the usual variance on G−2G - 2, the bias-reduced variance on G−2G - 2, and the bias-reduced variance on an effective count of clusters.

The usual interval on G−2G-2 misses in every setting, and by more as the clusters get fewer and the big ones move to the edges. At ten equal clusters it covers 91.1%. At ten clusters with sizes from 4 to 56 and the largest at the ends, it covers 84.9%. At six such clusters, 76.2%. The bias-reduced variance helps in every setting — 93.0%, 89.6% and 86.8% at those three — but on G−2G-2 it still misses wherever the edges are heavy.

The third column is the same bias-reduced variance read against an effective count, and it is where the figure lines up. Across all twelve settings it covers between 94.5% and 96.8%. The worst is six clusters with sizes in a 15-to-1 ratio, at 94.5%, and the best-covered settings are the ones with the big clusters in the middle, which over-cover by a point or two. Nothing about the estimate changed between the second column and the third. Only the number it is compared with did.

The edges carry the slope

The effective count used there is far below G−2G-2, and the reason is visible before any data are drawn.

Ten clusters laid out four ways: each cluster's share of the slope's variance, with its size beneath. Clusters placed evenly along the covariate from −1 to 1, each bar the cluster's share of the slope estimate's true variance, u²m(1 + (m − 1)ρ) normalised, with the cluster's size printed beneath. With equal sizes the two outermost clusters carry 49.1% between them and the two in the middle 0.6%. Effective counts and the counts the studies needed: equal sizes 4.45 and 5.15; 5:1, big at the edges 3.52 and 3.33; 5:1, big at the centre 6.07 and 10.43; 15:1, big at the edges 3.32 and 3.26.
Fig. 2 Ten clusters laid out four ways, with each cluster’s share of the slope’s true variance drawn as a bar and its size printed beneath, beside the effective count for the layout and the count the studies needed.

A slope is estimated from how the outcome changes across the covariate, so a cluster contributes in proportion to its squared distance from the covariate’s mean, as well as to its size and its own internal correlation. With equal sizes and evenly spaced positions, the two outermost of ten clusters carry 49.1% of the slope’s variance between them, and the two in the middle 0.6%. The cluster-robust variance is a sum of ten squared scores, but nearly half of it is two of them. A sum dominated by two terms behaves like a sum of about two, with a little help from the next pair. The tt reference should reflect that, and G−2G-2 does not.

There is a closed form for the evenly spread case. Weight each cluster by its squared position and count clusters the way Kish counts unequal weights, (∑w)2/∑w2(\sum w)^2/\sum w^2. For positions spread evenly over an interval, that count tends to 5G/95G/9. Ten equal clusters carry about 5.6 effective clusters for a slope, not ten, and twenty carry about eleven. That is the arithmetic of a design effect applied to the number of independent terms rather than the number of independent rows. The variance of a variance is set by how many terms it really has, which is the same thing that made a variance from two units uninformative.

Unequal sizes then move the count in whichever direction the layout pushes it. Putting the largest clusters at the edges concentrates the slope’s variance further: with sizes from 4 to 56, the outer two of ten carry two-thirds of it, and the effective count falls from 4.45 to 3.32. Putting the largest clusters in the middle spreads it out, because the big clusters there have little leverage on a slope and the edge clusters that do are small. The count rises to 6.07 — more than with equal sizes. Unequal sizes are not bad in themselves. What matters is whether the size sits where the leverage is.

Three ways to count

Three effective counts were measured, each computable from the design.

The Kish count weights each cluster by its true score variance, ug2mg(1+(mg−1)ρ)u_g^2 m_g (1 + (m_g - 1)\rho), and takes G/(1+CV2)G/(1 + CV^2) of those weights, less one, as its degrees of freedom. It needs the within-cluster correlation ρ\rho. The Bell–McCaffrey count is the Satterthwaite degrees of freedom of the bias-reduced variance: treat the estimate as a quadratic form in the outcomes, compute its mean and variance under a working model, and match a scaled chi-square. Under a working model of independent rows it needs nothing but the design, and that version is the one in the coverage figure. The same computation under the true correlation is an oracle. Its feasible version plugs in a correlation estimated from the residuals of each study.

The working model is wrong — rows within a cluster are correlated — and it is worth seeing why it works anyway. The correlation multiplies each cluster’s score variance by 1+(mg−1)ρ1 + (m_g - 1)\rho. With equal sizes that factor is the same for every cluster, it cancels from the ratio of traces, and the working and true counts agree exactly: 4.45 and 4.45 at ten clusters. With unequal sizes the factor grows with the cluster, so it tilts the weights towards the big clusters. Where the big clusters sit at the edges, that tilt concentrates the slope’s variance further, and the true count, 2.90, is below the working one, 3.32. Where they sit in the middle, it spreads the variance and the true count is higher, 6.61 against 6.07. The working model’s error is therefore in a known direction for each layout, and it is the reason the working count under-covers by a fraction of a point with heavy edges.

The degrees of freedom the studies needed, against what G − 2 and the effective count supply. For each of the 12 settings, the degrees of freedom whose t quantile is the 95th percentile of the bias-reduced interval's standardised error, against G − 2 and against the Bell–McCaffrey count: 6 clusters, equal sizes: needed 2.72, G − 2 = 4, effective 2.41; 6 clusters, 5:1, big at the edges: needed 2.07, G − 2 = 4, effective 2.03; 6 clusters, 5:1, big at the centre: needed 4.39, G − 2 = 4, effective 3.08; 6 clusters, 15:1, big at the edges: needed 1.91, G − 2 = 4, effective 1.96; 10 clusters, equal sizes: needed 5.15, G − 2 = 8, effective 4.45; 10 clusters, 5:1, big at the edges: needed 3.33, G − 2 = 8, effective 3.52; 10 clusters, 5:1, big at the centre: needed 10.43, G − 2 = 8, effective 6.07; 10 clusters, 15:1, big at the edges: needed 3.26, G − 2 = 8, effective 3.32; 20 clusters, equal sizes: needed 10.35, G − 2 = 18, effective 9.89; 20 clusters, 5:1, big at the edges: needed 7.33, G − 2 = 18, effective 7.97; 20 clusters, 5:1, big at the centre: needed 17.76, G − 2 = 18, effective 13.33; 20 clusters, 15:1, big at the edges: needed 6.77, G − 2 = 18, effective 7.56. Points on the diagonal supply exactly what was needed; points above it supply too many, and the interval undercovers.
Fig. 3 For each of the twelve settings, the degrees of freedom the studies needed — the tt whose 97.5% point is the 95th percentile of the bias-reduced interval’s standardised error — against G−2G - 2 and against the Bell–McCaffrey count.

The target can be measured directly. In each setting, the 95th percentile of the standardised error across five thousand studies is the critical value an exact interval would use, and the tt with that 97.5% point is the number of degrees of freedom the studies needed. At ten equal clusters they needed 5.15. The Bell–McCaffrey count supplied 4.45 and G−2G-2 supplied 8. At ten clusters with the big ones at the edges in a 15-to-1 ratio they needed 3.26, and the count supplied 3.32, against G−2G-2’s 8 again. In nine of the twelve settings the count lies within a degree of freedom of what was needed. In the other three it supplies fewer, so the interval errs on the safe side. Every one of those three has the big clusters at the centre, where the studies needed more degrees of freedom than G−2G-2 — 10.43 at ten clusters.

G−2G-2 is above the diagonal in ten of the twelve settings, and the two below it are centre layouts. It supplies more degrees of freedom than the studies had, which is precisely an interval read as narrower than its variance warrants. The rule was derived for clusters that contribute equally, and on a slope they never do.

The usual estimator, read without the bias reduction, needs fewer degrees of freedom again — 3.93 at ten equal clusters — because it is too small on average, and a heavier-tailed reference has to make up for that as well as for the few terms. That is the decomposition the essay on the reference made at equal sizes, now visible in the degrees of freedom. Part of the gap is the estimator, and the bias reduction closes it. The rest is the count, and the effective count closes that.

What each repair adds

Ten clusters of very unequal size: what each repair adds to the slope interval's coverage. Ten clusters of sizes 50, 39, 27, 15, 4, 10, 21, 33, 45, 56 along the covariate, the big ones at the edges, 5,000 studies. Coverage: usual variance, t on G − 2 84.88%; bias-reduced variance, t on G − 2 89.56%; bias-reduced, t on the Kish count 95.92%; bias-reduced, t on the working count 94.80%; bias-reduced, t on the estimated count 95.56%; bias-reduced, t on the true count 95.92%. The counts are 8 for G − 2, 2.89 from the Kish weighting, 3.32 under a working model of independent rows, a median of 2.94 with the correlation estimated, and 2.90 with it known.
Fig. 4 Ten clusters with sizes from 4 to 56 and the largest at the edges: coverage of the slope interval with each repair added in turn, and with each of the four effective counts as the reference.

In the hardest ten-cluster setting the steps can be read one at a time. The usual variance on G−2G-2 covers 84.88%. Bias reduction on the same reference takes it to 89.56%. Every effective count then takes it to the nominal level or past it: 95.92% on the Kish count, 94.80% on the working Bell–McCaffrey count, 95.56% on the count with the correlation estimated, and 95.92% on the oracle. The counts themselves are 2.89, 3.32, a median of 2.94 and 2.90.

The reference supplies a little over half of the repair: about five of the ten points. And the choice among effective counts matters much less than the choice of an effective count over G−2G-2. The working-model version is the most convenient, needing only the design and no estimate of anything, and it is the one that under-covers slightly, by a fifth of a point here. The versions that use the correlation, known or estimated, are a point more conservative.

What it costs

Reading the interval against three degrees of freedom rather than eight widens it. A tt on 3.32 has a 97.5% point of about 3.0 against 2.31 on 8, so the interval is about 30% wider. That is not a cost the method imposes. It is the honest width of an interval built from a variance that ten clusters with these sizes and positions can only estimate from about three effective terms. As the robust standard error showed for single rows, a promise about coverage that holds only asymptotically is paid for, at small samples, in either coverage or width, and here the effective count makes the payment in width.

The width is also paid only where the design concentrates the slope. With the big clusters at the centre, the effective count is 6.07 and the studies needed 10.43, so the count is conservative and the interval wider than strictly necessary. That is the one place a design-based count gives away width it did not need to, and it is the place a study is least at risk.

A layout chosen for its count

The effective count is fixed by the design, so it belongs in the design. A study placing ten clusters of unequal size along a covariate it controls — dose levels across sites, a treatment intensity assigned to schools — faces a choice that a formula for independent rows answers one way and the count answers another.

For independent rows the advice is to put the weight at the extremes, because leverage is where a slope’s information comes from. Under clustering that advice buys almost nothing. The true standard deviation of the slope is 0.2621 with ten equal clusters and 0.2629 with the largest at the edges, because a big cluster is also a strongly correlated one and its rows count for less than their number. The design effect rises from 6.80 to 9.31 and absorbs the leverage gained. What the edge layout does change is the count. It concentrates two-thirds of the slope’s variance in the outer two clusters — 66.6% at a 15-to-1 ratio — and the count falls to 3.32.

An interval read against the degrees of freedom the studies needed is the honest one, and its width combines both effects. With equal sizes it is 1.34 wide on average. With the largest clusters at the edges it is 1.58, and at a 15-to-1 ratio 1.61. With the largest in the middle, whose slope is less precise at 0.2926, it is 1.30, the narrowest of the four, because the count it carries makes up for the precision it lacks. At ten clusters the layout a textbook on design would reject gives the shortest honest interval.

That reverses at twenty clusters, where the count matters less. Equal sizes give the narrowest honest interval there, at 0.87, against 0.90 with the big clusters at the edges and 0.92 in the middle. The rule that survives both is that the variance of a variance enters the design whenever clusters are few. Ignoring it optimises a standard error that the study cannot estimate well. As the bread and the filling showed for single rows, which way an estimated variance errs depends on the design, and here the design also fixes how well it can be estimated at all.

What a clustered regression with few clusters should report

The effective number of clusters for the coefficient reported, not GG. For a slope it depends on where the clusters sit along the covariate as much as on their sizes. For evenly spread equal clusters it is about five-ninths of GG. It can be computed from the design before the outcome is collected, which also makes it a design criterion: a study choosing where to put its clusters can see what each placement buys.

The bias-reduced variance. Its inflation per cluster is one number when the covariate is constant within clusters, and it removes the part of the miss that is the estimator’s rather than the reference’s.

Which count the interval is read against. G−2G-2, a Kish count and a Satterthwaite count give 8, 2.89 and 3.32 in the same study, and the interval covers 89.6%, 95.9% and 94.8% on them.

What is claimed and what is not

Computed exactly. Every effective count, from the design’s cluster sizes and positions, with the Bell–McCaffrey count a ratio of traces of a G×GG \times G matrix of covariances between the clusters’ adjusted scores — never an n×nn \times n one, because the covariate is constant within each cluster.

Counted. Coverage under all six readings and the degrees of freedom each setting needed, on five thousand studies a setting, the same fits for every reading.

Not claimed. One regressor, constant within clusters, and normal errors with exchangeable within-cluster correlation. A covariate that varies within clusters gives each cluster’s block of the hat matrix more than one nonzero eigenvalue, and the bias reduction then needs the matrix square root it can avoid here. Heavy-tailed cluster effects make the scores’ variance itself more variable than the chi-square approximation assumes, and that is the next thing that would move the count.

Still open: clusters that are few and also skewed

The effective count corrects for how many terms the variance really has. It assumes each term is a normal score squared. When the clusters’ random effects are skewed or heavy-tailed, one cluster with an extreme effect at the edge of the covariate both dominates the slope and makes its own score’s square far more variable than a chi-square on one degree of freedom allows. The Satterthwaite count’s variance term would then be wrong in the direction that matters. Whether a wild cluster bootstrap — resampling the sign of each cluster’s residuals rather than relying on a reference at all — keeps its coverage where the effective count does not, and how many clusters it needs to do so, is a measurement this field has not made.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Cluster-robust standard errorCluster sizeClustered samplingCoverageDegrees of freedomEffective sample sizeHat matrixKish effective sizeLeverageMonte CarloReference distributionT interval