An effective number of clusters
Worth reading first: The bread and the filling · The count that is not the rows.
The reference the sandwich is read against separated two causes of a cluster-robust interval’s failure at few clusters. The estimator is too small on average, and the normal reference it is read against is too thin-tailed for a variance built from a handful of terms. Reading the same estimate against a on degrees of freedom recovered much of the miss: at five equal clusters, from 75.05% to 87.95%. That essay held the cluster sizes equal and varied their number. The count that is not the rows had done the opposite, varying the sizes and their arrangement at a comfortable cluster count. Both closed on the case studies actually present, clusters that are few and unequal at once, and on a specific worry. If a few clusters dominate the sum the variance is built from, the number of terms that matter is smaller than , and a on is too optimistic exactly where it is most needed.
The worry turns out to be right, and to understate the problem. The number of terms that matter is smaller than even when every cluster has the same size. What decides it is not the sizes alone but where each cluster sits along the covariate. It can be computed from the design before any outcome is seen, and a on that count closes the gap in every layout measured here.
The setting
Three hundred rows are split into six, ten or twenty clusters. The covariate is constant within a cluster, and clusters sit evenly from −1 to 1, which is the arrangement every clustered regression on this field has used. The within-cluster correlation is 0.2 and the slope is 0.6. Sizes are equal, or run in a straight line from smallest to largest with a ratio of 5 or 15, laid out two ways: the largest clusters at the ends of the covariate range or in its middle. Each setting is drawn five thousand times, and every interval below is computed on the same fits.
Two estimators of the slope’s variance are compared. The usual cluster-robust estimator is the squared sum of each cluster’s score with the small correction every package ships. The bias-reduced estimator inflates each cluster’s score by its own leverage, as the HC2 correction does for single rows. Because the covariate is constant within a cluster, a cluster’s block of the hat matrix is its leverage times a matrix of ones, so the inflation is one number per cluster: the score divided by , with the cluster’s total leverage.
Where the miss comes from
The usual interval on misses in every setting, and by more as the clusters get fewer and the big ones move to the edges. At ten equal clusters it covers 91.1%. At ten clusters with sizes from 4 to 56 and the largest at the ends, it covers 84.9%. At six such clusters, 76.2%. The bias-reduced variance helps in every setting — 93.0%, 89.6% and 86.8% at those three — but on it still misses wherever the edges are heavy.
The third column is the same bias-reduced variance read against an effective count, and it is where the figure lines up. Across all twelve settings it covers between 94.5% and 96.8%. The worst is six clusters with sizes in a 15-to-1 ratio, at 94.5%, and the best-covered settings are the ones with the big clusters in the middle, which over-cover by a point or two. Nothing about the estimate changed between the second column and the third. Only the number it is compared with did.
The edges carry the slope
The effective count used there is far below , and the reason is visible before any data are drawn.
A slope is estimated from how the outcome changes across the covariate, so a cluster contributes in proportion to its squared distance from the covariate’s mean, as well as to its size and its own internal correlation. With equal sizes and evenly spaced positions, the two outermost of ten clusters carry 49.1% of the slope’s variance between them, and the two in the middle 0.6%. The cluster-robust variance is a sum of ten squared scores, but nearly half of it is two of them. A sum dominated by two terms behaves like a sum of about two, with a little help from the next pair. The reference should reflect that, and does not.
There is a closed form for the evenly spread case. Weight each cluster by its squared position and count clusters the way Kish counts unequal weights, . For positions spread evenly over an interval, that count tends to . Ten equal clusters carry about 5.6 effective clusters for a slope, not ten, and twenty carry about eleven. That is the arithmetic of a design effect applied to the number of independent terms rather than the number of independent rows. The variance of a variance is set by how many terms it really has, which is the same thing that made a variance from two units uninformative.
Unequal sizes then move the count in whichever direction the layout pushes it. Putting the largest clusters at the edges concentrates the slope’s variance further: with sizes from 4 to 56, the outer two of ten carry two-thirds of it, and the effective count falls from 4.45 to 3.32. Putting the largest clusters in the middle spreads it out, because the big clusters there have little leverage on a slope and the edge clusters that do are small. The count rises to 6.07 — more than with equal sizes. Unequal sizes are not bad in themselves. What matters is whether the size sits where the leverage is.
Three ways to count
Three effective counts were measured, each computable from the design.
The Kish count weights each cluster by its true score variance, , and takes of those weights, less one, as its degrees of freedom. It needs the within-cluster correlation . The Bell–McCaffrey count is the Satterthwaite degrees of freedom of the bias-reduced variance: treat the estimate as a quadratic form in the outcomes, compute its mean and variance under a working model, and match a scaled chi-square. Under a working model of independent rows it needs nothing but the design, and that version is the one in the coverage figure. The same computation under the true correlation is an oracle. Its feasible version plugs in a correlation estimated from the residuals of each study.
The working model is wrong — rows within a cluster are correlated — and it is worth seeing why it works anyway. The correlation multiplies each cluster’s score variance by . With equal sizes that factor is the same for every cluster, it cancels from the ratio of traces, and the working and true counts agree exactly: 4.45 and 4.45 at ten clusters. With unequal sizes the factor grows with the cluster, so it tilts the weights towards the big clusters. Where the big clusters sit at the edges, that tilt concentrates the slope’s variance further, and the true count, 2.90, is below the working one, 3.32. Where they sit in the middle, it spreads the variance and the true count is higher, 6.61 against 6.07. The working model’s error is therefore in a known direction for each layout, and it is the reason the working count under-covers by a fraction of a point with heavy edges.
The target can be measured directly. In each setting, the 95th percentile of the standardised error across five thousand studies is the critical value an exact interval would use, and the with that 97.5% point is the number of degrees of freedom the studies needed. At ten equal clusters they needed 5.15. The Bell–McCaffrey count supplied 4.45 and supplied 8. At ten clusters with the big ones at the edges in a 15-to-1 ratio they needed 3.26, and the count supplied 3.32, against ’s 8 again. In nine of the twelve settings the count lies within a degree of freedom of what was needed. In the other three it supplies fewer, so the interval errs on the safe side. Every one of those three has the big clusters at the centre, where the studies needed more degrees of freedom than — 10.43 at ten clusters.
is above the diagonal in ten of the twelve settings, and the two below it are centre layouts. It supplies more degrees of freedom than the studies had, which is precisely an interval read as narrower than its variance warrants. The rule was derived for clusters that contribute equally, and on a slope they never do.
The usual estimator, read without the bias reduction, needs fewer degrees of freedom again — 3.93 at ten equal clusters — because it is too small on average, and a heavier-tailed reference has to make up for that as well as for the few terms. That is the decomposition the essay on the reference made at equal sizes, now visible in the degrees of freedom. Part of the gap is the estimator, and the bias reduction closes it. The rest is the count, and the effective count closes that.
What each repair adds
In the hardest ten-cluster setting the steps can be read one at a time. The usual variance on covers 84.88%. Bias reduction on the same reference takes it to 89.56%. Every effective count then takes it to the nominal level or past it: 95.92% on the Kish count, 94.80% on the working Bell–McCaffrey count, 95.56% on the count with the correlation estimated, and 95.92% on the oracle. The counts themselves are 2.89, 3.32, a median of 2.94 and 2.90.
The reference supplies a little over half of the repair: about five of the ten points. And the choice among effective counts matters much less than the choice of an effective count over . The working-model version is the most convenient, needing only the design and no estimate of anything, and it is the one that under-covers slightly, by a fifth of a point here. The versions that use the correlation, known or estimated, are a point more conservative.
What it costs
Reading the interval against three degrees of freedom rather than eight widens it. A on 3.32 has a 97.5% point of about 3.0 against 2.31 on 8, so the interval is about 30% wider. That is not a cost the method imposes. It is the honest width of an interval built from a variance that ten clusters with these sizes and positions can only estimate from about three effective terms. As the robust standard error showed for single rows, a promise about coverage that holds only asymptotically is paid for, at small samples, in either coverage or width, and here the effective count makes the payment in width.
The width is also paid only where the design concentrates the slope. With the big clusters at the centre, the effective count is 6.07 and the studies needed 10.43, so the count is conservative and the interval wider than strictly necessary. That is the one place a design-based count gives away width it did not need to, and it is the place a study is least at risk.
A layout chosen for its count
The effective count is fixed by the design, so it belongs in the design. A study placing ten clusters of unequal size along a covariate it controls — dose levels across sites, a treatment intensity assigned to schools — faces a choice that a formula for independent rows answers one way and the count answers another.
For independent rows the advice is to put the weight at the extremes, because leverage is where a slope’s information comes from. Under clustering that advice buys almost nothing. The true standard deviation of the slope is 0.2621 with ten equal clusters and 0.2629 with the largest at the edges, because a big cluster is also a strongly correlated one and its rows count for less than their number. The design effect rises from 6.80 to 9.31 and absorbs the leverage gained. What the edge layout does change is the count. It concentrates two-thirds of the slope’s variance in the outer two clusters — 66.6% at a 15-to-1 ratio — and the count falls to 3.32.
An interval read against the degrees of freedom the studies needed is the honest one, and its width combines both effects. With equal sizes it is 1.34 wide on average. With the largest clusters at the edges it is 1.58, and at a 15-to-1 ratio 1.61. With the largest in the middle, whose slope is less precise at 0.2926, it is 1.30, the narrowest of the four, because the count it carries makes up for the precision it lacks. At ten clusters the layout a textbook on design would reject gives the shortest honest interval.
That reverses at twenty clusters, where the count matters less. Equal sizes give the narrowest honest interval there, at 0.87, against 0.90 with the big clusters at the edges and 0.92 in the middle. The rule that survives both is that the variance of a variance enters the design whenever clusters are few. Ignoring it optimises a standard error that the study cannot estimate well. As the bread and the filling showed for single rows, which way an estimated variance errs depends on the design, and here the design also fixes how well it can be estimated at all.
What a clustered regression with few clusters should report
The effective number of clusters for the coefficient reported, not . For a slope it depends on where the clusters sit along the covariate as much as on their sizes. For evenly spread equal clusters it is about five-ninths of . It can be computed from the design before the outcome is collected, which also makes it a design criterion: a study choosing where to put its clusters can see what each placement buys.
The bias-reduced variance. Its inflation per cluster is one number when the covariate is constant within clusters, and it removes the part of the miss that is the estimator’s rather than the reference’s.
Which count the interval is read against. , a Kish count and a Satterthwaite count give 8, 2.89 and 3.32 in the same study, and the interval covers 89.6%, 95.9% and 94.8% on them.
What is claimed and what is not
Computed exactly. Every effective count, from the design’s cluster sizes and positions, with the Bell–McCaffrey count a ratio of traces of a matrix of covariances between the clusters’ adjusted scores — never an one, because the covariate is constant within each cluster.
Counted. Coverage under all six readings and the degrees of freedom each setting needed, on five thousand studies a setting, the same fits for every reading.
Not claimed. One regressor, constant within clusters, and normal errors with exchangeable within-cluster correlation. A covariate that varies within clusters gives each cluster’s block of the hat matrix more than one nonzero eigenvalue, and the bias reduction then needs the matrix square root it can avoid here. Heavy-tailed cluster effects make the scores’ variance itself more variable than the chi-square approximation assumes, and that is the next thing that would move the count.
Still open: clusters that are few and also skewed
The effective count corrects for how many terms the variance really has. It assumes each term is a normal score squared. When the clusters’ random effects are skewed or heavy-tailed, one cluster with an extreme effect at the edge of the covariate both dominates the slope and makes its own score’s square far more variable than a chi-square on one degree of freedom allows. The Satterthwaite count’s variance term would then be wrong in the direction that matters. Whether a wild cluster bootstrap — resampling the sign of each cluster’s residuals rather than relying on a reference at all — keeps its coverage where the effective count does not, and how many clusters it needs to do so, is a measurement this field has not made.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- What a two-unit study should report — both name coverage, degrees of freedom, effective sample size, monte carlo, t interval
- A block size that changes — both name coverage, degrees of freedom, monte carlo
- A coverage table with its own error — both name coverage, monte carlo, t interval
- A penalty is a trace — both name degrees of freedom, effective sample size, hat matrix
- A point ordinary on every axis — both name effective sample size, hat matrix, leverage
- A prior on the spread between two sites — both name coverage, degrees of freedom, monte carlo
Named objects
A flat tag is an object no other essay names yet.
Cluster-robust standard errorCluster sizeClustered samplingCoverageDegrees of freedomEffective sample sizeHat matrixKish effective sizeLeverageMonte CarloReference distributionT interval