Hierarchy past one number

Two groupings that cross

Pupils belong to a school and to a neighbourhood, and neither is nested in the other. There is then no design effect: the overall mean is worth 8.5 independent observations out of 240, a row difference 11.1 and a column difference 26.6, and which grouping matters depends on the question rather than on the study.

Worth reading first: The slope that borrows.

Two levels at once is about nesting: wards inside hospitals, pupils inside classes inside schools. Every unit belongs to exactly one group at each level, the levels sit inside one another, and the whole apparatus collapses into a single number — the design effect — that says how many independent observations the study is worth.

A great many studies are not nested. Pupils belong to a school and to a neighbourhood, and the two are not the same partition: one school draws from many neighbourhoods, one neighbourhood sends pupils to many schools. Measurements belong to a day and to an instrument. Ratings belong to a rater and to an item.

What 240 observations are worth, by which question is asked8 rows and 10 columns with 3 observations in each cell — 240 in all, each belonging to one row and one column, neither nested in the other. the overall mean: variance 0.1703 against a naive 0.0060, a design effect of 28.4 and 8.5 effective observations; a difference between two rows: variance 2.0682 against a naive 0.0960, a design effect of 21.5 and 11.1 effective observations; a difference between two columns: variance 1.0845 against a naive 0.1200, a design effect of 9.0 and 26.6 effective observations.the overall mean8.5design effect 28.4a difference between two rows11.1design effect 21.5a difference between two columns26.6design effect 9.0240 observations, 2,500 studiesone study, three effective sizes
Fig. 1 Eight rows and ten columns with three observations in each cell — 240 observations, each belonging to one row and one column. The overall mean is worth 8.5 independent observations; a difference between two rows is worth 11.1; a difference between two columns is worth 26.6.
16 groups shrunk towards a fitted line, at γ = 1.2. Hollow circles are the groups' own values, filled ones the estimates after pooling, and the diagonal is the line fitted through them with each group weighted by how well it is measured — slope 1.21, intercept 0.20. The horizontal rule is where the same 16 groups would have been shrunk to with no covariate. The spread left to borrow against is 0.64 with the covariate against 1.28 without, so every group is pulled further in than it would otherwise have been.
Fig. 2 The field’s other answer to a second structure, for contrast: a group-level predictor rather than a second grouping. Where the second thing a unit belongs to is a measured quantity rather than a set of unordered levels, it enters as a covariate with one parameter — and the design effect stays a single number, because the covariate is not a source of correlation between units.

The model gains one line and loses one number

The nested model writes an observation as a cluster effect plus a group effect plus noise. The crossed model writes it as

y=μ+arow+bcol+εy = \mu + a_{\text{row}} + b_{\text{col}} + \varepsilon

which is the same length and the same kind of object. Two variance components rather than one, both estimated the same way, nothing new to fit.

What is lost is the summary. In the nested case, two observations are correlated if they share a cluster and independent otherwise, so counting correlated pairs gives one number — 1+(m1)ρ1 + (m-1)\rho — and it describes the whole study.

Here two observations share a row, or a column, or both, or neither. Three kinds of pair rather than two, with different correlations, and how many of each a contrast involves depends on the contrast.

Three contrasts, three answers

The three questions a crossed study is usually asked have three different answers, and the arithmetic for each is short enough to write out.

The overall mean averages over all rows and all columns, so both grouping effects survive:

Var(yˉ)=σa2R+σb2C+σ2N\operatorname{Var}(\bar y) = \frac{\sigma_a^{2}}{R} + \frac{\sigma_b^{2}}{C} + \frac{\sigma^{2}}{N}

At eight rows, ten columns and three per cell that is 0.125+0.049+0.006=0.1800.125 + 0.049 + 0.006 = 0.180, against the 0.0060.006 an analysis ignoring both would claim — a design effect of 28.4.

A difference between two rows subtracts two averages that each cover every column, so every column effect appears in both and cancels exactly:

Var(yˉ1yˉ2)=2(σa2+σ2Cn)\operatorname{Var}(\bar y_{1\cdot} - \bar y_{2\cdot}) = 2\Bigl(\sigma_a^{2} + \frac{\sigma^{2}}{Cn}\Bigr)

The column variance is not in it at all. Only the row variance and the noise, and the design effect is 21.5.

A difference between two columns is the mirror image: the row effects cancel, the column variance survives, and the design effect is 9.0.

So one study produces design effects of 28.4, 21.5 and 9.0 depending on the question. There is no number to report.

The variance of each contrast, against what ignoring the groupings claims. 8 rows and 10 columns with 3 observations in each cell — 240 in all, each belonging to one row and one column, neither nested in the other. the overall mean: variance 0.1703 against a naive 0.0060, a design effect of 28.4 and 8.5 effective observations; a difference between two rows: variance 2.0682 against a naive 0.0960, a design effect of 21.5 and 11.1 effective observations; a difference between two columns: variance 1.0845 against a naive 0.1200, a design effect of 9.0 and 26.6 effective observations.
Fig. 3 The same three contrasts as variances rather than as counts, with what an analysis ignoring both groupings would claim printed beside each. The overall mean’s variance is twenty-eight times the naive one; the column difference’s is nine times. The three are the same three numbers in the form the fit actually produces.

What separates them

Holding the study fixed and moving one grouping’s spread shows the three quantities are not variations on one thing.

Three questions, three effective sample sizes, one study. 240 observations arranged in 8 rows and 10 columns. As the column grouping's spread grows from 0.1 to 3, the column contrast falls from 210.2 effective observations to 1.6 and the row contrast stays near 11.3, because the column effects cancel out of a row difference exactly.
Fig. 4 The same 240 observations as the column grouping’s spread grows from 0.1 to 3. The column contrast falls from 210.2 effective observations to 1.6. The row contrast stays near 11.3 throughout, because the column effects cancel out of a row difference exactly.

The flat line is the finding. A grouping that grows from being almost irrelevant to dominating the study changes one contrast by a factor of a hundred and thirty and leaves another untouched — not approximately, exactly, because the cancellation is algebraic rather than statistical.

That is why no single number works. A design effect is a statement about a study; here the correlation structure interacts with the contrast, and the two cannot be separated into a property of the design times a property of the question.

The three kinds of pair

The counting argument that gives a nested design its one number is worth doing here, because it says exactly where it breaks.

In a nested design, two observations either share a cluster — correlation ρ\rho — or do not. The variance of a mean is the average variance plus the average covariance over pairs, and with one kind of correlated pair that sum is 1+(m1)ρ1 + (m-1)\rho times what independence would give. One kind of pair, one number.

Here there are three. Two observations in the same cell share both groupings and have covariance σa2+σb2\sigma_a^{2} + \sigma_b^{2}. Two in the same row but different columns have σa2\sigma_a^{2}. Two in the same column but different rows have σb2\sigma_b^{2}. Two sharing neither have zero.

A contrast’s variance is a weighted sum over those four categories with weights set by the contrast’s own coefficients, and the weights are what changes between questions. For the overall mean every pair appears with a positive weight; for a row difference the same-column pairs appear with opposite signs from the two rows and cancel.

So the study’s correlation structure is one thing and the effective sample size is not a property of it. It is a property of the pair — the structure and the question — which is why one of the two can be held fixed and the other still changes the answer.

The cancellation is blocking

The exact cancellation in a row difference has a name in another field, and recognising it says what a crossed design is for.

Every row is observed at every column. So a comparison between rows is made within each column and averaged — which is exactly what blocking does: the block effect is common to the units being compared and subtracts out.

A crossed design is therefore two blocking structures at once, each blocking the other’s contrasts. The column grouping is a nuisance for the overall mean and a block for row comparisons, and which it is depends on what is being asked.

That also says when a crossed structure is good news. A study asking only about row differences is better off with a large column variance than a small one, because a large column variance means each column is a tighter block. The 210.2 effective observations at the left of the sweep and the 1.6 at the right are the same fact from the column contrast’s point of view, and the row contrast’s 11.3 is what blocking looks like when it works perfectly.

What 240 observations are worth, by which question is asked. 8 rows and 10 columns with 3 observations in each cell — 240 in all, each belonging to one row and one column, neither nested in the other. the overall mean: variance 0.9846 against a naive 0.0060, a design effect of 164.1 and 1.5 effective observations; a difference between two rows: variance 2.0682 against a naive 0.0960, a design effect of 21.5 and 11.1 effective observations; a difference between two columns: variance 17.8068 against a naive 0.1200, a design effect of 148.4 and 1.6 effective observations.
Fig. 5 The same study with the column grouping’s spread at three rather than 0.7. The column difference is now worth 1.6 effective observations out of 240 and the row difference is still worth 11.3. One grouping has become overwhelming and the contrast it does not touch has not moved.
What 240 observations are worth, by which question is asked. 8 rows and 10 columns with 3 observations in each cell — 240 in all, each belonging to one row and one column, neither nested in the other. the overall mean: variance 0.1297 against a naive 0.0060, a design effect of 21.6 and 11.1 effective observations; a difference between two rows: variance 2.0682 against a naive 0.0960, a design effect of 21.5 and 11.1 effective observations; a difference between two columns: variance 0.2009 against a naive 0.1200, a design effect of 1.7 and 143.4 effective observations.
Fig. 6 And with it nearly absent. The column difference is worth 210.2 of the 240 — close to what independence would give — while the row difference is, again, 11.3. The row contrast is the one quantity in this study that no amount of column structure changes.

How large the loss is, in plain terms

The numbers deserve restating in the units a reader of a study would meet, because “8.5 effective observations out of 240” sounds like an arithmetic curiosity and is not.

Two hundred and forty pupils, eight schools, ten neighbourhoods. An analysis that treats the 240 as independent quotes a standard error for the overall mean that is 28.4=5.3\sqrt{28.4} = 5.3 times too small, so an interval it draws is five times too narrow and a test it runs rejects far more often than it should. That is the ordinary consequence the nested field measured as coverage, and it is the same size here.

What is new is that the same analysis, asked for a difference between two schools, is wrong by 21.5=4.6\sqrt{21.5} = 4.6 times, and asked for a difference between two neighbourhoods by 9.0=3.0\sqrt{9.0} = 3.0 times. Three corrections, from one dataset, and an analyst who applies the first to all three over-corrects two of them.

The most common practical error is the opposite of that and worse: fitting one of the two groupings — usually the one the study is about — and leaving the other in the residual. The residual variance is then inflated by the omitted grouping, which widens every interval, and the widening is roughly right for one contrast and wrong for the other two.

Two routes

The three variance formulas are algebra about a simulation nobody ran, so they are checked against one.

Twenty-five hundred studies from the stated model, with each contrast computed and its variance taken across studies. The counted variances are 0.1703, 2.0682 and 1.0845 against closed forms of 0.1800, 2.0960 and 1.1000 — within the simulation’s own error at each, and sharing nothing but the four parameters.

The refusal beside it is the number a nested analysis would report. Applying the grand mean’s design effect of 28.4 to a row difference gives a variance of 2.72 against the true 2.10 — wrong by 30% — which is what “one design effect describes the study” would mean and is required to fail.

Where the nested case is a special case

Setting the two structures side by side says what nesting buys and what it is.

A nested design is a crossed design in which one of the two groupings is constant within the other: every pupil in a class is in the same school, so the class grouping refines the school grouping rather than crossing it. Then the “same row, different column” category is empty, there are two kinds of pair rather than three, and the counting argument gives one number.

That is not a small simplification. It is the reason the whole apparatus of design effects, intraclass correlations and effective sample sizes exists for nested studies and has no equivalent here.

It also says which real studies are which, and the distinction is not always the one intuition gives. Pupils in classes in schools is nested. Pupils in schools and neighbourhoods is crossed. Patients in wards in hospitals is nested; patients seen by doctors and in clinics is crossed if a doctor works in more than one clinic and nested if not — and which it is is a fact about the roster rather than about the analysis.

What an analysis should report

The measurements support a short list and none of it is difficult.

Report the variance components, not a design effect. σa2\sigma_a^{2}, σb2\sigma_b^{2} and σ2\sigma^{2} are three numbers that determine every contrast’s variance through the formulas above, and they are what the fit produces anyway. A design effect is a summary that exists because the nested case has one, and the crossed case does not.

Say which grouping the contrast is within. A row difference is made within columns and is therefore immune to the column variance; that sentence tells a reader more than any number, and it is checkable from the design rather than from the fit.

Fit both groupings even when one is a nuisance. A row difference is immune to the column variance only if the columns are balanced across the rows, which is a property of the design and not of the analysis; leaving the columns out of the model and relying on the balance works when the balance is exact and degrades when it is not. The cost of including them is one variance component.

And do not power a crossed study on one effective sample size. A study powered for a row difference at 11.1 effective observations has 26.6 for a column difference and 8.5 for the overall mean. Powering on the wrong one of the three is off by a factor of three, in either direction.

What 20 clusters of 20 correlated observations do to a 95% interval. Each study has 400 observations arranged as 20 clusters of 20. The lower points are the counted coverage of the usual interval, which treats them as 400 independent observations; the curve through them is 2Φ(1.96/√deff) − 1 with deff = 1 + 19ρ, computed before any data was drawn. At ρ = 0.81 the interval covers 36% rather than 95%. The upper points treat the cluster as the unit and need no variance components at all.
Fig. 7 What the nested case looks like, for comparison: one curve, one design effect, one effective sample size, and a closed form for the coverage of an interval that ignores it. The crossed case has no equivalent picture, because the axis would need to be labelled with a question.

What this does to a power calculation

The practical damage is at the design stage rather than at the analysis stage, and it is worth the arithmetic because the error is a factor rather than a margin.

A power calculation asks for a sample size that detects a stated effect. What it needs is the variance of the contrast the study is about, and a crossed study has three of them. Powering on the wrong one is off by the ratio.

At the setting drawn here, a study powered for a row difference needs enough for a variance of 2.096. Powered on the overall mean’s design effect it would size itself for 2.725 — 30% too large, which is a waste rather than a failure. Powered on the column difference’s design effect it would size itself for 0.868, which is 59% too small, and the study would be under-powered by a factor that no interim analysis would reveal as a design error.

The direction is what makes the second worse. An over-sized study produces a correct answer expensively; an under-sized one produces an inconclusive answer and the inconclusiveness is attributed to the effect being small.

The fix is to name the contrast before computing anything, which is good practice everywhere and is merely optional in a nested design where the three answers coincide.

The time-series version of the same statement

The cancellation above has an analogue in time series, and naming it says that the phenomenon is about contrasts rather than about groupings.

The observations that repeat each other computes the effective sample size of an autocorrelated series for a mean, and gets n(1φ)/(1+φ)n(1-\varphi)/(1+\varphi) — a single number, because a series has one correlation structure and the essay asks one question of it. Ask a different question and the number changes: a difference between the first half and the second half of the series has a variance that depends on the correlation completely differently, and a high-frequency contrast is nearly unaffected by a slow-moving dependence that ruins the mean.

So a design effect is never a property of a dataset. It is a property of a dataset and a contrast, and the nested case hides that because its three natural contrasts happen to give the same answer. A crossed study is the same statement in a setting where they do not, and the blocking field’s gain is the same statement again with the sign reversed.

What is claimed here, and what is not

The claim is what happens to a study’s effective sample size when two groupings cross: that three contrasts from the same 240 observations are worth 8.5, 11.1 and 26.6 independent ones; that the closed forms for the three variances match a count to within the simulation’s own error; that a grouping’s spread changes one contrast by a factor of a hundred and thirty and leaves another exactly unchanged; and that applying the overall mean’s design effect to a row contrast is wrong by 30%.

Every number is twenty-five hundred studies from a balanced crossed design with three observations in each cell.

What stays out: unbalanced crossed designs, where the cancellation is no longer exact and the variances have no closed form as short as these; a row-by-column interaction, which is a third variance component and is what the residual absorbs when it is present and unmodelled; and the estimation of the components themselves, which has the same atom at zero the nested case does and for the same reason — a variance estimated as a difference of mean squares can come out negative whatever it is a variance of.

The design is balanced deliberately, because the exact cancellation is the essay’s point and it is a property of balance. In an unbalanced crossed study the column effects cancel out of a row difference only approximately, and how approximately is a measurement nothing here makes.

Still open: a level with two units

Every variance component above is estimated from eight rows or ten columns, which is enough for the estimate to mean something. A great many studies have fewer.

A two-site trial has two units at its top level. The between-site mean square is then a scaled chi-square on one degree of freedom, whose interquartile range spans a factor of 13 and whose ten-to-ninety range spans a factor of 171 — and everything downstream of it, the intraclass correlation, the design effect and the effective sample size, inherits that spread. That is what a level with two units can say.

The check, and the refusal

Two claims are gated. That the counted variance of each of the three contrasts matches its closed form, which is the two-route habit applied to three quantities that share a derivation and nothing else. And that the three effective sample sizes differ by more than a factor of two, which is the essay’s claim stated as a threshold rather than as three values — if they agreed, a single design effect would describe the study and the essay would have no subject.

The refusal is that single number, required to fail: applying the overall mean’s design effect to a row contrast must get the variance wrong by more than a fifth. It is wrong by 30%. A check in which it came out right would mean the crossed structure was behaving like a nested one, which happens when one of the two variances is negligible — so the refusal is also a check that the figure is drawn at a setting where both groupings are doing something.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

BlockingCorrelationCrossed random-effectsDependenceEffective sample sizeHierarchical modelRandom-effectsStudy designVariance components