Concept

Study design — where it appears

The whole arrangement an experiment is run under — the units, the arms, the blocks, the stopping rule. Most of what can be concluded is fixed by it before any data exists, and none of it can be repaired by analysis afterwards.

Named by 21 essays across 10 fields — each of them below, with the objects they name alongside it.

What each rule leaves behind, at 120 patients. Four allocation rules over the same cohorts and the same seeds, each scored on three imbalances: the number of patients in each arm, the worst of the nine factor levels, and the worst of the 24 cells of the cross-classification. No rule holds all three. Permuted blocks hold the totals exactly and leave the margins near a coin's. Blocks inside every cell hold the cells and let the totals drift, because 24 part-filled blocks do not have to end level. Minimisation holds the margins and the totals and is at 83% of a coin's cell imbalance. Each of the three columns is somebody's definition of a balanced trial.

Balancing what is known in advance

Four allocation rules, three definitions of balance, and no rule that holds more than one of them. Minimisation keeps the worst factor margin near three patients whether the trial has forty or six hundred and forty — and lets the imbalance in the cross-classified cells climb to 86% of a coin's, because the cells are not what it is watching.

covadapt · Assignment
Six groups of 10, each fitting its own slope, then borrowing. Each faint line is one group's own least-squares slope through its own centre; each solid line is that slope after pooling towards the population slope of 0.79. Every group has the same 10 observations. The group whose x values span 0.4 has a slope standard error of 1.86 and moves 91% of the way in; the group spanning 2.0 has a standard error of 0.37 and moves 28%.

The slope that borrows

Pooling a mean makes it look as though how much a group borrows depends on how much data it has. Pool a slope instead and the illusion breaks — ten groups with ten observations each can borrow anything from 28% to 91%, decided entirely by where those ten observations were placed.

multilevel · Levels
Where one rule becomes three. Every arrival in 200 simulated trials is put to all three scores, and the picture is how often they would send that patient to different arms. The range and the pairwise sum are the same rule at two arms and at three — for sorted counts the pairwise sum is twice the range, so the arm that minimises one minimises the other — and they part company at four, where the pairwise sum is 3(d − a) + (c − b) and the range still sees only d − a. The variance disagrees with both from two arms onwards, on 5.4% of arrivals at two and 27.0% at five, because the scores are summed over 3 factors and a sum of squares does not order the candidates the way a sum of absolute values does. All three are called minimisation.

Three arms and three scores

Minimisation balances a trial by keeping the arms' counts even inside every prognostic factor. With two arms there is one way to measure how uneven two counts are. With three there are several, they are all called minimisation, and they send different patients to different arms.

multiarm · Assignment
A trial designed 2:1:1, and what two scores deliver. 500 trials of 180 patients, three arms, a target of 2:1:1. The shaded bars are a minimisation score that divides each arm's count by the share that arm is supposed to receive before measuring the spread; it delivers 49.9% : 25.1% : 25.1%. The others are the same rule with the counts left raw, which delivers 33.4% : 33.3% : 33.3% — the balance it enforces inside every factor level is equality, and equality is what it gets. The marks are the shares that were asked for.

Balancing towards unequal targets

A three-arm trial allocating two to one to one is the ordinary case, and a balancing rule built from raw counts does not know it. It balances the arms towards equality inside every factor level, delivers a third to each arm, and reports that it minimised imbalance.

multiarm · Allocation
What a rule gives away by being predictable. Minimisation run at every probability from a coin to fully deterministic, over 500 cohorts of 120 at each. The upper line is the share of assignments an investigator who knows the rule and the enrolled patients can name in advance: 49.7% at p = 0.5, which is a coin and cannot be beaten, and 87.6% at p = 1 — short of everything only where the two arms tie and the rule falls back on a coin. The lower line is what that is worth: an investigator who enrols a patient 0.5 of a standard deviation better than average whenever they predict their favoured arm produces a treatment effect of 0.75 where the truth is zero. Nothing about the randomisation was broken; the bias entered through who was enrolled, which is the one thing an allocation rule cannot control. The dashed line is the closed form 2δ(2g − 1).

The rule that can be guessed

A balancing rule improves as it becomes more deterministic, and a deterministic rule can be worked out in advance from information the person enrolling the patient already has. At full determinism 87.6% of assignments are guessable, and an investigator who acts on the guess produces a treatment effect of three quarters of a standard deviation where the truth is zero.

covadapt · Assignment
A cut at a quantile, and a cut at a value. Two rules that read identically in a protocol. One splits each covariate at its median; the other splits it at 1 on the covariate's own scale — a dose, a temperature, a clinical threshold. At a correlation of 0.5 the first removes exactly nothing of the interaction between its own two splits, under every marginal here, because a median split is a function of the sign of the latent normal whatever the marginal is. The second removes what the bars show, and it does so on a normal covariate too: the threshold sits at 1.000 on the latent scale rather than at zero, so it is 59.4% odd and 40.6% even. The exact zero was never about the cut; it was about the cut being at the median.

The cut that is not a quantile

A protocol that says split the covariate at a threshold and one that says split it at the median read the same and are different rules. One has an exact guarantee under every marginal and the other has none under any.

skew · Criterion
Power to find a real effect of 3 standard errors, 10 of 20 real. no correction finds 85.1%, Bonferroni finds 49.1%, Holm finds 52.5%, Benjamini–Hochberg finds 74.9%. The uncorrected procedure finds the most and controls nothing.

The price of control

Every correction is paid for in power, and the exchange rate can be measured. Holm buys familywise control for 33 percentage points of power; Benjamini–Hochberg buys a weaker guarantee for 10. Neither is free and neither is a matter of taste.

multiplicity · Multiplicity
A guarantee that stops being a number. The worst case of each dictionary over six outcome shapes, at a correlation of 0.5, against the skewness of the covariate. Under a symmetric marginal every rule made of odd functions has a worst case of exactly zero, and the rule holding a mean and a median split of each covariate — the two things every trial balances — is one of them. Under skew that zero becomes 0.74%, 1.83%, 2.24%, 2.49%: small numbers, each of which depends on a marginal nobody stated. The guarantee has not improved by becoming positive. It has stopped being a guarantee, because it can no longer be written down without the covariate's distribution in it.

Balancing a skewed covariate

The worst case of the rule every trial runs goes from exactly zero to somewhere between a quarter of a per cent and two and a half. Which is small, and is a number that cannot be stated without the covariate's distribution in it.

skew · Criterion
Where the set stops being one set. How many of the 3,432 equal splits of fourteen units a balancing rule admits, as the tolerance tightens, with the number of components single swaps leave it in. The set falls from 886 to 84 assignments, and somewhere in that fall it stops being connected: at 0.8 it is in 2 pieces and every assignment's complement is in the other one. Nothing about the rule changes at that point and nothing a chain reports changes either, which is the whole difficulty — the acceptance rate, the stationary distribution and the detailed balance are all in order on both sides of it.

Before the trial and after

The same diagnostic run at two moments answers two different questions. Before, a positive verdict changes the design. After, it changes which number gets reported — and only for the numbers the defect can reach.

after · Assignment
What a guesser gets, and what a guesser gets for nothing. 600 trials of 150 patients under a fully deterministic rule, with an investigator who knows the rule, the factors and every assignment so far. The guess rate falls with the number of arms — 87.8% at 2, 86.2% at 3, 81.0% at 4 — which reads like a trial getting safer and is not: what a guesser can trade on is the excess over the 50%, 33%, 25% they would get by naming an arm at random, and that goes the other way, from 1.76× chance at two arms to 3.24× at 4. The gap a guesser manufactures between the best and worst arm under a true null is 0.759, 0.733, 0.711 standard deviations — nearly unchanged.

Guessing one arm in three

A balancing rule is guessable because it is balancing. With three arms the next assignment is worked out less often than with two — and by more, relative to what a guesser gets for nothing, and the damage they can do is almost unchanged.

multiarm · Assignment
What 20 clusters of 20 correlated observations do to a 95% interval. Each study has 400 observations arranged as 20 clusters of 20. The lower points are the counted coverage of the usual interval, which treats them as 400 independent observations; the curve through them is 2Φ(1.96/√deff) − 1 with deff = 1 + 19ρ, computed before any data was drawn. At ρ = 0.81 the interval covers 36% rather than 95%. The upper points treat the cluster as the unit and need no variance components at all.

Two levels at once

A third level of grouping adds no new arithmetic and produces one number — the design effect — that decides how many independent observations a clustered study is worth. It is the same quantity the time-series field computes for autocorrelated data, arrived at from a completely different picture.

multilevel · Levels
A covariate that is prior to everything and still ruins it. A covariate measured before the treatment, caused by neither the treatment nor the outcome, and not a common cause of them. Two unmeasured variables sit behind it: one reaches the treatment, the other reaches the outcome, and both reach the covariate. Every rule of thumb for including a baseline variable is satisfied, and the regression that leaves the covariate out estimates the treatment's effect of 0.50 without bias, while the regression that includes it is off by −0.2000 — because the covariate is a common effect of the two unmeasured causes, and conditioning on a common effect makes its causes dependent. The path it opens runs from the treatment back through the first unmeasured cause, through the covariate, and out through the second to the outcome.

A collider before the treatment

A covariate measured before the treatment, on no causal path, and not a common cause of anything, still biases the estimate by exactly −0.2000 against an effect of 0.5 — while the regression that leaves it out is exact. The bias saturates at 0.3536, and the two paths that make it a collider do not appear in that bound.

collider · Conditioning
What a 8-run fraction of 4 factors confounds. The defining relation is I = ABCD, so the resolution is 4. A is estimated as A + BCD; B is estimated as B + ACD; C is estimated as C + ABD; D is estimated as D + ABC. Each of those is an identity about the design rather than an approximation about the data.

The word a fraction costs

A half fraction estimates each main effect as an exact sum of that effect and everything it is confounded with — no error term, no sample-size argument. With every interaction at 0.8 the design reports a true effect of −1 as −0.20, and the design cannot test the assumption that makes the number mean anything.

design · Factorial
What 240 observations are worth, by which question is asked. 8 rows and 10 columns with 3 observations in each cell — 240 in all, each belonging to one row and one column, neither nested in the other. the overall mean: variance 0.1703 against a naive 0.0060, a design effect of 28.4 and 8.5 effective observations; a difference between two rows: variance 2.0682 against a naive 0.0960, a design effect of 21.5 and 11.1 effective observations; a difference between two columns: variance 1.0845 against a naive 0.1200, a design effect of 9.0 and 26.6 effective observations.

Two groupings that cross

Pupils belong to a school and to a neighbourhood, and neither is nested in the other. There is then no design effect: the overall mean is worth 8.5 independent observations out of 240, a row difference 11.1 and a column difference 26.6, and which grouping matters depends on the question rather than on the study.

multilevel · Levels
Twenty hypotheses tested in a declared order, the ten real effects listed first. Effects of three standard errors, ten real, familywise 5%. fixed sequence: 85.3% at position 1, 45.0% at 5, 20.4% at 10; overall power 46.10%; fallback: 49.1% at position 1, 56.4% at 5, 57.9% at 10; overall power 55.87%; Holm: 52.5% at position 1, 52.2% at 5, 52.5% at 10; overall power 52.53%.

An order that spends the error rate

Test twenty hypotheses in a declared order, each at the full 5% and each only if every one before it was rejected, and the first is found 85.3% of the time where Holm finds it 52.5%. The tenth is found 20.4% of the time, the product of the powers before it. Move one true null to the head of the list and every real effect behind it is found no more than 4.3% of the time.

multiplicity · Multiplicity
What each group gains from being pooled, τ = 1. Eight groups whose sizes span a factor of 13.3. The smallest gains 2.076 of squared error, which is 36.8% of the total reduction; the largest gains 0.030, which is 0.5%. 2 of the eight account for half of everything pooling buys.

Where the borrowing goes

Pooling cuts the total squared error across eight groups by 56%. Two of the eight take 61% of that reduction, the four best-measured groups share 11% between them, and the largest group gets 1.5% of what the smallest does. The headline is a fact about the groups nobody was asking about.

hierarchical · Pooling
The Box–Behnken design in three factors: 15 runs, none at a corner. Twelve runs at the midpoints of the cube's edges and 3 at its centre. Every run holds one factor at zero, so no run puts all three factors at an extreme — which is what makes it runnable where a corner is not. The three panels are the design's coordinate projections, with repeated positions marked.

The design that refuses the corners

Box–Behnken runs three factors in fifteen runs and puts none of them at a corner, which is what makes it usable where a corner cannot be run. It predicts the corner 1.84 times worse than the seventeen-run design that goes there, and 1.31 times worse at the middle of a face, and all three numbers are matrix computations with no simulation in them.

design · Factorial
The certificate when 4 runs are already spent. a 2² factorial has already been run and 2 further runs are to be placed. The stationarity condition is no longer max d = p; it is max d = (p − λ·tr(M⁻¹M_fixed))/(1 − λ) with λ = 0.6667 the share of runs already spent, which is 7.0985 here. The search reaches 7.098508790 against it, and the largest value anywhere on a 41×41 grid is 7.098508790.

Augmenting a design that has already run

The equivalence theorem still certifies when some runs are already spent, and one number in it changes: the bound is no longer p but (p − λ·tr(M⁻¹M_fixed))/(1 − λ). It equals p again exactly when the runs already made can still be absorbed into the design that would have been chosen — so the certificate says whether the experiment is still recoverable.

optimality · Equivalence
What a variance estimated from K units is worth. The between-unit mean square is a scaled chi-square on K − 1 degrees of freedom, so the estimator's whole distribution is decided by the number of units. At two units its interquartile range spans a factor of 13.03 and its ten-to-ninety range a factor of 171.3, and it comes out exactly zero on 26.7% of studies. The closed form and 3,000 simulated studies agree to 0.051 at every quantile.

A level with two units

A variance estimated from two units is a scaled chi-square on one degree of freedom. Its interquartile range spans a factor of thirteen, its ten-to-ninety range a factor of a hundred and seventy-one, and it comes out exactly zero on 26.7% of studies — so the design effect it decides runs from 1.00 to 7.01 against a truth of 4.69.

multilevel · Levels
What one lost run costs a 16-run factorial fitting 11 coefficients. Every run is worth the same: dropping any one multiplies every coefficient's variance by 1.2000, which is 1 + 1/(N − p) with N = 16 and p = 11, and gives every pair of coefficients a correlation of 0.1667 where the complete design had exactly zero.

The run that did not happen

Lose one run from any orthogonal design and every coefficient's variance is multiplied by exactly 1 + 1/(N − p), and every pair of coefficients acquires a correlation of exactly 1/(N − p + 1) where there was none. The price is set by the design's spare capacity and by nothing else, and a saturated design cannot survive it at all.

design · Factorial
Each interval covers one question and not the other. Coverage of each interval for the overall mean, scored against both estimands, over 20,000 two-site studies of 10 observations apiece. The fixed-effect interval covers the mean of the two sites in hand 96.37% of the time and the population mean 54.77%. The random-effects interval covers the population mean 94.96% — exactly its level, from one degree of freedom — and over-covers the two sites in hand at 98.25%. Both are correct; they are answers to different questions printed in the same place.

What a two-unit study should report

The fixed-effect interval covers the mean of the two sites in hand 96.37% of the time and the population mean 54.77%. The random-effects interval covers the population mean 94.96% — exactly its level, from one degree of freedom — and is 11.6 times wider.

multilevel · Levels

Named alongside it

The objects these essays reach for when they reach for this one.

Covariate balanceExperimental designHierarchical modelCovariate-adaptive randomisationInteractionMinimisationMonte CarloPermuted blocksSample sizeVariance componentsEffective sample sizeRandom-effects

All concepts