Concept

Treatment effect — where it appears

The difference an intervention makes to an outcome, estimated by comparing arms that were assigned rather than chosen. It is defined by the assignment mechanism, so a comparison of arms that were chosen rather than assigned is estimating something else.

Named by 15 essays across 6 fields — each of them below, with the objects they name alongside it.

The six best bases of 2 functions, and what each protects. Every cell is R²(g | span B) — the share of the imbalance in that shape a rule balancing that basis removes — computed from exact inner products between Hermite functions and indicators, with nothing simulated. The rows are ordered by their worst cell, which is the number an experimenter who does not know the shape is exposed to. The best row here guarantees 26.8% against every shape in the list, and the worst of the six guarantees 15.1%: the difference between them is entirely which subspace was picked, at the same cost per arrival.

A basis is a subspace

A balancing rule cannot tell one basis from another with the same span, so choosing what to hand it is choosing a subspace — and then what it removes of any outcome shape is a projection, computable exactly, with no trial anywhere in it.

basis · Criterion
Four rules, three shapes, and no ordering that survives. The variance of the unadjusted treatment estimate under each rule, as a fraction of the variance a coin gives, over 450 trials of 200 units each. Against a covariate that enters linearly the rule that reads the number nearly halves it. Against a threshold at 1 it removes about a fifth. Against a quadratic every rule here is at or worse than a coin — they are all optimising a criterion that is one over the variance of an estimate in a model this outcome does not obey, and a constraint that helps nothing still costs something. Nothing in a trial says which column it is in.

Balanced on the wrong function

A rule that reads a covariate's numbers halves the variance of the treatment estimate, if the covariate enters the outcome as a straight line. If it enters as a threshold the rule is worth a fifth of that, and if it enters as a curve every rule here is worse than a coin.

shape · Blocking
Two diagnostics, one answer, two different moments. The two-chain statistic on a covariate probe and on the trial's own difference in arm means, at seven tolerances of a fourteen-unit rule, against the enumerated truth. Both are quiet wherever the set is one set and both fire wherever it is not, at every tolerance — which is what says the outcome probe is the same test rather than a resemblance of it. The difference between them is not accuracy and it is not power. It is when: the covariate probe can be run before a single outcome exists, when a practitioner can still loosen the rule or change the sampler, and it can be run again on a different function if it comes back quiet. The outcome probe runs after the trial, on the one column the trial produced, and what it can do with a positive verdict is repair the p-value rather than the design.

The statistic the p-value is about

The test for whether a balanced-assignment walk reaches its whole set is run on a covariate function chosen before the trial. Run on the difference in arm means it is the same test, and it is about the number the trial publishes.

after · Randomisation
What a mean split leaves, with both halves varying. The share of a mean split's interaction that survives the rule balancing it, at every copula and every marginal, matched at a Spearman correlation of 0.40. The three radially symmetric copulas leave exactly nothing with a symmetric covariate and rise steeply with the skew. The two asymmetric ones start at 7.707% and go opposite ways: the lower-tail copula falls to 0.002% at a skewness of 0.95 — the two failures cancel almost exactly, and a guarantee both fields report as broken is restored — while the upper-tail one climbs to 40.288%. And the heavy-tailed symmetric covariate, which leaks exactly nothing on its own, doubles what the asymmetric copulas leak: 14.229% against 7.707%.

Two failures that cancel

A mildly skewed covariate under a lower-tail copula leaks 0.002% of an interaction where each failure alone leaks eight and seven per cent. Turn the copula over and the same pair compounds.

compound · Adjustment
The probe a trial has is the probe a trial got. What the two-chain test says when it is run on the trial's own difference in arm means, over 24 outcomes on one fourteen-unit set. The set is in 2 mirror components — that is enumerated, not inferred — so every quiet reading is a miss. 29% of them are quiet. The reason is in the enumerated set rather than in the run: how far the two components are apart on a given probe ranges from 0.001 to 4.938 of a within-component spread across these outcomes, a factor of several thousand. Both covariate probes — chosen before any outcome existed, and replaceable if they had been quiet — report the split. An outcome cannot be chosen and cannot be replaced.

A probe nobody chose

On a set that is definitively in two pieces, seven of twenty-four outcomes report nothing at all. Every covariate probe reports it. What separates them is not accuracy — it is that one of them can be chosen and the other is what happened.

after · Randomisation
The same copula, turned over. A Clayton copula and its reflection, at the same Spearman correlation of 0.40 and the same Kendall tau of 0.275, against the covariate's marginal. With a symmetric covariate the two are the same number to nine decimals — 7.707% apiece — because the leak then depends on how much asymmetry the copula has and not on which way it points. Skew the covariate and they come apart: at a skewness of 2.26 the lower-tail copula leaves 3.431% and the upper-tail one 36.213%, a factor of 10.6. Both halves of the dependence are asymmetries and an asymmetry has a direction; a lower-tail copula concentrates the dependence where a right-skewed marginal is compressed and the two distortions partly undo each other, and an upper-tail one concentrates it where the marginal is stretched.

A symmetry that was not enough

A heavy-tailed symmetric covariate has a skewness of zero and leaks exactly nothing under three copulas. Under the two asymmetric ones it doubles the leak, from 7.707% to 14.229%.

compound · Adjustment
A threshold in the tail is a threshold nothing balances. The share of a coin's imbalance in an indicator 1{x > c} that survives a rule which balances the covariate itself. The smooth curve is 1 − ρ² with ρ = φ(c)/√(p(1−p)), a closed form with no trial in it; the points are counted over 500 trials of 200 units at each threshold. At the median the two agree that about a third survives — the removed share is exactly 2/π — and by two standard deviations 86.9% survives. The closed form is exact in the limit and optimistic by a few points at this many units, because the rule balances the sample's mean rather than the population's.

A threshold in the tail

How much of a threshold's imbalance a balanced covariate removes is a correlation, and the correlation is a closed form. At the median it is exactly 2/π — the same 2/π a median split throws away — and two standard deviations out it is an eighth.

shape · Blocking
The one thing a trial always reports is the one thing that survives. How wrong three p-values are when they are computed over the half of the admissible set a single walk can reach, rather than over all of it, at a fourteen-unit trial where the whole set can be enumerated. The two-sided p-value on the difference in arm means — the number a trial publishes — is wrong by exactly nothing, at every row, to machine precision. That is not luck: the two components are complement pairs and the difference in arm means is exactly negated by the complement, so the distribution of its absolute value is the same on both. A one-sided p-value on the same statistic is out by as much as 0.112, and the largest response observed in the treated arm — a safety reading rather than an effect, and the one statistic here that is not odd under the complement — by as much as 0.172. The defect survived because the commonest thing anybody computes is the one quantity it cannot touch.

Half a reference distribution

A walk that reaches half its admissible set reports the two-sided p-value exactly right, to the last digit, for ever. A one-sided one it puts on the wrong side of five per cent about once in thirty.

after · Reference
What a rule reads, against what the outcome uses. The variance of the treatment estimate relative to a coin's, for four things a rule might balance against three shapes the outcome might have, over 350 trials of 200 units. The diagonal is the easy part — a rule that reads the function the outcome uses removes about half the variance. What the table is for is the off-diagonal: reading the covariate alone is worth nothing against a quadratic (0.755), and reading all three is worth nearly as much against every shape as the matching rule is against its own (0.532, 0.493, 0.514).

Three functions of one number

A rule that balances the covariate is exposed to every shape the outcome might have. A rule that balances three functions of it costs two points of variance against the shape the first was built for and takes the worst case from a coin's to about half of it.

shape · Criterion
Where the set stops being one set. How many of the 3,432 equal splits of fourteen units a balancing rule admits, as the tolerance tightens, with the number of components single swaps leave it in. The set falls from 886 to 84 assignments, and somewhere in that fall it stops being connected: at 0.8 it is in 2 pieces and every assignment's complement is in the other one. Nothing about the rule changes at that point and nothing a chain reports changes either, which is the whole difficulty — the acceptance rate, the stationary distribution and the detailed balance are all in order on both sides of it.

Before the trial and after

The same diagnostic run at two moments answers two different questions. Before, a positive verdict changes the design. After, it changes which number gets reported — and only for the numbers the defect can reach.

after · Assignment
What each analysis does at a true null, by shape. Four analyses of the same trials — 500 of them at each shape, 120 units, assigned by the rule that reads the covariate. Every rejection is false. The unadjusted analysis is the one that moves: 1.60% against a linear outcome, where the design removed a great deal that the standard error still prices, and 5.20% against a quadratic, where it removed nothing and the standard error is right. Adjusting holds the level in all three columns, and so does the design's own reference distribution, which needs to be told the rule and nothing else.

The analysis and the shape

An unadjusted analysis after a rule that read the covariate is too cautious — by a third against a linear outcome, by nothing at all against a quadratic. And an adjustment for the wrong function recovers almost none of the precision the right one would.

shape · Randomisation
The one zero neither half of the dependence can touch. A median split's interaction leak at all 30 combinations of copula and marginal, on a log scale. Every one is under 10⁻¹⁶ and the largest is 1.74e-20, which is the quadrature's own noise rather than a leak. The reason is arithmetic and it is short: a centred median split takes the values ±½, so its square is a quarter identically — for every unit, on every draw, whatever the covariate's scale is and whatever joint law the ranks have. The interaction is then orthogonal to both main effects by construction, and there is nothing for either half of the dependence to break. Both of the fields this one joins report this zero holding under their own variation; running both variations at once is what establishes that it is not two coincidences.

The zero that survives both

A median split's interaction leak is under 10⁻¹⁶ at all thirty combinations of copula and marginal. It is the only guarantee in the collection that neither half of the dependence can touch.

compound · Adjustment
Free until the sums stop seeing what the differences see. Coverage with and without the block sums pooled into the interval's variance estimate. With one effect and one level they are free. With an effect that varies between blocks they are still free, because a block's sum picks that variation up exactly as its difference does. With a level that varies they make the interval 37% wider and conservative. And where the effect falls as the level rises — a ceiling, and not an exotic thing to suppose — the sums carry none of the between-block variation while the differences carry all of it, the pooled estimate is short, and the interval that uses it covers 88.75% on a width 20% narrower than the honest one.

What a two-arm rule may not pool

A spread computed "within the block" without the arm label carries a share of the effect, so the trial runs 173 observations at a null and 282 at an effect of 1.5. The stopping rule is reading the thing it exists to measure, and the phrase that produced it is one word long.

contrast · Allocation
What the guess is worth, when it is worth anything. The variance cost of an even split relative to the variance-minimising one for a risk difference, against the first arm's proportion, with the second at 0.3. The cost is a pure number: it does not depend on the trial's size. It is exactly zero at 0.3 and at 0.70, where the two arms have the same p(1 − p); it is 0.19% at a half and 4.36% at a tenth. Across the whole range from a tenth to nine tenths it never exceeds 4.36%, which is what the variance-minimising rule is worth here — and what it is worth is the reason it is safe to use with a guess.

The arm whose variance is its answer

With a binary outcome the allocation rule is a function of the proportions the trial exists to estimate. It costs at most 4.36% of variance to ignore it anywhere between a tenth and nine tenths, because √(p(1−p)) stays within a factor of two of its peak across 98% of the unit interval.

allocation · Allocation
Three contrasts on one dataset, three different splits. The variance-minimising allocation for each of three ways of reporting the same two-arm comparison, against the first arm's proportion, with the second at 0.1. A risk difference wants the arm with the larger p(1 − p) to get more units; a log odds ratio wants it to get fewer, and the two curves are exact reflections of each other in the half line. A log risk ratio wants something else again. At a first-arm proportion of 0.6 they ask for 62.0%, 21.4% and 38.0% of the units. A trial reporting more than one of them cannot be optimal for either.

Two contrasts, one split

A risk difference wants 62.0% of the units in the first arm, a log risk ratio wants 21.4% and a log odds ratio wants 38.0% — on one dataset, with one pair of proportions. The difference's rule and the odds ratio's are exact reflections of each other, so no split can be near-optimal for both.

allocation · Allocation

Named alongside it

The objects these essays reach for when they reach for this one.

Covariate balanceImbalanceAllocation ruleClosed formThresholdVariance reductionAssignment mechanismEfficiencyExperimental designModel misspecificationRandomisation testReference distribution

All concepts