Concept

Rerandomisation — where it appears

Drawing an assignment, checking whether it balances the covariates well enough, and drawing again if not. The reference distribution for any test afterwards is over the admissible assignments rather than all of them, which is what makes the tolerance part of the analysis rather than of the logistics.

Named by 27 essays across 9 fields — each of them below, with the objects they name alongside it.

The rule is parity, and it runs both ways. At a correlation of 0.5, four combinations of a dictionary and an outcome shape. The joint sign flip (X, Y) → (−X, −Y) leaves the bivariate normal alone at every correlation, so a function that changes sign under it is orthogonal to one that does not. A product of two odd functions is even; a product of an odd and an even one is odd. So an odd dictionary removes exactly none of the first and something of the second, and an even dictionary does the reverse — which it does, to machine precision, in both of the two rows that should be zero. This is one rule where there had been two: that a median split's square is constant, and that a polynomial dictionary contains the products a correlation generates.

A dictionary that is neither

A rule handed two median splits removes none of their interaction; a rule handed two covariates removes none of their product. Those were two results with two explanations, and they are one result with one — and finding it corrected the number underneath both.

dict · Criterion
A proposal that moves more, refused more often. The two halves of the trade, both exact, on the 410 admissible assignments of twelve units. The integrated autocorrelation time of an imbalance the rule was never handed falls from 7.30 at one swap to 3.97 at three, and the acceptance rate falls with it, from 58.8% to 40.8%. A rejected proposal costs one evaluation and leaves the chain where it was, so acceptance is not the price of anything and the ranking by acceptance is the reverse of the ranking by cost. Past three the family folds: exchanging k of six from each arm is the complement of exchanging six − k, so k = 5 has the same 36 proposals as k = 1 and k = 6 has 1.

A proposal that moves more than two units

The walk's autocorrelation is a fact about its step size and not about its acceptance rate. Exchanging three units from each arm mixes nearly twice as fast as exchanging one, and is refused a third more often.

blocks · Randomisation
The test, checked where the answer is known. A fourteen-unit trial at eight tolerances. At each one the admissible set is enumerated — 1534, 886, 304, 158, 126, 116, 102, 84 assignments — and its components counted, which is only possible because 3432 equal splits of fourteen units can be walked. The dots are the test, which walks none of them: two chains, one started at an assignment and one at its complement, compared on a statistic the rule was not handed. Filled marks are tolerances the enumeration says leave the set in more than one piece. The test fires on every one of them and on none of the others, 0 misses and 0 false alarms.

A test rather than a survey

A thin admissible set falls into an arrangement and its mirror image, and the walk that samples it is uniform on half the reference distribution for ever. That was found by enumerating fourteen units, and enumeration stops at twenty-four.

reach · Randomisation
The fourth-order expectation, by two routes. Every inner product in an eight-term slice of the dictionary at ρ = 0.5, computed from the linearisation and Mehler's formula and counted from two hundred thousand draws of a correlated pair. The entries that matter are the ones off the main effects: ⟨f(X)u(Y), g(X)v(Y)⟩ is a fourth-order expectation, which the independent-covariate field could not write down. The worst departure is 1.99 standard errors over 36 pairs, measured in each pair's own error because the entries differ in size by two orders of magnitude.

The fourth moment that was missing

Mehler's formula makes the main effects exact at any correlation and stops there, because the interactions need an expectation of four Hermite functions rather than two. A linearisation turns the four into two, and the whole geometry becomes closed again.

joint · Blocking
Two diagnostics, one answer, two different moments. The two-chain statistic on a covariate probe and on the trial's own difference in arm means, at seven tolerances of a fourteen-unit rule, against the enumerated truth. Both are quiet wherever the set is one set and both fire wherever it is not, at every tolerance — which is what says the outcome probe is the same test rather than a resemblance of it. The difference between them is not accuracy and it is not power. It is when: the covariate probe can be run before a single outcome exists, when a practitioner can still loosen the rule or change the sampler, and it can be run again on a different function if it comes back quiet. The outcome probe runs after the trial, on the one column the trial produced, and what it can do with a positive verdict is repair the p-value rather than the design.

The statistic the p-value is about

The test for whether a balanced-assignment walk reaches its whole set is run on a covariate function chosen before the trial. Run on the difference in arm means it is the same test, and it is about the number the trial publishes.

after · Randomisation
What the rule blocks is not where it splits. How much of the separating direction each probe carries, over 100 designs of 14 units whose admissible set is enumerated and split into two pieces. The deferral this field answers proposed the constraint's active set — which exchanges the tolerance box actually blocks — as a better probe than the design's own leverage, on the ground that leverage is a heuristic and the active set is the quantity. Modelled from the design and the tolerance, it reads 0.4272 against leverage's 0.5395, at 4.43 paired standard errors the wrong way. Counted exactly over the enumerated set — at a cost no trial can pay — it reads 0.3854, worse again. Both beat a random direction at 0.2622, so they are probes; neither beats the two the earlier field already had.

What the rule blocks

A balancing rule breaks the admissible set into pieces by refusing exchanges. Which exchanges it refuses is computable from the design and the tolerance alone, before any assignment exists — and it makes a probe.

blocked · Randomisation
A model of the active set, and the active set. Each of the 14 units of one design, at the share of its exchanges the tolerance box blocks — computed from the design's columns and the tolerance under a uniform position in the box, against counted over all 116 admissible assignments. The diagonal is where the two would agree. Over 192 designs they agree about the ordering of the units at a correlation of 0.8141 ± 0.0112, negative on 0.5% of them, and disagree about the level: 0.8442 counted against 0.8170 modelled, a gap of 0.0272 ± 0.0051. An admissible assignment does not sit uniformly in its box, and this is the size of that.

A model and a count

The share of a unit's exchanges a tolerance box refuses can be modelled from the design or counted over the admissible set. They order the units the same way at a correlation of 0.81 and disagree about the level by 0.027.

blocked · Randomisation
The probe a trial has is the probe a trial got. What the two-chain test says when it is run on the trial's own difference in arm means, over 24 outcomes on one fourteen-unit set. The set is in 2 mirror components — that is enumerated, not inferred — so every quiet reading is a miss. 29% of them are quiet. The reason is in the enumerated set rather than in the run: how far the two components are apart on a given probe ranges from 0.001 to 4.938 of a within-component spread across these outcomes, a factor of several thousand. Both covariate probes — chosen before any outcome existed, and replaceable if they had been quiet — report the split. An outcome cannot be chosen and cannot be replaced.

A probe nobody chose

On a set that is definitively in two pieces, seven of twenty-four outcomes report nothing at all. Every covariate probe reports it. What separates them is not accuracy — it is that one of them can be chosen and the other is what happened.

after · Randomisation
Where the two kinds of cut sit. Six covariates, each a monotone transformation of the same latent normal. The vertical line at zero is where every median split sits, on every one of them, because a monotone map preserves order: the median of the covariate is the image of the median of the latent normal. The marks on the curves are where a threshold at 1 on the covariate's scale falls — 1.000, 0.881, 0.875, 0.783, 0.713, 0.337 — and none of them is at zero. That is the whole of the difference. A function of the sign of the latent normal is odd, and a rule made of odd functions removes exactly nothing of an interaction between two of them; a threshold anywhere else is neither odd nor even and removes something.

A split survives what a mean does not

The two things every trial balances come apart on a skewed covariate. A median split is a function of the sign of the latent normal whatever the marginal is; a mean is not, and its exact zero is gone at a skewness of one.

skew · Criterion
The zero was a fact about independence. What a balancing rule handed every main effect of both covariates removes of a pure interaction, as the covariates are allowed to move together. At ρ = 0 it is exactly nothing — at machine precision, at any number of main effects — which is the independent-covariate result and is correct. It is not small anywhere else: the product of the two covariates loses 64.0% of itself by ρ = 0.5, because h₁h₁ = h₀ + √2·h₂ and Mehler pairs h₂ with h₂ at ρ². Four interactions are drawn and none of them keeps the zero.

A zero that was an assumption

A rule handed every main effect of both covariates removes exactly none of a pure interaction. That is true at machine precision, it is a fact about independence, and it dies as the square of the correlation.

joint · Criterion
Stationary is not the same as convergent. How far each k-swap walk is from uniform after t steps, started at the least balanced admissible assignment of 410. Every one of these chains has a symmetric proposal and rejects by standing still, so every one of them is doubly stochastic and every one preserves the uniform distribution exactly. Only five of the six get there. Exchanging all six units of each arm is a single proposal — the complement — and the admissible set is closed under complement, so the walk takes it every time and oscillates between two assignments for ever: after 160 steps it has visited 1 state and sits 0.9976 from uniform. Its stationary distribution is a fact about the matrix; its limit does not exist.

Stationary is not convergent

A walk that exchanges every unit in each arm preserves the uniform distribution exactly and never gets near it. Every doubly stochastic matrix has the same stationary distribution; only some of them have a limit.

blocks · Randomisation
What the worst case is worth, one function at a time. The smallest share each dictionary removes, over seven outcome shapes, at a correlation of 0.5. A rule balancing the mean of each covariate has a worst case of exactly zero — against the square, and against both products. Adding a median split to it, which is the second thing every trial balances, leaves the worst case at exactly zero, because a median split is odd and so is a mean. Adding the square instead moves it to 6.8%, and the extra functions after that move it to 7.4%. The worst case is decided by which parities the dictionary contains rather than by how many functions are in it.

What the extra function buys

A rule balancing the mean of each covariate has a worst case of exactly zero. Adding the median split — the other thing every trial balances — leaves it at exactly zero, and one square moves it.

dict · Criterion
A rate that does not know how large the trial is. The share of equal splits admitted by a tolerance of 1 coin-spreads on 3 functions, at six trial sizes. The first two are exact — 12,870 and 184,756 splits, walked, averaged over eight draws of the units — and the rest are sampled. From a hundred units on, the rate sits on (2Φ(1) − 1)^3 = 0.3182, which contains no n at all. The two small trials are 29.2% and 27.7% short of it, so the sixteen-unit measurement understates the rate rather than bracketing it. Meanwhile the admissible count — the rate times C(n, n/2) — goes from 2^11.5 to 2^393.7: the exhaustion a small trial runs into is a fact about small trials.

A count that has to be estimated

At sixteen units the admissible assignments can be counted by walking all 12,870 of them. At four hundred there are about 2^393.70, and the share admitted is 0.31885 against a closed form of 0.31818 that has no trial size in it at all. The exhaustion a small trial runs into is a fact about small trials.

product · Randomisation
Thinness stays put and reachability does not. The same balancing rule and the same tolerance at five trial sizes. The admitted share barely moves — one admissible assignment in 27, 30, 30, 20, 18 — because the acceptance rate of a rerandomisation is a fact about the basis rather than about the number of units. What does move is the number of single swaps available: 36, 49, 64, 81, 100, growing like a quarter of the square of the trial size. Filled marks are sizes whose admissible set falls into more than one piece. The set is in 4 pieces at 12 units, 2 at 14, and one piece from 16 upwards. The fourteen-unit result is a statement about fourteen units.

A defect that is about size

The admitted share of a rerandomisation barely moves with the number of units. The number of admissible neighbours grows like the square of it, and that is what decides whether the walk can go everywhere.

reach · Randomisation
What each probe can see. How far apart the two components of the admissible set are on each probe, over the spread inside a component, on 100 designs whose set is enumerated and split. It is the population quantity a chain is trying to report. The separating direction itself reads 10.5646; the projected fourth power 5.0800, the design's own leverage 3.8362, the modelled active set 1.9529, the counted active set 2.0170 and a random direction in the same subspace 0.9422. The two active-set probes beat the random direction and lose to both of the earlier field's, which is the field's answer to the question that opened it.

A quantity that loses to a heuristic

Leverage is a heuristic about which units a balancing rule has most to say about. The constraint's active set is the thing the rule actually does. As a probe, the heuristic wins by 4.4 paired standard errors.

blocked · Randomisation
Whichever dial made the set thin, the crossing is at the same thinness. Each curve is one dictionary, swept over eight tolerances at two hundred units; a point above the line is a set thin enough that walking beats hunting. The curves lie nearly on top of one another, which is the answer to whether the crossing is a fact about the tolerance or about the thinness it produces: the crossings sit between one admissible assignment in 176 and one in 268 for dictionaries of 3 to 6 functions. The mechanism is that a hunt costs exactly 1/p and a walk costs almost the same everywhere — between 82 and 394 evaluations per usable draw across the whole table — so the crossing is wherever 1/p reaches a number that does not move.

The set a dictionary leaves

A rule constrained on six functions at a loose tolerance leaves a set as thin as one constrained on three at a tight one. Both sampling methods cross over at the same thinness, and the tolerance where that happens moves by a factor of three.

dict · Assignment
How long a walk has to be given. Every equal split of twelve units is enumerated, the 410 admissible ones are found, the transition matrix is built, and the distance from uniform is computed exactly at each step — no simulation anywhere. The walk is started at the least balanced admissible assignment, which is the state a rejection sampler is least likely to have handed it and the one a burn-in has to cover. It is 0.0849 away after twenty steps and 0.00008 after ninety. A real cost, and a small one, and naming it is what stops it being assumed to be zero.

Walking the admissible set

A rerandomisation test hunts for admissible assignments and throws away the rest. A walk visits them instead — and it is exactly uniform only because it stands still when a proposal fails, which is the step that looks like waste.

joint · Randomisation
The crossing barely moves. Both methods' costs in one unit — assignments evaluated per usable draw — as the tolerance tightens. A hunt costs 1/p and rises without limit: from 2.22 at a tolerance of 1.2 to 357.14 at 0.18. A walk costs its autocorrelation time and barely moves. The two cross at a tolerance of 0.190 at one swap and 0.195 at eight — the whole family of proposal sizes crosses inside a band of about two hundredths, because where the crossing is, the large proposal has already lost its advantage. A multi-swap proposal is worth a factor of 5.65 in the regime where the walk should not be used at all.

Where the gain is, and where the decision is

A bigger proposal is worth a factor of six at a loose tolerance and nothing at a tight one. The tolerances where it helps are the ones where a hunt costs two evaluations a draw, and the crossing barely moves.

blocks · Assignment
A guarantee that stops being a number. The worst case of each dictionary over six outcome shapes, at a correlation of 0.5, against the skewness of the covariate. Under a symmetric marginal every rule made of odd functions has a worst case of exactly zero, and the rule holding a mean and a median split of each covariate — the two things every trial balances — is one of them. Under skew that zero becomes 0.74%, 1.83%, 2.24%, 2.49%: small numbers, each of which depends on a marginal nobody stated. The guarantee has not improved by becoming positive. It has stopped being a guarantee, because it can no longer be written down without the covariate's distribution in it.

Balancing a skewed covariate

The worst case of the rule every trial runs goes from exactly zero to somewhere between a quarter of a per cent and two and a half. Which is small, and is a number that cannot be stated without the covariate's distribution in it.

skew · Criterion
Where the set stops being one set. How many of the 3,432 equal splits of fourteen units a balancing rule admits, as the tolerance tightens, with the number of components single swaps leave it in. The set falls from 886 to 84 assignments, and somewhere in that fall it stops being connected: at 0.8 it is in 2 pieces and every assignment's complement is in the other one. Nothing about the rule changes at that point and nothing a chain reports changes either, which is the whole difficulty — the acceptance rate, the stationary distribution and the detailed balance are all in order on both sides of it.

Before the trial and after

The same diagnostic run at two moments answers two different questions. Before, a positive verdict changes the design. After, it changes which number gets reported — and only for the numbers the defect can reach.

after · Assignment
One of them is mostly leverage. How much of the design's own leverage direction each active-set probe carries, once both are standardised and projected off the rule's span — which is what a probe is, so it is the comparison that matters. Over 189 designs the modelled active set agrees with leverage at |r| = 0.8359 ± 0.0114 and the counted one at 0.4239 ± 0.0216. So the modelled probe is largely leverage under another name and the counted one is genuinely a different direction — and the counted one is the worse probe, at 0.3854 of alignment against 0.4272. What the active set contains beyond leverage points away from where the set splits.

Counting it exactly does not help

If a modelled active set lost because the model was crude, the exact one would win. It is computed at a cost no trial can pay, and it is worse — so the approximation was never what was costing the probe.

blocked · Randomisation
Where a walk is cheaper than a hunt. Both costs in the same unit. A rejection sampler evaluates 1/p assignments per independent draw and does not care how large the trial is; a walk evaluates one per step and yields an effective draw every τ steps, and τ is a property of the constraint and the statistic together. They cross at a tolerance of 0.194 standard deviations, where about one assignment in 396 is admissible — far tighter than any trial is designed at. And the walk does not remove the acceptance cost; it pays it once, hunting for somewhere to start.

Draws that repeat each other

A hunt costs 1/p evaluations per independent draw. A walk costs one per step and yields an effective draw every τ steps. Both are counted in the same unit, and the walk is dearer at every tolerance a trial is designed at.

joint · Reference
What the diagnostic says at two hundred units. The same test run 8 times on independent streams, at five tolerances of a two-hundred-unit trial, 40,000 steps each. At the loosest tolerance every run says the same thing — the walk reaches the whole set — and it keeps saying it as the set is thinned. Past a point the runs stop agreeing with each other: at the tightest tolerance here 6 of 8 report that the chains have not mixed and 2 report a split, which is a diagnostic disagreeing with itself rather than a property of the set. That disagreement is the honest answer at this size, and it is one nothing in this collection could give before: an enumeration stops at about twenty-four units.

The diagnostic at two hundred

Pointed at a trial size no enumeration reaches, the test gives three answers rather than one — and past a certain thinness it stops agreeing with itself, which is the honest reading and the one nothing could give before.

reach · Randomisation
A thin enough set is not one set. Every admissible set of 14 units this table can enumerate, by how much of the assignment space it admits and how many pieces it falls into under single swaps. A walk is uniform on the piece it starts in and never leaves it. The pieces are not fragments: at 522 admissible assignments the set splits into 3 halves of exactly 520 each, and every assignment's complement is in the other half — no sequence of admissible single swaps takes an assignment to its own mirror image. Two-swap proposals reconnect four of the six disconnected sets here — the two they do not are the thinnest, where a two-unit move rarely lands anywhere admissible either — which makes a bigger proposal a correctness repair rather than the speed dial it was measured as.

The walk that cannot cross

A thin enough admissible set is not one set. It splits into an assignment and its mirror image, no sequence of admissible single swaps joins them, and the walk that samples it is uniform on half the reference distribution for ever.

dict · Randomisation
The rate falls geometrically; the count does not fall at all. The share of equal splits of two hundred units that a tolerance of 1 coin-spreads admits, against the number of functions the tolerance is stated for, with the closed form (2Φ(1) − 1)^k drawn beside it. The rate falls by about two thirds with every constraint. The admissible count is that rate times C(200, 100), and it goes from 2^195 to 2^192 — it does not fall in any sense a trial cares about. What the falling rate costs is sampling: 9,878 draws to collect a thousand admissible ones at six constraints, against 1,465 at one.

What a reference distribution costs to sample

A randomisation test on a trial too large to enumerate has to sample its reference distribution, at 1/p attempts per draw and a p-value resolved to 1/(B + 1). Six constraints cost 9,878 attempts per thousand draws, and a thousand draws resolve p to 9.99·10⁻⁴ and not one digit finer.

product · Reference
Where the constraints exhaust the randomisation. At 16 units there are 12,870 equal splits, so the ones meeting a stated tolerance can be counted rather than estimated. With each of the first k standardised imbalances required to be within 0.4 of a coin's own spread, the admissible count runs 3874 → 1006 → 314 → 0 → 0 → 0 — and at 4 functions there is no admissible assignment at all. The count is the number of distinct answers a randomisation test can give: at 3 functions its finest attainable p-value is 1 in 314. Balance improves with every constraint and the reference distribution shrinks with it, and the two run out at different rates.

When the constraints run out

Every function added to a basis is a constraint the assignment has to satisfy with the same units. At sixteen units and a stated tolerance the admissible assignments run 3,874, then 1,006, then 314, then none — and the count is exact, because the assignment space is finite.

basis · Allocation
The pairing recovers most of it and passes nothing. How much of the separating direction each probe carries, over 100 designs of 14 units whose admissible set is enumerated and split in two. The four rows the pairing adds are the dominant direction of what each blocking matrix keeps past its degrees, and the cut that direction's signs induce. Counted, they read 0.5266 and 0.5258 against the counted per-unit share's 0.3854 — most of the gap between that share and the design's own leverage at 0.5395, closed. Modelled, they read 0.4274 and 0.4954 against 0.4272. Nothing built from the active set passes leverage, and the projected fourth power is still ahead of all of them at 0.6583.

A set of pairs, not a vector

The active set is a graph on the units, and every probe built from it so far has been its degree. Read as a graph it recovers 0.1326 of the alignment the summary lost — and draws level with leverage rather than passing it.

blocked · Randomisation

Named alongside it

The objects these essays reach for when they reach for this one.

Covariate balanceRandomisation testImbalanceReference distributionAssignment mechanismMarkov chain Monte CarloExact enumerationExperimental designConnected componentAcceptance rateBasis functionsCombinatorial search

All concepts