Series

Randomisation — the series

24 essays on one idea, from the one that introduces it to the one that assumes the rest.
  1. Every way of splitting 16 units into two halves. All 12,870 assignments, enumerated. The spread of the standardised imbalance is exactly 2/√n = 0.500, whatever the covariate's own distribution, and 33.3% of assignments differ by more than 0.5 standard deviations. Randomisation does not deliver balance; it delivers a known distribution of imbalance.

    Randomisation is not balance

    A third of all ways to split sixteen units leave the two halves more than half a standard deviation apart on a covariate. What randomisation delivers is not balance but a known reference distribution — and it makes a test exact with no assumption about the data's shape at all.

    part 1 · design
  2. 20 adaptive trials, 45% against 25%. Each line is one trial allocating patients one at a time by the arm's own posterior. The average final share on the better arm is 84.7%, with a standard deviation of 10.3 points across these 20 trials. The rule does not deliver a fixed advantage: it delivers one that depends on how the first few patients came out.

    Randomising towards the winner

    Allocating more patients to the arm that is doing better is the humane thing to want and it buys nothing statistically: at a fixed total it costs thirty points of power. And because the allocation is a function of the outcomes, the ordinary test on it rejects a true null 7.8% of the time before any time trend is applied — and 58% after one.

    part 2 · adaptive
  3. Three analyses of the same trials, none of them wrong about the data. 320 trials at n = 60 with no treatment effect at all, so every rejection counted is a false one, and a covariate that drives the outcome with coefficient 1. The unadjusted comparison is at 5.94% after a coin — its level — and at 0.00% after the rule that reads the covariate: the design removed the imbalance and the analysis is still pricing it. Adjusting for the covariate gives 4.06%, and the rule's own reference distribution — hold the outcomes, re-run the rule 199 times, count — gives 3.13% against the 4.5% that 199 draws can deliver. The last of the three has to be told the assignment rule and nothing else, which is the one thing the experimenter certainly knows.

    What the balanced trial is worth

    A rule that reads the covariate removes three quarters of the imbalance. An analysis that does not know it happened prices the imbalance anyway, rejects one true null in two hundred instead of one in twenty, and finds a real effect less often than a coin-tossed trial does.

    part 3 · continuous
  4. What each analysis does at a true null, by shape. Four analyses of the same trials — 500 of them at each shape, 120 units, assigned by the rule that reads the covariate. Every rejection is false. The unadjusted analysis is the one that moves: 1.60% against a linear outcome, where the design removed a great deal that the standard error still prices, and 5.20% against a quadratic, where it removed nothing and the standard error is right. Adjusting holds the level in all three columns, and so does the design's own reference distribution, which needs to be told the rule and nothing else.

    The analysis and the shape

    An unadjusted analysis after a rule that read the covariate is too cautious — by a third against a linear outcome, by nothing at all against a quadratic. And an adjustment for the wrong function recovers almost none of the precision the right one would.

    part 4 · shape
  5. A class that is a subspace has no guarantee below its own dimension. Each cell is the worst case over every unit-variance function in a class of dimension m, for a rule reading k functions: the smallest squared principal-angle cosine between the two subspaces. Wherever k is less than m the number is zero to machine precision, and that is not a weak guarantee but the absence of one — some direction of the class is orthogonal to the entire basis, and against an outcome in that direction the rule does exactly what a coin does. An experimenter who declines to name the shapes and asks instead to be protected against everything smooth is asking for the cells above the diagonal.

    Where the guarantee is exactly zero

    An experimenter who declines to name the shapes, and asks instead to be protected against anything in a class, is asking for a number that is not small but zero. Bounding the class is unavoidable, and the two ways of doing it choose different bases.

    part 5 · basis
  6. A rate that does not know how large the trial is. The share of equal splits admitted by a tolerance of 1 coin-spreads on 3 functions, at six trial sizes. The first two are exact — 12,870 and 184,756 splits, walked, averaged over eight draws of the units — and the rest are sampled. From a hundred units on, the rate sits on (2Φ(1) − 1)^3 = 0.3182, which contains no n at all. The two small trials are 29.2% and 27.7% short of it, so the sixteen-unit measurement understates the rate rather than bracketing it. Meanwhile the admissible count — the rate times C(n, n/2) — goes from 2^11.5 to 2^393.7: the exhaustion a small trial runs into is a fact about small trials.

    A count that has to be estimated

    At sixteen units the admissible assignments can be counted by walking all 12,870 of them. At four hundred there are about 2^393.70, and the share admitted is 0.31885 against a closed form of 0.31818 that has no trial size in it at all. The exhaustion a small trial runs into is a fact about small trials.

    part 6 · product
  7. How long a walk has to be given. Every equal split of twelve units is enumerated, the 410 admissible ones are found, the transition matrix is built, and the distance from uniform is computed exactly at each step — no simulation anywhere. The walk is started at the least balanced admissible assignment, which is the state a rejection sampler is least likely to have handed it and the one a burn-in has to cover. It is 0.0849 away after twenty steps and 0.00008 after ninety. A real cost, and a small one, and naming it is what stops it being assumed to be zero.

    Walking the admissible set

    A rerandomisation test hunts for admissible assignments and throws away the rest. A walk visits them instead — and it is exactly uniform only because it stands still when a proposal fails, which is the step that looks like waste.

    part 7 · joint
  8. A proposal that moves more, refused more often. The two halves of the trade, both exact, on the 410 admissible assignments of twelve units. The integrated autocorrelation time of an imbalance the rule was never handed falls from 7.30 at one swap to 3.97 at three, and the acceptance rate falls with it, from 58.8% to 40.8%. A rejected proposal costs one evaluation and leaves the chain where it was, so acceptance is not the price of anything and the ranking by acceptance is the reverse of the ranking by cost. Past three the family folds: exchanging k of six from each arm is the complement of exchanging six − k, so k = 5 has the same 36 proposals as k = 1 and k = 6 has 1.

    A proposal that moves more than two units

    The walk's autocorrelation is a fact about its step size and not about its acceptance rate. Exchanging three units from each arm mixes nearly twice as fast as exchanging one, and is refused a third more often.

    part 8 · blocks
  9. Stationary is not the same as convergent. How far each k-swap walk is from uniform after t steps, started at the least balanced admissible assignment of 410. Every one of these chains has a symmetric proposal and rejects by standing still, so every one of them is doubly stochastic and every one preserves the uniform distribution exactly. Only five of the six get there. Exchanging all six units of each arm is a single proposal — the complement — and the admissible set is closed under complement, so the walk takes it every time and oscillates between two assignments for ever: after 160 steps it has visited 1 state and sits 0.9976 from uniform. Its stationary distribution is a fact about the matrix; its limit does not exist.

    Stationary is not convergent

    A walk that exchanges every unit in each arm preserves the uniform distribution exactly and never gets near it. Every doubly stochastic matrix has the same stationary distribution; only some of them have a limit.

    part 9 · blocks
  10. A thin enough set is not one set. Every admissible set of 14 units this table can enumerate, by how much of the assignment space it admits and how many pieces it falls into under single swaps. A walk is uniform on the piece it starts in and never leaves it. The pieces are not fragments: at 522 admissible assignments the set splits into 3 halves of exactly 520 each, and every assignment's complement is in the other half — no sequence of admissible single swaps takes an assignment to its own mirror image. Two-swap proposals reconnect four of the six disconnected sets here — the two they do not are the thinnest, where a two-unit move rarely lands anywhere admissible either — which makes a bigger proposal a correctness repair rather than the speed dial it was measured as.

    The walk that cannot cross

    A thin enough admissible set is not one set. It splits into an assignment and its mirror image, no sequence of admissible single swaps joins them, and the walk that samples it is uniform on half the reference distribution for ever.

    part 10 · dict
  11. The test, checked where the answer is known. A fourteen-unit trial at eight tolerances. At each one the admissible set is enumerated — 1534, 886, 304, 158, 126, 116, 102, 84 assignments — and its components counted, which is only possible because 3432 equal splits of fourteen units can be walked. The dots are the test, which walks none of them: two chains, one started at an assignment and one at its complement, compared on a statistic the rule was not handed. Filled marks are tolerances the enumeration says leave the set in more than one piece. The test fires on every one of them and on none of the others, 0 misses and 0 false alarms.

    A test rather than a survey

    A thin admissible set falls into an arrangement and its mirror image, and the walk that samples it is uniform on half the reference distribution for ever. That was found by enumerating fourteen units, and enumeration stops at twenty-four.

    part 11 · reach
  12. What is left of a probe after the rule has had it. The share of each dictionary function a rule balancing x, x2, x3, cut0 has already taken, on trials of 14 units, averaged over 100 designs. Four of the eight functions are the basis, so their share is exactly one: a randomisation test run on one of them is asking about a quantity the rule forced to zero, and one of them is the default probe of the field this measurement comes from. The four that are not still read 0.919, 0.873, 0.903, 0.832 — between 0.832 and 0.919 of them is inside the span — against closed-form removed shares of 0.000, 0.692, 0.590, 0.692. At 14 units a rule with four functions in it takes most of anything it is shown.

    The part the rule already took

    A diagnostic that reports on what a balancing rule was not handed is run through a column that is 92% inside the span the rule balanced — because orthogonality in the population is not orthogonality on fourteen units.

    part 12 · aimed
  13. Two diagnostics, one answer, two different moments. The two-chain statistic on a covariate probe and on the trial's own difference in arm means, at seven tolerances of a fourteen-unit rule, against the enumerated truth. Both are quiet wherever the set is one set and both fire wherever it is not, at every tolerance — which is what says the outcome probe is the same test rather than a resemblance of it. The difference between them is not accuracy and it is not power. It is when: the covariate probe can be run before a single outcome exists, when a practitioner can still loosen the rule or change the sampler, and it can be run again on a different function if it comes back quiet. The outcome probe runs after the trial, on the one column the trial produced, and what it can do with a positive verdict is repair the p-value rather than the design.

    The statistic the p-value is about

    The test for whether a balanced-assignment walk reaches its whole set is run on a covariate function chosen before the trial. Run on the difference in arm means it is the same test, and it is about the number the trial publishes.

    part 12 · after
  14. Three readings, one verdict. Every tolerance of a fourteen-unit trial, with three comparisons on each. The first is between a chain started at an assignment and a chain started at its complement, which is what a mirror split separates. The second is between two chains started at the same assignment on different streams, which nothing about the set can separate — so a large reading there says the run is too short and not that the set is in pieces. The third is the same comparison on the statistic's absolute value, which is symmetric under the complement and therefore blind to the split by construction. The verdict is the pattern rather than any one line: the split is called only where the first fires and the other two do not, which happens at exactly the tolerances the enumeration calls disconnected — 0.8, 0.75, 0.7.

    The statistic that changes sign

    A test for an unreachable half needs a quantity that tells one half from the other. Every symmetric reading of a mirror pair is identical, and a magnitude is the natural thing to reach for.

    part 12 · reach
  15. How far apart the two components are, on each probe. The median separation between the two components of the admissible set — the difference in their mean probe values, over the spread inside a component — over the 100 of 200 designs whose set is enumerated and found split. The separating direction carries 10.565 and needs the enumeration. The fourth power as the earlier fields use it carries 1.543; projected off the span the rule balances, 5.080. The design's own leverage, which uses no dictionary and no outcome, carries 3.836. A random direction in the same subspace carries 0.942, and a direction chosen by looking for concentrated structure carries 0.543 — below random, and the one heuristic here that is worse than not choosing at all.

    A probe chosen from the design

    The design's own leverage aligns with the separating direction four times better than a random direction in the same subspace. The concentrated direction the argument invites is worse than random.

    part 13 · aimed
  16. The probe a trial has is the probe a trial got. What the two-chain test says when it is run on the trial's own difference in arm means, over 24 outcomes on one fourteen-unit set. The set is in 2 mirror components — that is enumerated, not inferred — so every quiet reading is a miss. 29% of them are quiet. The reason is in the enumerated set rather than in the run: how far the two components are apart on a given probe ranges from 0.001 to 4.938 of a within-component spread across these outcomes, a factor of several thousand. Both covariate probes — chosen before any outcome existed, and replaceable if they had been quiet — report the split. An outcome cannot be chosen and cannot be replaced.

    A probe nobody chose

    On a set that is definitively in two pieces, seven of twenty-four outcomes report nothing at all. Every covariate probe reports it. What separates them is not accuracy — it is that one of them can be chosen and the other is what happened.

    part 13 · after
  17. Thinness stays put and reachability does not. The same balancing rule and the same tolerance at five trial sizes. The admitted share barely moves — one admissible assignment in 27, 30, 30, 20, 18 — because the acceptance rate of a rerandomisation is a fact about the basis rather than about the number of units. What does move is the number of single swaps available: 36, 49, 64, 81, 100, growing like a quarter of the square of the trial size. Filled marks are sizes whose admissible set falls into more than one piece. The set is in 4 pieces at 12 units, 2 at 14, and one piece from 16 upwards. The fourteen-unit result is a statement about fourteen units.

    A defect that is about size

    The admitted share of a rerandomisation barely moves with the number of units. The number of admissible neighbours grows like the square of it, and that is what decides whether the walk can go everywhere.

    part 13 · reach
  18. How often each probe finds a split that is there. The share of 34 designs — every one of them enumerated to be in two components — on which a two-chain test of 800 draws declares the split, by probe. The fourth power as the earlier fields use it finds it on 55.9%, so it misses 44.1% of the sets that have one. The same column projected off the rule's span finds it on 88.2%, and the separating direction itself on 91.2%. The design's own leverage, chosen without any dictionary, gets 79.4%. A random direction in the same subspace gets 44.1%, and the direction chosen for being concentrated gets 38.2% — worse than random, which is what a heuristic that finds the wrong structure looks like from the outside.

    What a chosen probe finds

    On a chain of eight hundred draws the probe the earlier fields use misses 44% of the sets that are split. Its own residual off the rule's span misses 12%, for one least-squares fit.

    part 14 · aimed
  19. What the diagnostic says at two hundred units. The same test run 8 times on independent streams, at five tolerances of a two-hundred-unit trial, 40,000 steps each. At the loosest tolerance every run says the same thing — the walk reaches the whole set — and it keeps saying it as the set is thinned. Past a point the runs stop agreeing with each other: at the tightest tolerance here 6 of 8 report that the chains have not mixed and 2 report a split, which is a diagnostic disagreeing with itself rather than a property of the set. That disagreement is the honest answer at this size, and it is one nothing in this collection could give before: an enumeration stops at about twenty-four units.

    The diagnostic at two hundred

    Pointed at a trial size no enumeration reaches, the test gives three answers rather than one — and past a certain thinness it stops agreeing with itself, which is the honest reading and the one nothing could give before.

    part 14 · reach
  20. What the rule blocks is not where it splits. How much of the separating direction each probe carries, over 100 designs of 14 units whose admissible set is enumerated and split into two pieces. The deferral this field answers proposed the constraint's active set — which exchanges the tolerance box actually blocks — as a better probe than the design's own leverage, on the ground that leverage is a heuristic and the active set is the quantity. Modelled from the design and the tolerance, it reads 0.4272 against leverage's 0.5395, at 4.43 paired standard errors the wrong way. Counted exactly over the enumerated set — at a cost no trial can pay — it reads 0.3854, worse again. Both beat a random direction at 0.2622, so they are probes; neither beats the two the earlier field already had.

    What the rule blocks

    A balancing rule breaks the admissible set into pieces by refusing exchanges. Which exchanges it refuses is computable from the design and the tolerance alone, before any assignment exists — and it makes a probe.

    part 15 · blocked
  21. A model of the active set, and the active set. Each of the 14 units of one design, at the share of its exchanges the tolerance box blocks — computed from the design's columns and the tolerance under a uniform position in the box, against counted over all 116 admissible assignments. The diagonal is where the two would agree. Over 192 designs they agree about the ordering of the units at a correlation of 0.8141 ± 0.0112, negative on 0.5% of them, and disagree about the level: 0.8442 counted against 0.8170 modelled, a gap of 0.0272 ± 0.0051. An admissible assignment does not sit uniformly in its box, and this is the size of that.

    A model and a count

    The share of a unit's exchanges a tolerance box refuses can be modelled from the design or counted over the admissible set. They order the units the same way at a correlation of 0.81 and disagree about the level by 0.027.

    part 16 · blocked
  22. What each probe can see. How far apart the two components of the admissible set are on each probe, over the spread inside a component, on 100 designs whose set is enumerated and split. It is the population quantity a chain is trying to report. The separating direction itself reads 10.5646; the projected fourth power 5.0800, the design's own leverage 3.8362, the modelled active set 1.9529, the counted active set 2.0170 and a random direction in the same subspace 0.9422. The two active-set probes beat the random direction and lose to both of the earlier field's, which is the field's answer to the question that opened it.

    A quantity that loses to a heuristic

    Leverage is a heuristic about which units a balancing rule has most to say about. The constraint's active set is the thing the rule actually does. As a probe, the heuristic wins by 4.4 paired standard errors.

    part 17 · blocked
  23. One of them is mostly leverage. How much of the design's own leverage direction each active-set probe carries, once both are standardised and projected off the rule's span — which is what a probe is, so it is the comparison that matters. Over 189 designs the modelled active set agrees with leverage at |r| = 0.8359 ± 0.0114 and the counted one at 0.4239 ± 0.0216. So the modelled probe is largely leverage under another name and the counted one is genuinely a different direction — and the counted one is the worse probe, at 0.3854 of alignment against 0.4272. What the active set contains beyond leverage points away from where the set splits.

    Counting it exactly does not help

    If a modelled active set lost because the model was crude, the exact one would win. It is computed at a cost no trial can pay, and it is worse — so the approximation was never what was costing the probe.

    part 18 · blocked
  24. The pairing recovers most of it and passes nothing. How much of the separating direction each probe carries, over 100 designs of 14 units whose admissible set is enumerated and split in two. The four rows the pairing adds are the dominant direction of what each blocking matrix keeps past its degrees, and the cut that direction's signs induce. Counted, they read 0.5266 and 0.5258 against the counted per-unit share's 0.3854 — most of the gap between that share and the design's own leverage at 0.5395, closed. Modelled, they read 0.4274 and 0.4954 against 0.4272. Nothing built from the active set passes leverage, and the projected fourth power is still ahead of all of them at 0.6583.

    A set of pairs, not a vector

    The active set is a graph on the units, and every probe built from it so far has been its degree. Read as a graph it recovers 0.1326 of the alignment the summary lost — and draws level with leverage rather than passing it.

    part 19 · blocked

All series