Concept

Exact enumeration — where it appears

Walking every possibility rather than sampling from them, which turns a probability into a count. It is available only where the space is small — every equal split of fourteen units is 3,432 arrangements, and of two hundred units it is more than there are atoms in anything.

Named by 22 essays across 11 fields — each of them below, with the objects they name alongside it.

The test, checked where the answer is known. A fourteen-unit trial at eight tolerances. At each one the admissible set is enumerated — 1534, 886, 304, 158, 126, 116, 102, 84 assignments — and its components counted, which is only possible because 3432 equal splits of fourteen units can be walked. The dots are the test, which walks none of them: two chains, one started at an assignment and one at its complement, compared on a statistic the rule was not handed. Filled marks are tolerances the enumeration says leave the set in more than one piece. The test fires on every one of them and on none of the others, 0 misses and 0 false alarms.

A test rather than a survey

A thin admissible set falls into an arrangement and its mirror image, and the walk that samples it is uniform on half the reference distribution for ever. That was found by enumerating fourteen units, and enumeration stops at twenty-four.

reach · Randomisation
What is left of a probe after the rule has had it. The share of each dictionary function a rule balancing x, x2, x3, cut0 has already taken, on trials of 14 units, averaged over 100 designs. Four of the eight functions are the basis, so their share is exactly one: a randomisation test run on one of them is asking about a quantity the rule forced to zero, and one of them is the default probe of the field this measurement comes from. The four that are not still read 0.919, 0.873, 0.903, 0.832 — between 0.832 and 0.919 of them is inside the span — against closed-form removed shares of 0.000, 0.692, 0.590, 0.692. At 14 units a rule with four functions in it takes most of anything it is shown.

The part the rule already took

A diagnostic that reports on what a balancing rule was not handed is run through a column that is 92% inside the span the rule balanced — because orthogonality in the population is not orthogonality on fourteen units.

aimed · Randomisation
Two diagnostics, one answer, two different moments. The two-chain statistic on a covariate probe and on the trial's own difference in arm means, at seven tolerances of a fourteen-unit rule, against the enumerated truth. Both are quiet wherever the set is one set and both fire wherever it is not, at every tolerance — which is what says the outcome probe is the same test rather than a resemblance of it. The difference between them is not accuracy and it is not power. It is when: the covariate probe can be run before a single outcome exists, when a practitioner can still loosen the rule or change the sampler, and it can be run again on a different function if it comes back quiet. The outcome probe runs after the trial, on the one column the trial produced, and what it can do with a positive verdict is repair the p-value rather than the design.

The statistic the p-value is about

The test for whether a balanced-assignment walk reaches its whole set is run on a covariate function chosen before the trial. Run on the difference in arm means it is the same test, and it is about the number the trial publishes.

after · Randomisation
What the rule blocks is not where it splits. How much of the separating direction each probe carries, over 100 designs of 14 units whose admissible set is enumerated and split into two pieces. The deferral this field answers proposed the constraint's active set — which exchanges the tolerance box actually blocks — as a better probe than the design's own leverage, on the ground that leverage is a heuristic and the active set is the quantity. Modelled from the design and the tolerance, it reads 0.4272 against leverage's 0.5395, at 4.43 paired standard errors the wrong way. Counted exactly over the enumerated set — at a cost no trial can pay — it reads 0.3854, worse again. Both beat a random direction at 0.2622, so they are probes; neither beats the two the earlier field already had.

What the rule blocks

A balancing rule breaks the admissible set into pieces by refusing exchanges. Which exchanges it refuses is computable from the design and the tolerance alone, before any assignment exists — and it makes a probe.

blocked · Randomisation
The coverage is exact and it is not the nominal rate. ⌈(m+1)(1−α)⌉/(m+1) against m, the number of calibration points, at α = 0.05. It is a closed form and needs no data. It never falls below 95.0% and never reaches 1−α+1/(m+1), the two bounds the rank argument gives. It equals 95.0% exactly at 10 of the 182 sizes drawn — the sizes where (m+1)α is a whole number, which are 20 apart — and sits above it everywhere else, worst at 38 points where it is 97.4359%, or 2.4359% of coverage nobody asked for. Below 19 points there is no such order statistic and the interval is the whole line, which is where the curve starts.

Coverage from exchangeability alone

A conformal interval's coverage is a fact about the ranks of m+1 numbers, so it can be enumerated before any data arrive — all 40,320 orderings of eight values, agreeing with the closed form to machine precision. What that exactness delivers is not 95%.

conformal · Exchangeability
A model of the active set, and the active set. Each of the 14 units of one design, at the share of its exchanges the tolerance box blocks — computed from the design's columns and the tolerance under a uniform position in the box, against counted over all 116 admissible assignments. The diagonal is where the two would agree. Over 192 designs they agree about the ordering of the units at a correlation of 0.8141 ± 0.0112, negative on 0.5% of them, and disagree about the level: 0.8442 counted against 0.8170 modelled, a gap of 0.0272 ± 0.0051. An admissible assignment does not sit uniformly in its box, and this is the size of that.

A model and a count

The share of a unit's exchanges a tolerance box refuses can be modelled from the design or counted over the admissible set. They order the units the same way at a correlation of 0.81 and disagree about the level by 0.027.

blocked · Randomisation
How far apart the two components are, on each probe. The median separation between the two components of the admissible set — the difference in their mean probe values, over the spread inside a component — over the 100 of 200 designs whose set is enumerated and found split. The separating direction carries 10.565 and needs the enumeration. The fourth power as the earlier fields use it carries 1.543; projected off the span the rule balances, 5.080. The design's own leverage, which uses no dictionary and no outcome, carries 3.836. A random direction in the same subspace carries 0.942, and a direction chosen by looking for concentrated structure carries 0.543 — below random, and the one heuristic here that is worse than not choosing at all.

A probe chosen from the design

The design's own leverage aligns with the separating direction four times better than a random direction in the same subspace. The concentrated direction the argument invites is worse than random.

aimed · Randomisation
The probe a trial has is the probe a trial got. What the two-chain test says when it is run on the trial's own difference in arm means, over 24 outcomes on one fourteen-unit set. The set is in 2 mirror components — that is enumerated, not inferred — so every quiet reading is a miss. 29% of them are quiet. The reason is in the enumerated set rather than in the run: how far the two components are apart on a given probe ranges from 0.001 to 4.938 of a within-component spread across these outcomes, a factor of several thousand. Both covariate probes — chosen before any outcome existed, and replaceable if they had been quiet — report the split. An outcome cannot be chosen and cannot be replaced.

A probe nobody chose

On a set that is definitively in two pieces, seven of twenty-four outcomes report nothing at all. Every covariate probe reports it. What separates them is not accuracy — it is that one of them can be chosen and the other is what happened.

after · Randomisation
Three readings, one verdict. Every tolerance of a fourteen-unit trial, with three comparisons on each. The first is between a chain started at an assignment and a chain started at its complement, which is what a mirror split separates. The second is between two chains started at the same assignment on different streams, which nothing about the set can separate — so a large reading there says the run is too short and not that the set is in pieces. The third is the same comparison on the statistic's absolute value, which is symmetric under the complement and therefore blind to the split by construction. The verdict is the pattern rather than any one line: the split is called only where the first fires and the other two do not, which happens at exactly the tolerances the enumeration calls disconnected — 0.8, 0.75, 0.7.

The statistic that changes sign

A test for an unreachable half needs a quantity that tells one half from the other. Every symmetric reading of a mirror pair is identical, and a magnitude is the natural thing to reach for.

reach · Randomisation
Thinness stays put and reachability does not. The same balancing rule and the same tolerance at five trial sizes. The admitted share barely moves — one admissible assignment in 27, 30, 30, 20, 18 — because the acceptance rate of a rerandomisation is a fact about the basis rather than about the number of units. What does move is the number of single swaps available: 36, 49, 64, 81, 100, growing like a quarter of the square of the trial size. Filled marks are sizes whose admissible set falls into more than one piece. The set is in 4 pieces at 12 units, 2 at 14, and one piece from 16 upwards. The fourteen-unit result is a statement about fourteen units.

A defect that is about size

The admitted share of a rerandomisation barely moves with the number of units. The number of admissible neighbours grows like the square of it, and that is what decides whether the walk can go everywhere.

reach · Randomisation
What each probe can see. How far apart the two components of the admissible set are on each probe, over the spread inside a component, on 100 designs whose set is enumerated and split. It is the population quantity a chain is trying to report. The separating direction itself reads 10.5646; the projected fourth power 5.0800, the design's own leverage 3.8362, the modelled active set 1.9529, the counted active set 2.0170 and a random direction in the same subspace 0.9422. The two active-set probes beat the random direction and lose to both of the earlier field's, which is the field's answer to the question that opened it.

A quantity that loses to a heuristic

Leverage is a heuristic about which units a balancing rule has most to say about. The constraint's active set is the thing the rule actually does. As a probe, the heuristic wins by 4.4 paired standard errors.

blocked · Randomisation
The one thing a trial always reports is the one thing that survives. How wrong three p-values are when they are computed over the half of the admissible set a single walk can reach, rather than over all of it, at a fourteen-unit trial where the whole set can be enumerated. The two-sided p-value on the difference in arm means — the number a trial publishes — is wrong by exactly nothing, at every row, to machine precision. That is not luck: the two components are complement pairs and the difference in arm means is exactly negated by the complement, so the distribution of its absolute value is the same on both. A one-sided p-value on the same statistic is out by as much as 0.112, and the largest response observed in the treated arm — a safety reading rather than an effect, and the one statistic here that is not odd under the complement — by as much as 0.172. The defect survived because the commonest thing anybody computes is the one quantity it cannot touch.

Half a reference distribution

A walk that reaches half its admissible set reports the two-sided p-value exactly right, to the last digit, for ever. A one-sided one it puts on the wrong side of five per cent about once in thirty.

after · Reference
How often each probe finds a split that is there. The share of 34 designs — every one of them enumerated to be in two components — on which a two-chain test of 800 draws declares the split, by probe. The fourth power as the earlier fields use it finds it on 55.9%, so it misses 44.1% of the sets that have one. The same column projected off the rule's span finds it on 88.2%, and the separating direction itself on 91.2%. The design's own leverage, chosen without any dictionary, gets 79.4%. A random direction in the same subspace gets 44.1%, and the direction chosen for being concentrated gets 38.2% — worse than random, which is what a heuristic that finds the wrong structure looks like from the outside.

What a chosen probe finds

On a chain of eight hundred draws the probe the earlier fields use misses 44% of the sets that are split. Its own residual off the rule's span misses 12%, for one least-squares fit.

aimed · Randomisation
Where the set stops being one set. How many of the 3,432 equal splits of fourteen units a balancing rule admits, as the tolerance tightens, with the number of components single swaps leave it in. The set falls from 886 to 84 assignments, and somewhere in that fall it stops being connected: at 0.8 it is in 2 pieces and every assignment's complement is in the other one. Nothing about the rule changes at that point and nothing a chain reports changes either, which is the whole difficulty — the acceptance rate, the stationary distribution and the detailed balance are all in order on both sides of it.

Before the trial and after

The same diagnostic run at two moments answers two different questions. Before, a positive verdict changes the design. After, it changes which number gets reported — and only for the numbers the defect can reach.

after · Assignment
One of them is mostly leverage. How much of the design's own leverage direction each active-set probe carries, once both are standardised and projected off the rule's span — which is what a probe is, so it is the comparison that matters. Over 189 designs the modelled active set agrees with leverage at |r| = 0.8359 ± 0.0114 and the counted one at 0.4239 ± 0.0216. So the modelled probe is largely leverage under another name and the counted one is genuinely a different direction — and the counted one is the worse probe, at 0.3854 of alignment against 0.4272. What the active set contains beyond leverage points away from where the set splits.

Counting it exactly does not help

If a modelled active set lost because the model was crude, the exact one would win. It is computed at a cost no trial can pay, and it is worse — so the approximation was never what was costing the probe.

blocked · Randomisation
A thin enough set is not one set. Every admissible set of 14 units this table can enumerate, by how much of the assignment space it admits and how many pieces it falls into under single swaps. A walk is uniform on the piece it starts in and never leaves it. The pieces are not fragments: at 522 admissible assignments the set splits into 3 halves of exactly 520 each, and every assignment's complement is in the other half — no sequence of admissible single swaps takes an assignment to its own mirror image. Two-swap proposals reconnect four of the six disconnected sets here — the two they do not are the thinnest, where a two-unit move rarely lands anywhere admissible either — which makes a bigger proposal a correctness repair rather than the speed dial it was measured as.

The walk that cannot cross

A thin enough admissible set is not one set. It splits into an assignment and its mirror image, no sequence of admissible single swaps joins them, and the walk that samples it is uniform on half the reference distribution for ever.

dict · Randomisation
Where the constraints exhaust the randomisation. At 16 units there are 12,870 equal splits, so the ones meeting a stated tolerance can be counted rather than estimated. With each of the first k standardised imbalances required to be within 0.4 of a coin's own spread, the admissible count runs 3874 → 1006 → 314 → 0 → 0 → 0 — and at 4 functions there is no admissible assignment at all. The count is the number of distinct answers a randomisation test can give: at 3 functions its finest attainable p-value is 1 in 314. Balance improves with every constraint and the reference distribution shrinks with it, and the two run out at different rates.

When the constraints run out

Every function added to a basis is a constraint the assignment has to satisfy with the same units. At sixteen units and a stated tolerance the admissible assignments run 3,874, then 1,006, then 314, then none — and the count is exact, because the assignment space is finite.

basis · Allocation
The pairing recovers most of it and passes nothing. How much of the separating direction each probe carries, over 100 designs of 14 units whose admissible set is enumerated and split in two. The four rows the pairing adds are the dominant direction of what each blocking matrix keeps past its degrees, and the cut that direction's signs induce. Counted, they read 0.5266 and 0.5258 against the counted per-unit share's 0.3854 — most of the gap between that share and the design's own leverage at 0.5395, closed. Modelled, they read 0.4274 and 0.4954 against 0.4272. Nothing built from the active set passes leverage, and the projected fourth power is still ahead of all of them at 0.6583.

A set of pairs, not a vector

The active set is a graph on the units, and every probe built from it so far has been its degree. Read as a graph it recovers 0.1326 of the alignment the summary lost — and draws level with leverage rather than passing it.

blocked · Randomisation
The squared estimate 1 standard errors from the flat point, exact and linearised. At δ = √n·μ/σ = 1 the exact law of the squared estimate has mean 2.00, variance 6.00 and skewness 2.177; the delta method's normal has mean 1.00, variance 4.00, no skewness, and 30.85% of its mass below zero, where a square cannot go. The Kolmogorov distance between them is 0.3085.

Where the derivative is zero

The delta method reads a standard error off a tangent line, and at a flat point the tangent says the spread is zero. The interval built on it for a squared mean covers 99.991% there and 85.978% one and a half standard errors away, with nearly every miss on the same side — and the law it should have used is a χ², not a normal.

normal · Clt
Three ways to reject with two studies, drawn where the two z statistics live. Two one-sided studies, each summarised by its z statistic. Fisher's combination rejects outside a curve that runs parallel to both axes, so one study past z = 2.378 decides it alone; Stouffer's rejects above the straight line z₁ + z₂ = 2.326; Tippett's rejects when either z passes 1.955. Each region holds exactly 5% of the standard bivariate normal — Fisher's in closed form, e^(−c/2)(1 + c/2) at c = 9.488 — and of 100,000 counted null pairs they catch 4.95%, 4.88% and 5.04%. Two alternatives carry the same Stouffer evidence: one study at 2.326 and the other at nothing, where the powers are 62.7%, 50.0% and 65.4%; and both at 1.163, where they are 47.7%, 50.0% and 38.3%.

Two ways to combine p-values

Fisher's and Stouffer's combinations are both exactly right when every null is true, for the single reason that each p-value is flat. Under a real effect they disagree about which evidence counts: with Stouffer held at 50% power across ten studies, Fisher is the more powerful while the signal sits in six or fewer of them and the less powerful from seven.

testing · Uniformity
Three promises, and no procedure keeps all three. Average coverage and worst-case coverage for four 95% intervals for a proportion at n = 40, computed exactly. Their expected widths are 0.2418, 0.2417, 0.2472, 0.2641 in the same order. The textbook interval and the score interval have the same expected width to four digits — 0.2418 and 0.2417 — and worst-case coverages of 55.31% and 92.21%. The exact interval never breaks its promise and is 9.3% wider than the score interval to do it. Each of the three columns orders the four procedures differently.

An interval that covers and says nothing

A procedure returning the whole line 95% of the time and the empty set otherwise has coverage exactly 95% at every parameter value. Two real intervals at forty observations have expected widths of 0.2418 and 0.2417 and worst-case coverages of 55.31% and 92.21%.

intervals · Coverage
Two companions on one simulation, two hundredfold apart. How many times as many draws each companion is worth, on the same 4,000 simulated samples of 40 observations. The coverage of the interval is estimated with the observed count as its companion, whose expectation is 12 exactly; they correlate at 0.2665 and the companion is worth 1.08 times the draws. The expected width is estimated with p̂(1 − p̂) as its companion, whose expectation is 0.20475 exactly; they correlate at 0.9977 because the width is a monotone function of it, and the companion is worth 214 times the draws — 856 thousand simulated samples' worth of precision from four thousand.

The check worth more than the check

The same exactly known companion that verifies a simulation can sharpen it. On one set of four thousand draws, one companion is worth 1.08 times the draws and another is worth 214 times them, and the factor is 1 − ρ² with nothing else in it.

method · Routes

Named alongside it

The objects these essays reach for when they reach for this one.

Covariate balanceRandomisation testAssignment mechanismConnected componentMarkov chain Monte CarloImbalanceReference distributionRerandomisationExperimental designLeverageProjectionClosed form

All concepts