When the constraints run out
Worth reading first: Not half and half · Randomisation is not balance.
An earlier essay in a neighbouring field ends by naming the thing it has not measured: a basis rich enough to cover every shape the outcome might have is a basis with as many functions as units, and a rule that balances n quantities with n units has no freedom left and is not a randomisation at all. It measures the endpoint — a deterministic rule, whose reference distribution is a point mass — and nothing in between.
The curve between three functions and n is available, and it is exact. At sixteen units the space of equal splits has 12,870 points in it. They can be counted.
The construction, and why it is a counting problem
The experimenter states a tolerance: each of the k standardised imbalances must be no larger than a fraction a of what a coin’s own spread of that imbalance is. Any assignment meeting all k conditions is admissible; the design draws uniformly from the admissible ones.
That is rerandomisation, and its whole content is which assignments survive. With sixteen units and an equal split there are C(16, 8) = 12,870 candidates and every one can be tested, so the admissible count is a number rather than an acceptance rate estimated from draws — which matters, because the interesting part of this curve is where the count is small and an acceptance rate would be estimated from almost nothing.
At a tolerance of four tenths, the counts are
| functions balanced | admissible assignments | bits of randomisation |
|---|---|---|
| 1 | 3,874 | 11.92 |
| 2 | 1,006 | 9.97 |
| 3 | 314 | 8.29 |
| 4 | 0 | — |
At four functions there is no admissible assignment at all. Not few: none. The experimenter’s own stated requirement is unsatisfiable, and a rerandomisation procedure written the obvious way would loop for ever.
What the count is, in the units that matter
The number of admissible assignments is not an abstraction. It is the number of distinct answers a randomisation test can give.
A randomisation test compares the observed statistic with its value under every admissible reassignment. If there are 314 of them, the reference distribution has 314 points and the finest attainable p-value is 1 in 314 — 3.18 × 10⁻³. At two functions it is 9.94 × 10⁻⁴ and at one 2.58 × 10⁻⁴.
Those floors are comfortable. The point is the rate at which they rise: each function added costs about two bits, so the reference distribution loses three quarters of its resolution every two functions, and the fall is faster than the fall in the number of units would be. Three functions on sixteen units leaves as much randomisation as a coin would leave on about nine.
And the floor is the optimistic reading, because a p-value floor of 1/314 assumes the test can use all 314. A one-sided test at 5% needs the top 5% of the reference distribution to be a set the observed statistic can be compared with, which is fifteen assignments — a discreteness that shows up as a conservative test long before the floor becomes binding.
Each constraint costs about 1.8 bits, until one costs everything
The counts and the bits are the same column read two ways, and the differences between the rows are where the shape of the collapse is.
The unconstrained space is bits. So the first constraint costs 1.73 bits, the second 1.95, and the third 1.68.
As conditional admission rates that is 30.1%, 26.0% and 31.2% — three constraints in a row each letting through very close to three assignments in ten, with no sign of the rate falling.
Then the fourth admits 0%.
The trend does not merely fail; it fails by forty-nine orders of magnitude
That collapse is worth pricing against what the three preceding rows predict, because the gap is the whole content of the essay.
Extrapolating the conditional rate, the fourth constraint should have left about 94 admissible assignments out of 314. It leaves none.
And if the constraints were independent — if each of the 314 survivors had an independent 30% chance of meeting the fourth condition — the probability of no survivor would be , which is about .
So the zero is not a small count arrived at by bad luck, and it is not the tail of the trend the first three rows establish. It is a statement that the four conditions are geometrically incompatible at this tolerance: the region they jointly define is empty, and the three-constraint region’s 314 points all sit outside it.
Nothing in the first three rows carries a warning about that, which is the reason the curve is worth enumerating rather than extrapolating. A designer reading 3,874 → 1,006 → 314 has a well-behaved geometric sequence and every reason to expect ninety or so at the next step.
What eight bits of randomisation is, as a reference distribution
The last non-empty row deserves one conversion, because “8.29 bits” is not a quantity anybody acts on.
Three hundred and fourteen admissible assignments is the entire reference distribution a randomisation test on this design has to work with. The smallest p-value it can produce is 1/314 = 0.0032, and a test at five per cent is decided by whether the observed statistic beats the sixteenth largest of three hundred and fourteen numbers.
That is a usable test and it is not a comfortable one: the 5% point is estimated from fifteen assignments, so the granularity of the verdict is a third of a per cent and its position depends on a handful of points at the edge of an enumerated set. One more balancing function and there is no test at all.
The tolerance is the other dial, and it is not gentle
The tolerance is the experimenter’s other choice, and moving it moves the whole curve rather than shifting it.
At a quarter of a coin’s spread the counts are 2,462, 380, 34 and then none. Thirty-four admissible assignments is a reference distribution with a floor at one in thirty-four, which is not a test.
At six tenths the counts are 5,660, 2,324, 1,294, 1,294, 928 and 264 — six functions all feasible, with 264 assignments left at the end. At eight tenths, 7,272 down to 618.
So the feasibility boundary is a curve in two variables, and neither of them is the number of units. An experimenter who wants three functions balanced to a quarter of a coin’s spread on sixteen units is asking for something that exists — thirty-four ways — and one who wants four is asking for something that does not.
More units, and how much they buy
Twenty units instead of sixteen puts C(20, 10) = 184,756 assignments in the space, and at a tolerance of four tenths every basis size up to six is feasible: 56,080, 17,346, 10,268, 9,622, 7,390 and 1,408.
The cost in bits across those six is 15.78 down to 10.46, or a little under one bit per function — less than the two bits per function at sixteen units, because the space is larger relative to the constraints.
That comparison is the practical one. Four extra units — a 25% larger trial — turn an infeasible design into a feasible one and halve the cost per constraint. The binding resource is the number of units, not the number of functions, and the exchange rate between them is steep at the small end.
The extra units do not abolish the collapse, though; they move it. Hold the twenty units and tighten the tolerance from four tenths to a quarter, and the sequence still runs all the way to six functions — 35,596, 7,196, 4,036, 3,942, 3,420, 602 — but the last step is the familiar cliff. The three steps before it divide the admissible set by 1.8, then by 1.02, then by 1.15; the sixth function divides it by 5.7. It is the same shape the sixteen-unit run showed at four functions: a sequence that looks like it is levelling off right up to the point where it is not.
Which is why the count is worth having exactly rather than approximately. Six hundred and two assignments is still a usable reference distribution — a one-sided test at 5% has thirty of them to work with — but nothing in the shape of the first five steps predicts it, and a design chosen by extrapolating that shape would have expected two thousand. The assignment space is finite and the admissible set can be enumerated, so there is no reason to extrapolate at all.
One detail in the twenty-unit sequence is worth pointing at because it is not noise. The counts at three and four functions are 10,268 and 9,622 — almost the same. The fourth function in the dictionary is a cut at the median, and by the time the covariate, its square and its cube are balanced a median split is very nearly balanced too. A constraint that is nearly implied by the ones already imposed costs nearly nothing, which is the same correlation between basis functions that makes the insurance cheap in the field next door.
The sixth function is a cut at 2, and it costs a great deal — 7,390 down to 1,408 — because at twenty units only two or three fall above 2 and balancing an indicator on two units is nearly a constraint on which arm a single unit goes to.
The two curves, and where they cross
Balance improves with every constraint and randomisation falls with every constraint, and the interesting question is whether they run out together.
They do not. Measured on the assignments that survive at sixteen units and a tolerance of four tenths, the imbalance achieved in the functions the rule was told about is essentially flat: 0.0523 of a coin’s variance at one function, 0.0543 at two, 0.0633 at three. Balance is not improving as the basis grows — it was already near the tolerance’s own limit at one function, and the tolerance is what sets it.
What is improving is coverage. The imbalance left in the functions the rule was not told about falls from 0.6942 to 0.5913 to 0.5750 across the same three steps.
So one curve is nearly flat, one falls slowly, and the third — the randomisation — falls fast and then to zero. The trade is not balance against randomisation; it is which functions are covered against randomisation, and the covered ones are already as balanced as they are going to get.
That reframes the choice. Adding a function does not make the design better at what it was doing; it makes it protective against a shape it was not protecting against, at a fixed price in bits. Whether that is worth it is the maximin question, and this essay supplies the price.
Why an acceptance rate is the wrong quantity to report
Rerandomisation is usually described by its acceptance probability — accept one draw in a hundred, say — and that description is sufficient in the regime where the procedure is normally used and useless in the regime this essay is about.
The two are related by one multiplication: the admissible count is the acceptance probability times the size of the space. At sixteen units and a tolerance of four tenths, three functions accept 2.44% of the 12,870 splits, which sounds like a healthy design and is 314 assignments. Four functions accept 0%, which no simulation would ever establish: a program drawing a million assignments and rejecting all of them has evidence that the rate is below about three in a million, which is not the same statement as there is none.
The distinction matters because the two quantities behave differently as the trial grows. Hold the acceptance probability fixed and grow the trial, and the admissible count grows exponentially — everything is fine. Hold the tolerance fixed and grow the number of functions, and the acceptance probability falls geometrically while the space stays the same size, so the count falls off a cliff. An experimenter thinking in acceptance rates sees a small number get smaller; an experimenter thinking in counts sees a reference distribution disappear.
Report the count. It is the size of the reference distribution, it is what the p-value floor is one over, and it is a number rather than a rate — so it says whether the design exists.
What a deterministic design cannot do
The far end of the curve is worth naming precisely, because it is the reason any of this matters rather than being an interesting fact about counting.
A design with one admissible assignment has a reference distribution that is a point mass. A randomisation test on it compares the observed statistic with itself and returns a p-value of one, always. It cannot reject at a true null, which is fine, and it cannot reject at a false one, which is not: the neighbouring field measures a fully deterministic balancing rule finding a real effect 0.0% of the time.
Everything between that endpoint and a coin is a matter of degree, and the degree is the count. Fifteen assignments in the top tail is a test that can produce a p-value of 0.05 and nothing smaller. Three hundred and fourteen assignments is a test with a floor at 0.003. Twelve thousand is a coin.
There is no threshold in that sequence and no gate could impose one. What there is, is a number the experimenter can compute before the trial and almost never does.
Three quantities an experimenter can compute in advance
Everything in this essay is available before a single unit is treated, from the covariates alone, and none of it needs a simulation. It is worth listing what to compute, because the whole point of an exactly countable design space is that the answers are answers rather than estimates.
The admissible count. Enumerate if the trial is small enough; sample and multiply if it is not, remembering that a sampled count is unreliable exactly where it matters. If it is small, the design does not exist in the form it was written.
The p-value floor. One over the count, and then a second look at the top tail: a test at 5% needs one in twenty of the reference distribution above the observed statistic, so a reference distribution of forty is a test with two useful outcomes.
And the imbalance in the functions nobody constrained. The design’s coverage, measured on the admissible set: 0.6942, 0.5913 and 0.5750 of a coin’s for the unconstrained functions at one, two and three constraints here. This is the number that says what the constraints are buying, and it is the one that is nearly always left uncomputed because it requires naming functions the design is not balancing.
The third of those is the one worth insisting on. A design report that lists the balanced covariates and their achieved imbalances is reporting the cells the design was optimising — the same flattering reading a neighbouring field refuses, where every rule removes nearly all of the imbalance in the covariate it reads, including the rules whose worst case is no better than a coin’s.
What this says about the analysis
One consequence reaches past the design, and it is the reason a rerandomised trial needs its own reference distribution rather than a t table.
An assignment drawn uniformly from 314 admissible ones is not an assignment drawn uniformly from 12,870. The treatment estimate’s sampling distribution is over those 314, its variance is smaller than the coin’s — that is the whole point of the design — and a standard error computed as though the assignment were a coin toss is too large. That is the same conservatism the balancing fields measure from the other direction, and here it has an exact source: the design’s own reference set, whose size this essay counts.
So the analysis has to re-randomise over the admissible set, and it can only do that if the admissible set is written down. Which means the tolerance, the basis and the unit covariates all have to be recorded before the trial — not as good practice but because the analysis is arithmetic over a set that cannot otherwise be reconstructed.
What is claimed here, and what is not
This essay takes how much randomisation a basis leaves, and the claims are the exact admissible counts at two trial sizes and four tolerances, the p-value floors they imply, and the separation between the balance curve and the coverage curve.
What stays out and is named as a decision: a Mahalanobis criterion instead of a coordinate-wise one, which is the standard form of rerandomisation and whose acceptance probability is chosen rather than implied — the coordinate-wise version is used here precisely because its acceptance rate is a consequence of a stated tolerance rather than a parameter; trials large enough that the space cannot be enumerated, where the counts become estimates and the small-count end is exactly where estimation fails; and unequal arms, which change the space from C(n, n/2) to a sum over sizes.
The boundary against the shape field is what is being counted. That reading three functions costs about two points of variance, and that the constraints compete, is established there by simulation. What is new is that the competition has an exact currency and the currency runs out.
The checks, and the refusals that make them mean something
One claim is gated in this field’s library, in two parts. The admissible count is required to fall with every function added — monotone, on every tolerance — and at sixteen units with a tolerance of four tenths the fourth function is required to leave exactly zero admissible assignments, which is the endpoint this essay exists to locate and which would be invisible to any procedure that estimated an acceptance rate. The second part is that twenty units keep all six basis sizes feasible while still costing more than four bits across them, which is what says the exhaustion is about the ratio rather than about either number alone.
The refusal that bears on this essay is the deterministic rule from the neighbouring field, whose reference distribution is a point mass and which finds a real effect 0.0% of the time. It is the last row of the table above, reached by a different route and measured on outcomes rather than on counts.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A probe nobody chose — both name assignment mechanism, covariate balance, exact enumeration, imbalance, randomisation test, reference distribution, rerandomisation, statistical power
- Before the trial and after — both name assignment mechanism, covariate balance, exact enumeration, imbalance, p-value, randomisation test, reference distribution, rerandomisation
- A defect that is about size — both name assignment mechanism, covariate balance, exact enumeration, imbalance, randomisation test, reference distribution, rerandomisation
- A model and a count — both name assignment mechanism, covariate balance, exact enumeration, experimental design, imbalance, randomisation test, rerandomisation
- A probe chosen from the design — both name assignment mechanism, basis functions, covariate balance, exact enumeration, experimental design, randomisation test, reference distribution
- A quantity that loses to a heuristic — both name assignment mechanism, covariate balance, exact enumeration, experimental design, imbalance, randomisation test, rerandomisation
Named objects
A flat tag is an object no other essay names yet.
Allocation ruleAssignment mechanismBasis functionsCovariate balanceDesign criterionEntropyExact enumerationExperimental designImbalancep-valueRandomisation testReference distributionRerandomisationStatistical power