A probe from what the rule blocks

A model and a count

The share of a unit's exchanges a tolerance box refuses can be modelled from the design or counted over the admissible set. They order the units the same way at a correlation of 0.81 and disagree about the level by 0.027.

Worth reading first: Randomisation is not balance · Balancing what is known in advance.

The probe the previous essay builds is a model. It assumes an admissible assignment sits uniformly inside its tolerance box, which is how a displacement of a known size turns into a probability of being refused.

At fourteen units the assumption can be checked against the thing it approximates, because the admissible set can be enumerated and every exchange out of every member of it can be tried. This essay is that comparison, and it is worth making for a reason beyond diligence: the exact version is also a probe, and it is a probe no trial can build.

The two quantities

Modelled. For each unit, one minus the average over the other units of Π_k max(0, 1 − |Δz_k|/2a), where Δz is the displacement that unit’s exchange produces. Reads the design’s columns and the tolerance. Costs n² products.

Counted. For each unit, the share of the single exchanges it takes part in, over every admissible assignment, that leave the set. Reads the enumerated admissible set. Costs the enumeration, which is 3,432 splits at fourteen units and is not available at forty.

They are the same quantity in the same sense that a probability and a frequency are, and they are not the same number.

One more difference is worth naming because it is not an approximation at all. The modelled share averages over every other unit, since it is built before any assignment exists and there is no other arm to average over. The counted share averages over the exchanges that actually occur, which are exchanges between the two arms of admissible assignments. So the two are averaging over different sets of pairs, and even an exact model of the blocking probability would not reproduce the count.

A model of the active set, and the active set. Each of the 14 units of one design, at the share of its exchanges the tolerance box blocks — computed from the design's columns and the tolerance under a uniform position in the box, against counted over all 116 admissible assignments. The diagonal is where the two would agree. Over 192 designs they agree about the ordering of the units at a correlation of 0.8141 ± 0.0112, negative on 0.5% of them, and disagree about the level: 0.8442 counted against 0.8170 modelled, a gap of 0.0272 ± 0.0051. An admissible assignment does not sit uniformly in its box, and this is the size of that.
Fig. 1 The fourteen units of one design, at the two shares. The diagonal is where the two would agree. The slider changes how many designs the summary is read over.

Why the enumeration is affordable here and nowhere else

The whole comparison rests on being able to compute the exact quantity, and that is possible only in the setting this line of fields deliberately works in.

Fourteen units split evenly is 3,432 assignments. Checking each against a box on four columns is four sums and four comparisons, so the admissible set is a few milliseconds’ work; at the design the figures are drawn from it holds 116 members. Every single exchange out of each of those is 49 pairs, so the tally walks 5,684 exchanges and finds 4,748 of them refused.

That is the entire population, not a sample of it. There is no standard error on the counted share for one design, because nothing was estimated — which is why the summary’s standard error, 0.0051 on the level gap, is a standard error across designs rather than within them.

The field that established this setting is explicit about why fourteen: it is large enough for the admissible set to fall into pieces and small enough for the pieces to be counted, and both halves are needed. A field that could only sample the set would be comparing a model against an estimate, and the interesting part of this comparison is that one side of it is exact.

What agrees

The ordering does, which is what a probe needs.

A probe is a direction, so what matters is not whether the model gets the level right but whether it ranks the units the way the count does — a column and the same column shifted or scaled are the same probe. Over 192 designs the per-unit correlation between the two is 0.8141 ± 0.0112, with a median of 0.8564, and it is negative on 0.5% of them.

That is a good model of an ordering. It is not a perfect one: a correlation of 0.81 across fourteen points means the model routinely swaps neighbouring units, and on one design in two hundred it gets the ordering backwards.

What does not

The level does not, and the direction of the disagreement is the interesting half.

Averaged over designs, the counted share is 0.8442 and the modelled share is 0.8170 — a gap of 0.0272 ± 0.0051, more than five standard errors. The model under-states how much the box refuses.

That is the opposite of what the assumption predicts. An imbalance is a signed sum of n standardised terms, so it is nearly Gaussian; conditioning on it lying inside the box concentrates it towards the middle rather than spreading it uniformly; and a point nearer the middle survives more displacements than a uniform one. So a uniform-position model should be too pessimistic about survival — should over-state blocking — and it is too optimistic.

What 0.81 means across fourteen points

A correlation reported over designs needs its own reading, because the number is an average of 192 correlations each computed on fourteen points.

Fourteen points is few. The standard error of a single correlation of 0.8 on fourteen observations is about 0.1, so an individual design’s reading is not precise — which is why the summary is a mean over designs with a standard error of 0.0112 rather than a single number, and why the median, 0.8564, is reported beside it.

The mean sitting below the median says the distribution is left-skewed: most designs agree better than 0.81 and a tail of them agree much worse. That tail is where the 0.5% of negative correlations live, and it is worth knowing that they exist rather than reporting the median alone.

A design on which the model gets the ordering backwards is a design on which the probe built from it points away from the one built from the count. Those designs are rare, they are not identified by anything computable in advance, and they are part of what a probe built on a model costs.

What a correlation of 0.8141 leaves out

Squared, 0.6628: the model accounts for two thirds of the variation in the counted shares across the fourteen units and misses a third.

That third is not the level error. A model that were the count minus a constant would correlate at exactly one however wrong its constant was, so the 0.0272 offset and the 0.8141 correlation are independent failures — one in where the shares sit and one in how they are ordered relative to each other.

The offset is small as a share of what it is an offset in. The enumerations this field reports put the average counted share near 0.835 — 4,748 of 5,684 exchanges refused at the design the figures are drawn from — so 0.0272 is about three per cent of the level. The model is low, and it is low by a few per cent rather than by a factor.

So the two defects are of very different sizes: a three per cent error in the level and a third of the shape. And it is the shape error, not the level error, that the field goes on to find is doing the probe some good.

What the exact version costs, and why nobody will ever run it

The two columns’ costs are worth writing out at two sizes, because the ratio between them is not the kind of thing a constant factor covers.

At fourteen units the enumeration is 3,432 splits, of which 116 are admissible, and 49 exchanges out of each admissible one — under two hundred thousand operations counting the admissibility checks. The model is n2=196n^2 = 196 products. A factor of about a thousand.

At forty units the enumeration is (4020)=1.38×1011\binom{40}{20} = 1.38\times10^{11} splits and 400 exchanges out of each admissible one, which is of order 101310^{13} operations. The model is 1,600 products. A factor of about 10¹⁰.

The gap between those two ratios is the whole reason the comparison is worth making here rather than anywhere. At fourteen units the exact quantity is a thousand times dearer and available; at forty it is ten billion times dearer and is not. So this is the largest design on which the model can be checked at all — and, since the model’s cost grows like n2n^2 while the count’s grows like (nn/2)\binom{n}{n/2}, no increase in patience moves that boundary by more than a couple of units.

Why the sign runs that way

The reason is that a swap does not move one coordinate; it moves all of them, and it moves them in a correlated way.

The model multiplies four independent survival probabilities, one per basis column. That treats the four displacements as four separate chances to leave the box. In fact the displacement vector for a given exchange is a fixed direction — 2(c_i − c_j) scaled — and an admissible point that is near the boundary in one coordinate is, on this design’s correlated columns, often near it in another. So the coordinates fail together rather than independently, and failures that coincide are counted twice by a product and once by reality.

Counting them twice makes survival look more likely than it is, because a product of independent survivals over-counts the ways of dying only when the deaths are disjoint. Here they overlap, so the product over-states survival and under-states blocking.

Two wrong assumptions in opposite directions, and the independence one wins. That is not a satisfying resolution — it is a story consistent with the sign and the size — and the honest form of it is that the model is low by 0.027 and neither of its two assumptions is separately measured.

A third quantity the two share

Both columns are built from the same design, so it is worth asking how much of either is simply a restatement of something the design already carries.

The design’s own leverage — the diagonal of the hat matrix of the rule’s basis — is a per-unit number too, and it is the probe the earlier field found worked. A unit with high leverage carries a lot of the rule’s span, and a unit whose exchanges are refused is a unit whose columns are large. Those are not the same statement but they are close relatives.

Projected off the rule’s span and standardised, the modelled active set agrees with leverage at |r| = 0.8359 ± 0.0114 and the counted one at 0.4239 ± 0.0216. So the two active-set columns differ from each other by about as much as one of them differs from leverage, and the one that is nearer to leverage is the model.

That is a fact worth carrying into the next two essays, because it means the three columns are not three independent attempts at one thing. They are one direction and two departures from it, and the departures go in different directions and by different amounts.

What a level error costs a probe

Nothing directly, and that is worth being explicit about.

The column is standardised before it is used: centred and scaled to unit variance. A constant offset in every unit’s share disappears at the centring step, and a uniform scaling disappears at the scaling step. So a model that is systematically low by a constant produces exactly the same probe as one that is right.

What a level error does cost is interpretation. The number 0.8170 is not a good estimate of the share of exchanges a box refuses; the number 0.8442 is, and only one of them is available on a real trial. A practitioner who wanted to report how constrained their randomisation is would be reporting the modelled number and would be three points low.

The probe is unaffected. The report is wrong. Those are different failures and only the second is repaired by knowing about it.

One of them is mostly leverage. How much of the design's own leverage direction each active-set probe carries, once both are standardised and projected off the rule's span — which is what a probe is, so it is the comparison that matters. Over 189 designs the modelled active set agrees with leverage at |r| = 0.8359 ± 0.0114 and the counted one at 0.4239 ± 0.0216. So the modelled probe is largely leverage under another name and the counted one is genuinely a different direction — and the counted one is the worse probe, at 0.3854 of alignment against 0.4272. What the active set contains beyond leverage points away from where the set splits.
Fig. 2 How much of the leverage direction each of the two active-set columns carries, which is the field’s last essay.

What the comparison is not

Two things this essay deliberately does not do, and both would change what it means.

It does not fit the model to the count. The uniform-position product has no free parameter, so there is nothing to tune, and that is the point: a model with a constant fitted to make its level match would match its level and would have stopped being computable in advance. The gap of 0.027 is the model’s, not a residual after fitting.

And it does not repair the model. A better one is available in principle — replace the uniform position by a Gaussian one conditioned on the box, or compute the coordinates’ joint survival rather than their product — and either would presumably reduce the gap. Neither is here, because the last essay of this field measures what an exact active set is worth as a probe, and the answer makes a better approximation of it pointless.

That ordering is deliberate. Measure what the exact thing buys before improving an approximation to it, because the approximation can only ever get as good as the thing it approximates.

The units the model gets right

Reading one design closely is worth doing, because the errors are not uniform across units.

At the design the figures are drawn from, three of the fourteen units have a counted share of exactly one: every exchange they take part in is refused, so they never move between arms within the admissible set. The model gives them 0.993, 0.987 and 0.831.

The first two are nearly right and the third is not. All three are pinned by the count; the model sees the third as merely rather constrained.

The remaining eleven have counted shares between 0.704 and 0.961 and modelled shares between 0.739 and 0.869 — so the model compresses the range, which is what an averaging assumption does. The counted shares span 0.296 and the modelled ones span 0.254 across the mobile units, and the model’s floor is above the count’s floor while its ceiling is below.

A model that compresses towards the middle and gets the extremes half right is a reasonable description of what a uniform-position assumption should do, and the unit it gets badly wrong is a unit the count says is pinned.

Which unit that is, is not predictable from the model. The three pinned units on this design have modelled shares of 0.993, 0.987 and 0.831, and the fourth-largest modelled share is 0.869 — belonging to a unit whose counted share is 0.867 and which is not pinned at all. So a practitioner reading the modelled column would identify two of the three pinned units and would promote one that is not pinned above one that is.

That is the concrete form of a correlation of 0.81 across fourteen points. It gets the shape right, it gets most of the ordering right, and it makes a specific and unrecoverable mistake about one unit in a design where three units matter more than the other eleven.

What the rule blocks is not where it splits. How much of the separating direction each probe carries, over 100 designs of 14 units whose admissible set is enumerated and split into two pieces. The deferral this field answers proposed the constraint's active set — which exchanges the tolerance box actually blocks — as a better probe than the design's own leverage, on the ground that leverage is a heuristic and the active set is the quantity. Modelled from the design and the tolerance, it reads 0.4272 against leverage's 0.5395, at 4.43 paired standard errors the wrong way. Counted exactly over the enumerated set — at a cost no trial can pay — it reads 0.3854, worse again. Both beat a random direction at 0.2622, so they are probes; neither beats the two the earlier field already had.
Fig. 3 Where both columns land as probes, in the first essay of this field.

What the disagreement does to the probe, measured

The argument above says a level error cannot affect a standardised column. It is worth checking that against the numbers rather than leaving it as arithmetic.

If the model differed from the count only by a constant offset and a scale, the two columns would be the same direction and their correlation after standardising would be exactly one. It is 0.8141. So the disagreement is not a level error; the level error is the part that does not matter, and what is left over is the part that does.

The two probes therefore land in different places, and they do: the modelled column aligns with the separating direction at 0.4272 and the counted one at 0.3854, which is a real difference in what each can see even though it is not a large one.

A correlation of 0.81 between two columns is two probes, not one probe measured twice. Everything the field concludes about them is a statement about two directions that share four fifths of themselves, and both of those directions lose to a third one.

Why the count is not the answer

The counted share is exact, and it is on the table as a benchmark rather than as a recommendation, for a reason that should be stated before anybody wants it.

It needs the enumeration. At fourteen units there are 3,432 equal splits and the admissible set holds 116 of them; at twenty there are 184,756, which is where the field that enumerates these sets stops; at forty the count is beyond anything. And the diagnostic this probe feeds exists because the set cannot be enumerated — a trial that could enumerate its admissible set would not need a chain and would not need a test of whether the chain reaches everything.

So the counted share is a quantity available exactly where it is not needed, which is the same shape as the separating direction itself. Both are in the table as the thing feasible probes are short of.

And it turns out to be worse. The last essay of this field is that result, and it is the one that decides what the whole field means: if the modelled probe lost because the model is crude, the counted one would win.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

  • A probe nobody chose — both name assignment mechanism, connected component, covariate balance, exact enumeration, imbalance, markov chain monte carlo, randomisation test, rerandomisation
  • A test rather than a survey — both name assignment mechanism, connected component, covariate balance, exact enumeration, imbalance, markov chain monte carlo, randomisation test, rerandomisation
  • Before the trial and after — both name assignment mechanism, connected component, covariate balance, exact enumeration, imbalance, markov chain monte carlo, randomisation test, rerandomisation
  • The part the rule already took — both name assignment mechanism, connected component, covariate balance, exact enumeration, experimental design, leverage, markov chain monte carlo, randomisation test
  • The statistic the p-value is about — both name assignment mechanism, connected component, covariate balance, exact enumeration, imbalance, markov chain monte carlo, randomisation test, rerandomisation
  • What a chosen probe finds — both name assignment mechanism, connected component, covariate balance, exact enumeration, experimental design, leverage, markov chain monte carlo, randomisation test

Named objects

A flat tag is an object no other essay names yet.

Approximation errorAssignment mechanismBasisClosed formConnected componentCorrelationCovariate balanceExact enumerationExperimental designImbalanceLeverageMarkov chain Monte CarloNormal approximationRandomisation testRerandomisation