What a chosen probe finds
Worth reading first: Randomisation is not balance · Balancing what is known in advance.
Everything two essays have measured is a population quantity: how far apart the two components of an admissible set are on a given probe, computed by enumerating the set. A practitioner has no enumeration and no population. What they have is a chain of some length and a verdict.
Those are different things, and the gap between them is the whole difficulty of the diagnostic. The field that built the two-chain test spends most of itself on it: a statistic computed from B draws rather than from B/τ effective ones reports a disconnection on every run of a perfectly ergodic chain, and the integrated autocorrelation time of the signed imbalance runs into the hundreds at two hundred units.
So a separation of 5.080 against 1.543 is a promise rather than a result. This essay is the result.
What a short chain reports
Thirty-four designs of fourteen units, every one enumerated and found to be in two mirror components, and a two-chain test of eight hundred draws run through each of the six probes on each of them.
The fourth power as the earlier fields use it declares the split on 55.9% of them. It misses 44.1% of the sets that have one — not because the chain is too short in general, but because the statistic it is watching carries a median separation of 1.543 and the threshold is four.
The same column projected off the span the rule balances declares it on 88.2%.
The design’s own leverage, projected, gets 79.4%. A random direction in the same subspace gets 44.1%. The concentrated direction — the projection-pursuit heuristic — gets 38.2%, below random on the verdict as it was below random on the separation. And the separating direction itself, which needs the enumeration, gets 91.2%.
A missed split is the worst outcome this diagnostic has. It is a quiet verdict on a walk that is confined to half its support, and a quiet verdict is what a practitioner reads as the chain is fine. Going from 44% missed to 12% missed is the whole of what the projection buys, and it costs one least-squares fit.
Why eight hundred draws
The chain length is a choice and it is the choice that makes the comparison legible, so it is worth defending rather than presenting.
At twenty thousand draws every probe but two finds nearly every split, and the ordering that the separations predict is compressed into the top of the scale. At two hundred draws almost nothing is found and the ordering is compressed into the bottom. Eight hundred is where the six probes are spread out, and it is spread out because the threshold of four sits near the middle of the weakest probes’ distributions and in the tail of the strongest’s.
Choosing the length to make the differences visible is not choosing it to make them large. The separations are a population quantity measured without any chain at all, and the ordering here is theirs. What the chain length decides is only how much of the ordering a verdict can express — and the next section is what happens at lengths where it expresses none of it.
The statistic underneath the verdict
A detection rate compresses everything into a threshold, so it is worth looking at what is being thresholded.
The medians are 6.16 for the raw column, 24.61 for its residual, 15.26 for leverage, 3.44 for a random direction, 2.02 for the concentrated one, and 43.48 for the separating direction.
Two things read off that which the rates hide.
The projected probe is not marginally better, it is four times louder. Its median statistic is 24.61 against a threshold of 4, so it is not near the boundary on a typical design; the raw column at 6.16 is. That is why the rates differ so much for what looks like a modest change in separation — the threshold sits in the middle of the raw column’s distribution and in the tail of the projected one’s.
And the ordering is the separation’s ordering, exactly. 43.48 > 24.61 > 15.26 > 6.16 > 3.44 > 2.02 against 10.565 > 5.080 > 3.836 > 1.543 > 0.942 > 0.543. The chain reports what the population says it should, in the same order, at every chain length in the slider’s range. The population quantity transports to the verdict, which is not automatic and is the thing this essay exists to check.
Thirty-four designs is thirty-four designs
Every rate in the paragraph above is a count out of thirty-four, and converting back makes the comparison legible in a way the percentages hide.
The fourth power finds 19 of the thirty-four splits. The projected column finds 30. The design’s leverage finds 27, a random direction 15, the concentrated direction 13, and the separating direction — which needs the enumeration — finds 31.
So the whole distance between the best feasible probe and the oracle is one design. Not a margin that shrinks with effort; a single set out of thirty-four on which the enumeration’s own direction fires and the projected column does not.
That is the field’s result stated at its natural resolution. A probe costing one least-squares fit reaches within a design of a probe that requires knowing the answer.
What thirty-four designs cannot say
The same arithmetic bounds the other comparisons, and it bounds them harder.
A rate near 85% on thirty-four designs carries a binomial standard error of √(0.85 × 0.15 / 34) = 6.1 points, which is two designs. So the ordering of 79.4%, 88.2% and 91.2% — 27, 30 and 31 designs — sits inside the noise of an unpaired reading, and only the pairing across designs makes it a ranking at all: the same thirty-four sets are put to every probe, so what matters is whether a probe that misses is missing the same sets, not how many it misses.
The wide gaps survive that scrutiny easily. Nineteen against thirty is eleven designs, and thirteen against thirty is seventeen, which is half the sample. Those are the comparisons the field is about, and they are the ones that would survive being read off a much cruder instrument.
The narrow ones do not, and it is worth being explicit about which claims this rests on. That projecting off the balanced span is worth a great deal is measured. That leverage beats a random direction is measured. That the separating direction beats the projected column is one design, and one design is what a reader should treat it as.
Two probes that were never going to work
Two of the six rows are below the random baseline on the verdict as they were on the separation, and both are worth a sentence because both are things somebody would do.
The concentrated direction gets 38.2%, against 44.1% for a direction picked at random from the same subspace. It is the outcome of an optimisation — eight restarts of a fixed-point iteration, kept by its own index — and it is worse than not optimising. The reason is in the previous essay: the index finds a spike on one unit and the separating direction is spread over several.
And the raw fourth power gets 55.9%, which is the row that matters, because it is not a straw man. It is the probe the earlier fields use, chosen by an argument that is correct in the population, and it is missing nearly half the splits it is pointed at. Nothing about the run reports this: the statistic comes back under four, the control comes back quiet, the acceptance rate and the stationary distribution are in order, and the verdict is the walk reaches the whole set.
That is the same shape as the defect the whole line of fields is about. A diagnostic that is quiet for the wrong reason looks exactly like a diagnostic that is quiet for the right one, and the only way to tell them apart is to know how much the probe could have seen.
Where the rates run out
A detection rate is bounded above by one, and at a long chain everything gets there.
Running the same comparison at twenty thousand draws instead of eight hundred, the rates are 88.2% for the raw column, 94.1% for the projected one and 85.3% for a random direction. The chain has become long enough that a weak probe is nearly enough, and the spread that ran from 38% to 91% at eight hundred draws has closed to nine points.
That is not a reason to prefer the long chain. It is a statement about where the choice of probe matters, and it is the useful half of the finding:
A better probe buys chain length. The projected probe at eight hundred draws finds 88.2% of the splits; the raw probe at twenty thousand finds 88.2% of them. The same number, for a twenty-fifth of the computation, from a projection that costs one least-squares fit. At two hundred units, where the integrated autocorrelation time is in the hundreds and a run of twenty thousand draws is a few dozen effective ones, that is the difference between a diagnostic and a ritual.
What a practitioner does not have
Every rate above is conditioned on the set being split, which is knowable here and is not knowable on a trial.
Of two hundred designs drawn from one law at one tolerance, a hundred have a set that splits and a hundred do not. A practitioner has one design and cannot tell which. So the number they need is not the detection rate on split sets — it is what the diagnostic tells them, given that it is quiet.
That calculation needs the false-alarm rate as well, and it is measured against the enumerated truth in the field that built the test rather than here: the test agrees with the enumeration at every tolerance, quiet on connected sets and firing on split ones, for both probes it was run with. Combined with a split prevalence of a half at this tolerance and a false-alarm rate near zero, a quiet verdict through the raw probe still leaves about three chances in ten that the set is split; through the projected probe it leaves about one in nine.
That is the sentence a practitioner needs and it is arithmetic on two rates neither of which they can measure. The prevalence is a fact about the tolerance and the covariate law, and the detection rate is a fact about the probe and the chain. Both are computable in simulation for a design in hand, which is what makes the diagnostic usable at all, and neither comes out of the run itself.
The control, which stays quiet throughout
One reading has to be checked before any of this counts, and it is the reading that separates a better probe from a noisier one.
The two-chain test carries its own control: the same comparison on the statistic’s absolute value, which is symmetric under the complement and therefore blind to the mirror split by construction. If a probe fired more often because it was noisier rather than because it saw more, the blind statistic would rise with it.
It does not. Across every probe in the table the median blind statistic sits below one, against a threshold of four and against firing statistics that run to forty. The projected probe’s blind reading is no larger than the raw column’s.
So the extra firing is signal. That is the check that makes a rise in detection a rise in power rather than a rise in false alarms, and it is available on every run because the control is computed from the same draws.
The two things a chain length cannot fix
It is worth separating the two ways a probe fails, because only one of them is about running longer.
A probe that carries no separation cannot be rescued by any chain. If the two components have the same mean on a column, the two chains are estimating one number and the difference is noise however many draws are taken. That is the seven-of-twenty-four case the field that asked for a chosen probe found: its worst outcomes carry a separation of 0.001 of a within-component spread, and no run length reaches it.
A probe that carries some separation is a matter of arithmetic. The statistic grows like the square root of the effective number of draws, so a probe with half the separation needs four times the chain. The raw column’s 1.543 against the projected column’s 5.080 is a factor of 3.3, which is eleven times the chain — and the measured ratio, 88.2% at eight hundred draws against 81.3% at twenty thousand, says twenty-five times and still short.
The gap between eleven and twenty-five is the autocorrelation. A longer run of this walk does not buy proportionally more effective draws, because consecutive states differ by one swap out of seven, so the arithmetic that says four times the chain for half the separation understates the price at every length. A better probe is cheaper than a longer chain by more than the square-root law suggests, and that margin grows with the trial.
What would have to be true for this to reach a real trial
The measurements are at fourteen units because that is where the truth is computable. A trial is at two hundred or two thousand, and three things change on the way there.
The pinned share falls. At two hundred units the rule has taken about a third of the fourth power rather than 92% of it, so the projection is worth less — a third of a probe rather than nine tenths. It is still worth something and it still costs nothing.
The chain gets much worse. At two hundred units the integrated autocorrelation time of the signed imbalance is in the hundreds rather than the tens, so twenty thousand draws are a few dozen effective ones and the statistic’s standard error is computed from B/τ rather than from B. Everything that makes a stronger probe valuable at eight hundred draws is more valuable there, by the same arithmetic and more of it.
And the benchmark disappears. There is no enumeration at two hundred units, so there is no separating direction, no separation to measure, and no way to check that a chosen probe is better than the one being replaced. The recommendation transports on its mechanism — the rule pins part of any probe, and the pinned part carries nothing — and the verification does not.
That is the ordinary position for a method calibrated where the truth is available and used where it is not, and it is worth stating rather than implying. What is measured here is that the projection helps at the one size where helping can be seen.
What the field comes to
Three sentences, and they are all about one line of code.
A probe is mostly the rule’s own span, on a trial the size trials are. The fourth power is exactly orthogonal to a rule balancing the first three powers and a median cut, and on fourteen units the rule has taken 0.9192 of it. That is a finite-sample fact with a closed form of zero behind it, and it is invisible from the population argument that chose the probe.
Projecting it back triples the separation and halves the misses. From 1.543 to 5.080 in the population quantity, and from 44% missed to 12% missed in the verdict, for a Gram–Schmidt pass against columns the implementation already has.
And choosing the direction is worth having and is not the main effect. Leverage, which needs no dictionary, is four times a random direction in the same subspace; a projection-pursuit direction that looks right is worse than random on both readings. The projection is the finding; the choice is a smaller second one; and searching for a good direction, on the one search tried, made it worse.
What is left is the half of the separation the enumeration keeps: 5.080 against 10.565, with an eighth of the separating direction lying inside the span no projected probe can reach. Nothing here closes that, and nothing here says it can be closed.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A quantity that loses to a heuristic — both name assignment mechanism, connected component, covariate balance, exact enumeration, experimental design, leverage, markov chain monte carlo, model diagnostics, orthogonality, projection, randomisation test
- Counting it exactly does not help — both name assignment mechanism, connected component, covariate balance, exact enumeration, experimental design, leverage, markov chain monte carlo, model diagnostics, orthogonality, projection, randomisation test
- A defect that is about size — both name assignment mechanism, connected component, covariate balance, exact enumeration, markov chain monte carlo, mixing time, randomisation test, reference distribution
- A model and a count — both name assignment mechanism, connected component, covariate balance, exact enumeration, experimental design, leverage, markov chain monte carlo, randomisation test
- The statistic that changes sign — both name connected component, covariate balance, effective sample size, exact enumeration, markov chain monte carlo, mixing time, randomisation test, reference distribution
- The statistic the p-value is about — both name assignment mechanism, connected component, covariate balance, effective sample size, exact enumeration, markov chain monte carlo, randomisation test, reference distribution
Named objects
A flat tag is an object no other essay names yet.
Assignment mechanismConnected componentCovariate balanceEffective sample sizeExact enumerationExperimental designLeverageMarkov chain Monte CarloMixing timeModel diagnosticsOrthogonalityProjectionRandomisation testReference distributionStatistical power