The diagnostic at two hundred
Worth reading first: Randomisation is not balance · Balancing what is known in advance.
The test agrees with the enumeration at every tolerance of a fourteen-unit trial, with no misses and no false alarms. The size sweep predicts that a two-hundred-unit trial’s admissible set will be in one piece, because the expected number of admissible neighbours there is over a hundred against a crossing at about two.
This essay points the instrument at two hundred units, and the answer is not the one either of those sets up. It is three answers, and which one comes back depends on how long the chains are run.
Three verdicts
The test returns one of three things. Quiet — the comparison against the complement is inside the threshold, and the walk reaches the whole set. Split — that comparison fires and both controls stay quiet, which has one explanation. Unmixed — a control fires, and the reading is about the run rather than about the set.
At forty thousand steps a chain, over eight independent runs at each of five tolerances, the answers are:
At a box of 1.0, admitting about one assignment in three: quiet, eight times out of eight. At 0.6, one in twelve: quiet, eight out of eight. At 0.4, one in thirty-seven: quiet six times and unmixed twice. At 0.3, one in seventy-two: split twice, unmixed five times, quiet once. At 0.22, one in two hundred: split twice and unmixed six times.
Past a point the diagnostic stops agreeing with itself. Three verdicts from eight runs of the same test on the same set is not a finding about the set; it is a finding about the instrument, and it is the honest answer at a size where nothing else can be said at all.
Why forty thousand steps is the interesting length
The eight-run sweep is at forty thousand steps a chain, and the choice is not arbitrary.
At two hundred units with three balancing functions, the chain’s statistic has an estimated autocorrelation time of between 140 and 385 depending on the tolerance. Forty thousand steps is therefore between one hundred and three hundred relaxation times — which by any ordinary standard is a long run, and is more than a practitioner reading a draw count would think to take.
It is also about what a randomisation test would actually use. A p-value quoted to three decimal places wants a few thousand draws; a cautious practitioner might take ten or twenty thousand and feel generous. Forty thousand is past that, and it is where the instrument disagrees with itself seven times out of eight at a tolerance a trial might plausibly run.
The length at which the diagnostic becomes reliable is not near the length at which the sampler feels finished, and there is nothing in the sampler’s own output that says so.
The sets are certainly connected, and the arithmetic says so
The verdicts include two “split” readings, and before anything is said about the instrument it is worth establishing what the truth must be, because at two hundred units it can be bounded without enumerating anything.
An assignment of two hundred units into equal arms has single exchanges available to it. Multiplying by the admitted share at each tolerance gives the expected number of admissible neighbours:
3,333, 833, 270, 139 and 50, at boxes of 1.0, 0.6, 0.4, 0.3 and 0.22.
The size sweep puts the crossing from several components to one at an expected degree near two. The tightest tolerance here is at fifty, twenty-five times that, and the independence model that produces these figures is known to under-count the true degree by a factor of about five.
So none of these sets is in pieces, and the two “split” verdicts are false alarms. That is not an inference from the diagnostic’s own inconsistency; it is arithmetic done before the diagnostic ran, and it is what licenses reading the whole sweep as a statement about the instrument.
The two failure modes arrive together
Counted as shares of eight runs, the “unmixed” verdicts run 0%, 0%, 25%, 62.5% and 75% across the five tolerances, and the “split” verdicts run 0%, 0%, 0%, 25% and 25%.
The splits appear exactly where the unmixed readings pass a quarter, and never before. That is the pattern a single underlying failure produces: a run that has not mixed reports “unmixed” when a control happens to fire and “split” when the controls happen to stay quiet, so the two verdicts are two outcomes of one event rather than two different things going wrong.
Which is the more useful reading of the sweep than “three verdicts from eight runs”. There are two verdicts and one of them is unreliable in a way the other reports.
The admission rate is a volume, checked
One arithmetic aside, because it confirms the tolerances are doing what a box on three constraints should do.
Fitting a power law to the five admitted shares against the five tolerances gives an exponent of about 2.8, which is the three balancing functions counted from outside — a box on k constraints admits a share proportional to once the region is small enough for the density inside it to be flat.
Extrapolating that fit down to a tolerance of half a standard deviation predicts an admitted share near one in twenty, against the one in nineteen the field that built the walk reports at exactly that setting. Two measurements taken for different purposes, agreeing to five per cent.
What a longer run says
Run the same tolerance at four lengths and the disagreement resolves in a direction.
At five thousand steps the verdict is unmixed: the signed comparison reads 5.77 and the magnitude control reads 4.21, so both are loud and nothing can be concluded. At twenty thousand it is split: 5.92 with both controls quiet, which is the pattern that means a mirror pair. At eighty thousand it is quiet: −1.15, with the replicate at 3.48. At three hundred and twenty thousand it is quiet again: 1.90, replicate 2.68.
The settled answer is that the walk reaches the whole set, which is what the size sweep predicts. The interesting part is the middle: at twenty thousand steps the test reported a split, with both controls quiet, on a set that does not have one.
A false alarm at one run length and the right answer at sixteen times that length is not a defect in the verdict rule. It is what happens when the standard errors under the comparison are wrong, and at two hundred units they are.
What the three verdicts mean together
The five tolerances are not five independent measurements; they are a ladder, and the shape of the ladder says something the individual cells do not.
At the two loosest tolerances every run agrees and the answer is that the walk reaches everything. That is not merely the test being confident — it is the test being easy, because a set admitting one assignment in three or twelve gives a chain a dense graph to move on, so it mixes quickly, its autocorrelation time is short, and the estimate of that time is believable.
At the three tightest, the acceptance rate falls, the chain moves less per step, the relaxation time rises, and the estimate of it degrades — all three at once, because they are the same thing. So the disagreement appears exactly where the instrument’s own assumption stops holding, and it appears gradually rather than at a threshold: two runs out of eight at one in thirty-seven, seven out of eight at one in seventy-two.
The instrument degrades where the sampler does, which is unfortunate and is not a coincidence. Both are functions of how well the chain moves, and there is no arrangement of chains that measures a set the chain cannot explore.
The quantity that fails is the effective sample size
The chains’ statistic is autocorrelated, and the comparison divides by a standard error computed from an estimated integrated autocorrelation time. At fourteen units that estimate is about thirty on twenty thousand draws — six hundred and change of effective sample — and it is believable.
At two hundred units the same estimator reports 351 and, at three hundred and twenty thousand steps, 911 effective draws. That estimate is what every reading in this field divides by, and it is the thing that is wrong at the short lengths: the initial-positive-sequence rule truncates its sum at the first pair of lags whose autocorrelations sum to a negative number, which on a short run happens early and reports too little dependence.
Too little dependence means too small a standard error, which means a comparison that fires when it should not. That is exactly what the twenty-thousand-step reading is.
The replicate control cannot fix that — it uses the same wrong standard error — but it does something better: it is wrong in the same way. When the standard error is too small, the replicate comparison fires too, and the verdict rule refuses to say anything. It does that at five thousand steps and it fails to at twenty thousand, which is the gap the eight-run spread above is measuring.
What this says about the p-value
The failure is not confined to the diagnostic. Every chain-based reference distribution in this collection divides by the same estimate.
What a reference distribution costs to sample establishes that a chain’s p-value resolves to one over its effective draws rather than one over its draw count, and the refusal that goes with it is a p-value quoted finer than the draws support. Both of those take the effective count as computable.
Here it is not, at the sizes trials actually run. A two-hundred-unit chain reporting 911 effective draws from three hundred and twenty thousand is reporting a number whose own estimate two chains from the same starting point disagree about by more than three standard errors at eighty thousand. The p-value’s resolution is therefore not one over 911; it is one over something smaller and unknown.
That is a sharper statement than the field started with and it is not the one it went looking for. The instrument built to find an unreachable half found that the quantity every chain-based p-value is quoted against cannot be estimated from the run.
The two numbers a practitioner would have quoted
It is worth being concrete about what a reasonable person would have done here, because the failure is not exotic.
A practitioner sampling a rerandomisation’s reference distribution at two hundred units would take a few thousand draws from the walk, look at the acceptance rate — 46% at a tolerance of 0.3, which is unremarkable — and report a p-value to three decimal places from the draw count. Nothing in that sequence is careless.
The two numbers that would have gone into the report are the draw count and the acceptance rate, and neither is the quantity that matters. The draw count overstates the resolution by the autocorrelation time, which is about 284 at forty thousand steps and 351 at three hundred and twenty thousand, and is itself unreliable. The acceptance rate says how often a proposal is taken and says nothing at all about whether the states taken cover the set.
What this field adds is a third number that can be computed in the same few thousand draws: the admitted share times a quarter of the square of the trial size. At a tolerance of 0.3 and two hundred units that is one in seventy-two times ten thousand, or about a hundred and forty admissible neighbours per assignment — which is the reassurance the chain could not supply.
What to do about it
Three things, in order of how much they cost.
Run longer than seems necessary. The verdict settles between eighty and three hundred and twenty thousand steps at two hundred units, which is one to two orders of magnitude past what a practitioner would choose from a draw count.
Run the replicate. It costs a third more and it is the only reading that distinguishes a structural fact from a short run. A diagnostic without it reported a split on a set that does not have one, at a run length a practitioner would plausibly use.
Read the degree. The count from the size sweep — the admitted share times a quarter of the square of the trial size — costs a few thousand rejection draws and answers the connectivity question directly at two hundred units, where it comes out over a hundred against a crossing at two. It is a prediction rather than a measurement, and at this size it is a much better instrument than the chain.
That last one is worth stating plainly because it inverts the field’s own premise. The test was built because enumeration stops at twenty-four units. At two hundred units the arithmetic does not stop, and it is the chain that has run out of reach.
Where this leaves the field
Four essays that set out to find an unreachable half at a realistic trial size end by reporting that at a realistic trial size there probably is not one, and that the instrument built to check cannot say so reliably in the number of steps anybody would run.
That is a smaller result than the field hoped for and it is not a null one, and the difference is worth naming.
The defect is real and it is small-trial. Fourteen units, four balancing functions, a tight tolerance — those are the conditions, and they are the conditions of a pilot study or a mechanistic experiment rather than of a clinical trial. The enumeration establishes it and the size sweep explains it.
The instrument is real and it is validated. Eight tolerances, no misses, no false alarms, at the size where the answer is known.
And the instrument’s reach is shorter than the defect’s absence. At two hundred units it needs one to two orders of magnitude more steps than a draw count would suggest before it stops contradicting itself, and the reason is a quantity every chain-based procedure in this collection already depends on.
Those three together are the field, and the third is the one that generalises. A test whose standard errors are estimated from the same run it is testing is a test that fails quietly where the run fails, and a diagnostic built from where a chain goes is not the only kind of instrument that has that shape.
What is claimed here, and what is not
This essay takes what the diagnostic says at a trial size nothing can enumerate. The claims are that at two hundred units and forty thousand steps the test is unanimous at loose tolerances and returns three different verdicts across eight runs at tight ones; that lengthening the run at one tolerance takes the verdict from unmixed to split to quiet and leaves it quiet, so the settled answer is that the walk reaches the whole set; that the estimated integrated autocorrelation time is 351 and the effective draws 911 out of three hundred and twenty thousand; and that the same estimate is what every chain-based p-value in this collection divides by.
What stays out, and is named as a decision: a better estimator of the autocorrelation time. There are several and none is used here. The reason is that the failure is not the estimator’s arithmetic — it is that a run of forty thousand steps with a relaxation time in the hundreds does not contain enough independent information to estimate its own relaxation time, and no estimator repairs that. The repair is a longer run or a different instrument, and both are named above.
Also out: whether the set is connected at two hundred units. The settled verdict says yes and the degree count predicts yes, and neither is a proof. What is established is that a diagnostic run at the length a practitioner would choose cannot tell, and that it will say something anyway.
The boundary against the essay that validates the test is that it checks the instrument where the answer is known and this one runs it where it is not — which is what the instrument was for, and is where its limits are.
The checks, and the refusals that make them mean something
Two claims are gated, and one of them is unusual for this collection: the sweep is required not to agree with itself. At the loosest tolerance every run must give the same verdict, and at the tightest they must not, because a diagnostic that were unanimous everywhere would mean the disagreement this essay reports is a seed rather than a property of the run length.
The other is the length sweep, which is required to return at least two different verdicts across the four lengths — which is what says the short runs are reporting the run.
The refusal is carried from the field’s first essay and now has a second reading. A chain’s own diagnostics are refused as evidence of reachability — and at two hundred units the test’s own standard errors are one of those diagnostics, which is why the verdict is a pattern across three chains rather than a number from one.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The statistic the p-value is about — both name connected component, covariate balance, effective sample size, ergodicity, imbalance, markov chain monte carlo, randomisation test, reference distribution, rerandomisation
- Before the trial and after — both name connected component, covariate balance, ergodicity, imbalance, markov chain monte carlo, randomisation test, reference distribution, rerandomisation
- Draws that repeat each other — both name burn-in, critical value, effective sample size, markov chain monte carlo, monte carlo, randomisation test, reference distribution, rerandomisation
- The set a dictionary leaves — both name burn-in, covariate balance, effective sample size, granularity, markov chain monte carlo, randomisation test, reference distribution, rerandomisation
- A proposal that moves more than two units — both name covariate balance, effective sample size, imbalance, monte carlo, randomisation test, reference distribution, rerandomisation
- Where the gain is, and where the decision is — both name covariate balance, effective sample size, imbalance, monte carlo, randomisation test, reference distribution, rerandomisation
Named objects
A flat tag is an object no other essay names yet.
Burn-inConnected componentCovariate balanceCritical valueEffective sample sizeErgodicityGranularityImbalanceMarkov chain Monte CarloMixing timeMonte CarloRandomisation testReference distributionRerandomisation