A schedule that reads the mean
Worth reading first: When the looking happens · What the exactness buys.
The condition that makes an adaptive block size exact is one line: the sizes may be anything at all provided they are functions of the within-block contrasts. Those are independent of every block mean, so conditioning on them leaves the block means where they were and the interval is a t interval on the realised number of blocks.
A condition stated that permissively invites the question of what it excludes, and the answer is short: any block mean, the overall mean, the between-block sum of squares, the interval’s current half-width, and anything computed from them.
Three rules that cross the line are measured here. Two are schedules and one is a stopping rule, and the stopping rule is much the worst.
A schedule that shrinks on the spread
The first offending rule is a natural one to write and hard to feel guilty about. The interval will be built from the between-block spread; the run can watch that spread as it goes; if it is running above what the contrasts say σ² is, take a small block, and otherwise take a large one.
The stated motivation would be that a run whose block means are scattered needs more blocks to pin them down. The actual effect is that the rule is adding heavily weighted means precisely when the spread is low and lightly weighted ones when it is high, which biases the estimate the interval is built from downwards.
It does. The between-block estimate lands at 0.8373σ² where the honest rule’s lands at 0.9831σ², and the coverage goes with it: 93.73% against 94.87%, with a counting error of 0.56 points.
A point and a bit of coverage is not a catastrophe, and that is the uncomfortable part. A rule that lost twenty points would be caught by anybody who checked. This one produces an interval that is slightly too narrow, in a procedure whose entire selling point is exactness, and no diagnostic on the output would show it.
What 0.8373 does to the interval, and what it should have cost
The estimate and the coverage are two readings of the same failure, and connecting them says whether anything else is going on.
The offending schedule’s between-block estimate is 0.8373σ² against the honest rule’s 0.9831σ², a ratio of 0.852. An interval’s half-width goes as the square root of that, so it is of the honest one — 7.7% too narrow.
Near 95% the sensitivity of coverage to width is points of coverage per per cent of width. So 7.7% of narrowing ought to cost about 1.8 points.
Measured, it costs 1.14 — from 94.87% to 93.73%. So the coverage is about a third less sensitive than the first-order calculation says, which is what a t interval does when the quantity being understated is the same quantity its multiplier is built from: the estimate falls and the multiplier rises against it, recovering part of the width.
That is worth knowing in the direction that matters. The bias in the estimate is a much larger number than the loss in coverage, 14.8% against 1.14 points, so a diagnostic that reads the coverage is a diagnostic reading the muffled version of the defect.
Which is why the violation is barely detectable
The counting error is 0.56 points, so the 1.14-point shortfall reads at 2.0 standard errors.
Two standard errors is a result nobody would build a case on. It is the threshold at which a careful reader says “possibly”, and it is what a genuine and mechanical violation of the exactness condition looks like on this many runs.
Getting it to three would take about 2.2 times the runs; getting it to four, four times. That is affordable in a simulation and it is not available at all to somebody looking at one trial, which had one realisation of one schedule and no honest comparison beside it.
So the honest summary of this rule is not that it is caught. It is that the estimate it produces is biased by fifteen per cent and the only symptom is a one-point coverage loss that needs thousands of runs to see, and the rule looks entirely reasonable written down.
A schedule that reads the running mean
The second offending rule reads the running mean directly: take a small block when the mean is near zero and a large one otherwise.
Here the coverage goes the other way — 96.27%, with the between-block estimate at 1.1510σ². The rule over-covers.
That is worth dwelling on, because over-covering is easy to mistake for safety. It is not a repair; it is the same violation with the sign reversed. The level is no longer a property of the procedure: it is a property of the procedure and of where the mean happened to be, and a rule whose coverage depends on the parameter it is estimating has no level at all in any useful sense. At a different mean it would land somewhere else, and the direction is not predictable from the rule’s description.
The two schedules together make the point better than either would alone. Reading the block means does not have a sign. It has whichever sign the particular rule happens to produce, and an experimenter who reasons about the direction of the bias from the shape of the rule will be right about half the time.
The one that costs eight points
The third rule is not a schedule at all. The blocks are fixed at two throughout; what reads the means is the stopping rule.
Instead of stopping when the observations in hand reach z²σ̂²/d², it stops when the interval it is going to report is narrower than d. That is a natural thing to write — it appears to be exactly what a fixed-width procedure is for, stated directly rather than through an estimate of σ — and it is the most expensive mistake in this field.
Coverage: 86.87%, against the honest rule’s 94.87%. The between-block estimate lands at 0.7954σ².
And it looks like a saving while it does it. The offending rule stops after 56.40 observations where the honest one takes 58.79 — two and a half fewer, which is very nearly the whole of what a schedule buys legitimately. An experimenter comparing the two on observations spent would prefer the broken one.
Why the mechanism is the same in all three
The three rules look different and the failure is one failure, which is worth stating in a form that covers a rule nobody has written yet.
The interval is , and its exactness rests on S_B² being σ² times a χ² independent of x̄. That holds because the block means are iid normal with the weights the interval uses, which in turn holds because the sizes were fixed before those means were seen.
Any rule that lets a block mean influence a block size breaks the second step. The realised weights are then correlated with the deviations they are weighting, S_B² is no longer an unbiased estimate of σ², and the t statistic is no longer a t statistic. Whether the estimate ends up too small or too large depends on the sign of the correlation the particular rule induces — 0.8373, 1.1510 and 0.7954 in the three cases here, against a truth of 1.
So the diagnostic is the estimate itself. A simulation of a proposed schedule that reports the mean of S_B² and finds it away from σ² has found the violation without needing a coverage count, and it will find it far more precisely: the estimate’s mean is measured to three decimals in a few thousand runs where a coverage rate is measured to half a point.
That is the useful output of this essay for anyone extending the construction. Do not check the coverage of a new schedule. Check that its between-block estimate is unbiased, which is the thing that actually has to hold and is far easier to see.
Why the stopping rule is worse than the schedules
The size of the third failure — eight points against one — has an explanation and it generalises.
A schedule that reads the means influences the weights in S_B². Each block still contributes its own mean, and the correlation is between a mean and the size of the next block, which is a second-order effect.
A stopping rule that reads the interval influences when the sum stops, and it stops it exactly when S_B² is small. That is direct selection on the quantity being estimated, of the same kind as stopping a test when the p-value is small, and it is a first-order effect. The estimate is not slightly biased; it is the minimum of a random walk, observed at the moment it went low.
The field that built the rule measures the same shape from a different angle, on an interval built from the estimate that stopped it. Everything about this construction is an attempt to keep the two quantities apart, and the third rule here puts them back together in the most direct way available.
How large the violations are, in the units that matter
Coverage is the headline and it is the noisiest thing on the page, so it is worth putting the three violations in the currency the mechanism is measured in.
The between-block estimate should be σ², and it comes out at 0.8373, 1.1510 and 0.7954. Those are 16% low, 15% high and 20% low, measured over fifteen hundred runs where the mean of an estimate is resolved to a few thousandths. Beside them the coverage differences — 1.13 points, 1.40 points and 8.00 points, against a counting error of 0.56 — are a much blunter instrument.
The relation between the two columns is not one to one, and the reason is instructive. A biased estimate of σ² feeds a width through a square root, and a width feeds coverage through a normal integral near its shoulder, so a 16% error in the estimate becomes an 8% error in the width and roughly a point of coverage. The stopping rule’s 20% is worse not because 20 is bigger than 16 but because the selection also changes the shape of the estimate’s distribution — it is a minimum rather than a shifted average, so the interval is narrow exactly on the runs where the mean has wandered.
That distinction is the reason the third failure is a different kind of object from the first two, and it is why the ratio of coverage losses is eight to one where the ratio of biases is five to four.
What an honest schedule may still do
The permissive half of the condition deserves restating, because the three failures above make the construction sound fragile and it is not.
An honest schedule may read the pooled within-block estimate of σ², the number of blocks so far, the sizes already used, the observations in hand, and the current target z²σ̂²/d². It may be adaptive, non-monotone, randomised, or chosen by an optimiser. It may take a block of two hundred and then a block of two. Every schedule in this field does some of that and every one of them covers at its nominal level.
The target in particular is worth naming again, because it is the one that feels like it must be excluded and is not. It is z²σ̂²/d² with σ̂² pooled from contrasts alone, so it carries no information about any block mean — and the stopping decision itself is a comparison of N with it, which is why the whole construction works at all.
The line is not between “adaptive” and “fixed”. It is between the contrasts and the means, and those are two disjoint functions of the same data.
Why nobody would catch this in practice
The three rules above are shown failing because they were simulated against a known truth. It is worth asking what an experimenter running one of them once would see, because the answer is nothing.
A single run produces a sample size, a set of block means, an estimate and an interval. Every one of those is a perfectly ordinary number. The interval is not obviously narrow — it is 8% narrow on average, which is well inside the run-to-run variation of an interval on a dozen degrees of freedom. The estimate is not obviously low. The sample size is, if anything, reassuringly small.
Nor would a replication help much. Two experimenters running the same broken rule would report intervals that overlap, and a meta-analysis of twenty of them would show a spread of point estimates slightly larger than the reported standard errors imply — which is the same signature as mild heterogeneity, is the commonest observation in any literature, and would be attributed to anything before it was attributed to the block-sizing rule.
That is the general shape of a failure that lives in a procedure rather than in a number. It does not produce an outlier, it produces a whole distribution that is slightly wrong, and the only way to see it is to simulate the procedure against a truth. Which is available, costs a few seconds, and requires somebody to have written the schedule down as a function first.
What a protocol has to record
The practical consequence is short and it is about writing rather than statistics.
The schedule has to be stated as a function. “Blocks were adapted as the trial progressed” is not a schedule; it is a description of one, and it is consistent with every rule in this essay including the three that do not cover. The function’s arguments are the part that matters.
Its arguments have to be listed. A schedule that takes the current target and the block count is exact. One that takes the running mean is not. Nothing downstream can tell them apart.
And the interval has to be the one the construction specifies. The weighted between-block sum of squares over b − 1, not the ordinary t interval on all N observations, which is right there and is narrower and does not cover.
None of that is onerous, and all three are things a protocol would have to contain anyway for the trial to be reproducible. The reason to insist is that all three failures above are invisible in the output: the interval looks like an interval, the width looks like a width, and the coverage is a property of a procedure nobody wrote down.
What is claimed here, and what is not
This essay takes the boundary of what a schedule may read, and the claims are the coverage and the between-block estimate for three offending rules, the observation that the two schedules fail in opposite directions, and the account of why a stopping rule that reads the interval fails by an order more than a schedule that reads the means.
What stays out and is named as a decision: any repair for the offending rules — a randomisation test over the schedule’s own reference set would restore a level and is a different construction with a different cost; the size of the violation as a function of how strongly the rule reads the means, which is a dial nobody turned because the rules here were written to be natural rather than extreme; and rules that read the means of other arms in a comparative trial, which is a larger question and is where this construction would go next.
The boundary against the field that built the blinded rule is the direction of the argument. That the contrasts and the means are independent, and that an interval built from the stopping estimate covers far below its level, are established there. What is new is that the independence licenses far more adaptivity than that field used, and that the licence has a sharp edge.
The checks, and the refusals that make them mean something
Both of this field’s refusals are in this essay, and they are stated as refusals rather than as findings because a construction whose validity rests on what a rule is allowed to read has to be shown failing when the rule reads more.
The first is the schedule that chooses the next block size from the between-block spread. It is required to cover below its nominal level by more than two standard errors, and its between-block estimate is required to be below 0.92σ² — the second half being the mechanism, which is measured far more precisely than the first and is what a new schedule should be checked on.
The second is the rule that stops when the interval it is going to report is narrow enough. It is required to cover below 90%, and it is required to spend fewer observations than the honest rule while doing so — because a refusal that only showed the failure would miss the reason nobody notices it.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- Stopping on the arms — both name blinding, blocking, coverage, fixed-width interval, independence, optional stopping, stopping rule
- The interval after a stop it chose — both name confidence interval, coverage, fixed-width interval, monte carlo, selection effect, sequential design, stopping rule
- A width rule on skewed outcomes — both name blinding, coverage, fixed-width interval, independence, monte carlo, stopping rule
- A width the trial has to stop for — both name blinding, blocking, coverage, fixed-width interval, optional stopping, stopping rule
- Stopping when it is precise enough — both name coverage, fixed-width interval, monte carlo, sample size, sequential design, stopping rule
- The bias that lands in the slope — both name blinding, blocking, coverage, degrees of freedom, fixed-width interval, nuisance parameter
Named objects
A flat tag is an object no other essay names yet.
BlindingBlockingConfidence intervalCoverageDegrees of freedomFixed-width intervalIndependenceMonte CarloNuisance parameterOptional stoppingSample sizeSelection effectSequential designStopping rule