The block size as a schedule

A schedule that reads the mean

The block sizes may be anything at all provided they are functions of the contrasts. Two natural schedules break that, in opposite directions — and the most natural mistake of the three is not a schedule at all but a stopping rule, at 86.87% coverage and fewer observations.

Worth reading first: When the looking happens · What the exactness buys.

The condition that makes an adaptive block size exact is one line: the sizes may be anything at all provided they are functions of the within-block contrasts. Those are independent of every block mean, so conditioning on them leaves the block means where they were and the interval is a t interval on the realised number of blocks.

A condition stated that permissively invites the question of what it excludes, and the answer is short: any block mean, the overall mean, the between-block sum of squares, the interval’s current half-width, and anything computed from them.

Three rules that cross the line are measured here. Two are schedules and one is a stopping rule, and the stopping rule is much the worst.

What a schedule is allowed to read, and what happens when it reads more. The construction allows the block sizes to be anything at all as long as they are functions of the within-block contrasts, which are independent of every block mean. A schedule that shrinks the block whenever the between-block spread is running above what the contrasts say is a direct attempt to hold down the quantity the interval will be built from, and it succeeds: the estimate lands at 0.8373σ² against the honest 0.9831, and the coverage goes with it. Reading the running mean instead pushes the other way and over-covers — which is not a repair, it is the same violation with the sign reversed, and the level is no longer a property of the procedure at all.
Fig. 1 Three ways of sizing a block, one of them legitimate. The estimate the interval is built from is printed beside each bar, because that is where the coverage goes.

A schedule that shrinks on the spread

The first offending rule is a natural one to write and hard to feel guilty about. The interval will be built from the between-block spread; the run can watch that spread as it goes; if it is running above what the contrasts say σ² is, take a small block, and otherwise take a large one.

The stated motivation would be that a run whose block means are scattered needs more blocks to pin them down. The actual effect is that the rule is adding heavily weighted means precisely when the spread is low and lightly weighted ones when it is high, which biases the estimate the interval is built from downwards.

It does. The between-block estimate lands at 0.8373σ² where the honest rule’s lands at 0.9831σ², and the coverage goes with it: 93.73% against 94.87%, with a counting error of 0.56 points.

A point and a bit of coverage is not a catastrophe, and that is the uncomfortable part. A rule that lost twenty points would be caught by anybody who checked. This one produces an interval that is slightly too narrow, in a procedure whose entire selling point is exactness, and no diagnostic on the output would show it.

What 0.8373 does to the interval, and what it should have cost

The estimate and the coverage are two readings of the same failure, and connecting them says whether anything else is going on.

The offending schedule’s between-block estimate is 0.8373σ² against the honest rule’s 0.9831σ², a ratio of 0.852. An interval’s half-width goes as the square root of that, so it is 0.852=0.923\sqrt{0.852} = 0.923 of the honest one — 7.7% too narrow.

Near 95% the sensitivity of coverage to width is 2φ(1.96)×1.96=0.2292\varphi(1.96) \times 1.96 = 0.229 points of coverage per per cent of width. So 7.7% of narrowing ought to cost about 1.8 points.

Measured, it costs 1.14 — from 94.87% to 93.73%. So the coverage is about a third less sensitive than the first-order calculation says, which is what a t interval does when the quantity being understated is the same quantity its multiplier is built from: the estimate falls and the multiplier rises against it, recovering part of the width.

That is worth knowing in the direction that matters. The bias in the estimate is a much larger number than the loss in coverage, 14.8% against 1.14 points, so a diagnostic that reads the coverage is a diagnostic reading the muffled version of the defect.

Which is why the violation is barely detectable

The counting error is 0.56 points, so the 1.14-point shortfall reads at 2.0 standard errors.

Two standard errors is a result nobody would build a case on. It is the threshold at which a careful reader says “possibly”, and it is what a genuine and mechanical violation of the exactness condition looks like on this many runs.

Getting it to three would take about 2.2 times the runs; getting it to four, four times. That is affordable in a simulation and it is not available at all to somebody looking at one trial, which had one realisation of one schedule and no honest comparison beside it.

So the honest summary of this rule is not that it is caught. It is that the estimate it produces is biased by fifteen per cent and the only symptom is a one-point coverage loss that needs thousands of runs to see, and the rule looks entirely reasonable written down.

A schedule that reads the running mean

The second offending rule reads the running mean directly: take a small block when the mean is near zero and a large one otherwise.

Here the coverage goes the other way96.27%, with the between-block estimate at 1.1510σ². The rule over-covers.

That is worth dwelling on, because over-covering is easy to mistake for safety. It is not a repair; it is the same violation with the sign reversed. The level is no longer a property of the procedure: it is a property of the procedure and of where the mean happened to be, and a rule whose coverage depends on the parameter it is estimating has no level at all in any useful sense. At a different mean it would land somewhere else, and the direction is not predictable from the rule’s description.

The two schedules together make the point better than either would alone. Reading the block means does not have a sign. It has whichever sign the particular rule happens to produce, and an experimenter who reasons about the direction of the bias from the shape of the rule will be right about half the time.

What a schedule is allowed to read, and what happens when it reads more. The construction allows the block sizes to be anything at all as long as they are functions of the within-block contrasts, which are independent of every block mean. A schedule that shrinks the block whenever the between-block spread is running above what the contrasts say is a direct attempt to hold down the quantity the interval will be built from, and it succeeds: the estimate lands at 0.8158σ² against the honest 0.9691, and the coverage goes with it. Reading the running mean instead pushes the other way and over-covers — which is not a repair, it is the same violation with the sign reversed, and the level is no longer a property of the procedure at all.
Fig. 2 The same three rules at a looser requirement. The direction of each violation is a property of the rule and not of the setting.

The one that costs eight points

The third rule is not a schedule at all. The blocks are fixed at two throughout; what reads the means is the stopping rule.

Instead of stopping when the observations in hand reach z²σ̂²/d², it stops when the interval it is going to report is narrower than d. That is a natural thing to write — it appears to be exactly what a fixed-width procedure is for, stated directly rather than through an estimate of σ — and it is the most expensive mistake in this field.

Coverage: 86.87%, against the honest rule’s 94.87%. The between-block estimate lands at 0.7954σ².

And it looks like a saving while it does it. The offending rule stops after 56.40 observations where the honest one takes 58.79 — two and a half fewer, which is very nearly the whole of what a schedule buys legitimately. An experimenter comparing the two on observations spent would prefer the broken one.

Stopping on the interval you are going to report. The same blocks, the same data, and one change: the rule waits until the interval built from the block means is narrower than 0.25, rather than until the observations in hand reach what the contrasts say is enough. Coverage goes from 94.87% to 86.87% — and it looks like a saving while it happens, because the offending rule stops after 56.4 observations where the honest one takes 58.8. The quantity that decides when to stop and the quantity the interval is built from have to be independent, and here they are the same number.
Fig. 3 Stopping on the interval you are going to report. Eight points of coverage, and it looks like a saving of two and a half observations.

Why the mechanism is the same in all three

The three rules look different and the failure is one failure, which is worth stating in a form that covers a rule nobody has written yet.

The interval is xˉ±tb1SB2/N\bar{x} \pm t_{b-1}\sqrt{S_B^2/N}, and its exactness rests on S_B² being σ² times a χ² independent of x̄. That holds because the block means are iid normal with the weights the interval uses, which in turn holds because the sizes were fixed before those means were seen.

Any rule that lets a block mean influence a block size breaks the second step. The realised weights are then correlated with the deviations they are weighting, S_B² is no longer an unbiased estimate of σ², and the t statistic is no longer a t statistic. Whether the estimate ends up too small or too large depends on the sign of the correlation the particular rule induces — 0.8373, 1.1510 and 0.7954 in the three cases here, against a truth of 1.

So the diagnostic is the estimate itself. A simulation of a proposed schedule that reports the mean of S_B² and finds it away from σ² has found the violation without needing a coverage count, and it will find it far more precisely: the estimate’s mean is measured to three decimals in a few thousand runs where a coverage rate is measured to half a point.

That is the useful output of this essay for anyone extending the construction. Do not check the coverage of a new schedule. Check that its between-block estimate is unbiased, which is the thing that actually has to hold and is far easier to see.

Why the stopping rule is worse than the schedules

The size of the third failure — eight points against one — has an explanation and it generalises.

A schedule that reads the means influences the weights in S_B². Each block still contributes its own mean, and the correlation is between a mean and the size of the next block, which is a second-order effect.

A stopping rule that reads the interval influences when the sum stops, and it stops it exactly when S_B² is small. That is direct selection on the quantity being estimated, of the same kind as stopping a test when the p-value is small, and it is a first-order effect. The estimate is not slightly biased; it is the minimum of a random walk, observed at the moment it went low.

The field that built the rule measures the same shape from a different angle, on an interval built from the estimate that stopped it. Everything about this construction is an attempt to keep the two quantities apart, and the third rule here puts them back together in the most direct way available.

Stopping on the interval you are going to report. The same blocks, the same data, and one change: the rule waits until the interval built from the block means is narrower than 0.35, rather than until the observations in hand reach what the contrasts say is enough. Coverage goes from 95.47% to 86.47% — and it looks like a saving while it happens, because the offending rule stops after 30.6 observations where the honest one takes 29.6. The quantity that decides when to stop and the quantity the interval is built from have to be independent, and here they are the same number.
Fig. 4 The same rule at a looser requirement, where a run is shorter and the selection has fewer opportunities. The failure is smaller and it does not go away.

How large the violations are, in the units that matter

Coverage is the headline and it is the noisiest thing on the page, so it is worth putting the three violations in the currency the mechanism is measured in.

The between-block estimate should be σ², and it comes out at 0.8373, 1.1510 and 0.7954. Those are 16% low, 15% high and 20% low, measured over fifteen hundred runs where the mean of an estimate is resolved to a few thousandths. Beside them the coverage differences — 1.13 points, 1.40 points and 8.00 points, against a counting error of 0.56 — are a much blunter instrument.

The relation between the two columns is not one to one, and the reason is instructive. A biased estimate of σ² feeds a width through a square root, and a width feeds coverage through a normal integral near its shoulder, so a 16% error in the estimate becomes an 8% error in the width and roughly a point of coverage. The stopping rule’s 20% is worse not because 20 is bigger than 16 but because the selection also changes the shape of the estimate’s distribution — it is a minimum rather than a shifted average, so the interval is narrow exactly on the runs where the mean has wandered.

That distinction is the reason the third failure is a different kind of object from the first two, and it is why the ratio of coverage losses is eight to one where the ratio of biases is five to four.

What a schedule is allowed to read, and what happens when it reads more. The construction allows the block sizes to be anything at all as long as they are functions of the within-block contrasts, which are independent of every block mean. A schedule that shrinks the block whenever the between-block spread is running above what the contrasts say is a direct attempt to hold down the quantity the interval will be built from, and it succeeds: the estimate lands at 0.8002σ² against the honest 0.9499, and the coverage goes with it. Reading the running mean instead pushes the other way and over-covers — which is not a repair, it is the same violation with the sign reversed, and the level is no longer a property of the procedure at all.
Fig. 5 The estimates at the loosest requirement measured. The biases are properties of the rules and they survive the setting.

What an honest schedule may still do

The permissive half of the condition deserves restating, because the three failures above make the construction sound fragile and it is not.

An honest schedule may read the pooled within-block estimate of σ², the number of blocks so far, the sizes already used, the observations in hand, and the current target z²σ̂²/d². It may be adaptive, non-monotone, randomised, or chosen by an optimiser. It may take a block of two hundred and then a block of two. Every schedule in this field does some of that and every one of them covers at its nominal level.

The target in particular is worth naming again, because it is the one that feels like it must be excluded and is not. It is z²σ̂²/d² with σ̂² pooled from contrasts alone, so it carries no information about any block mean — and the stopping decision itself is a comparison of N with it, which is why the whole construction works at all.

The line is not between “adaptive” and “fixed”. It is between the contrasts and the means, and those are two disjoint functions of the same data.

A schedule is exact for the same reason a fixed size is. Coverage of the interval built from the block means, over 1500 runs at a requirement of 0.25. The block means are independent of every within-block contrast whatever the sizes are, so a rule that stops on the contrasts and reports the weighted mean with Σm_b(ȳ_b − x̄)²/(b − 1) behind it is exact conditional on the sizes — and the sizes may then be chosen from the contrasts, adaptively, without touching the argument. Each bar is within 2.3% of the level it claims. Nothing here is an approximation that improves with the sample size.
Fig. 6 The honest side of the line: three fixed sizes and three schedules, all exact.
What a schedule is allowed to read, and what happens when it reads moreThe construction allows the block sizes to be anything at all as long as they are functions of the within-block contrasts, which are independent of every block mean. A schedule that shrinks the block whenever the between-block spread is running above what the contrasts say is a direct attempt to hold down the quantity the interval will be built from, and it succeeds: the estimate lands at 0.8373σ² against the honest 0.9831, and the coverage goes with it. Reading the running mean instead pushes the other way and over-covers — which is not a repair, it is the same violation with the sign reversed, and the level is no longer a property of the procedure at all.sizes from the contrasts94.87%sizes from the between-block spread93.73%sizes from the running mean96.27%spread 0.983σ²spread 0.837σ²spread 1.151σ²how the next block size is chosen1500 runs, requirement 0.25, nominal 95%±1.13% on each bar
Fig. 7 Drag the requirement. The two offending schedules keep their directions and the honest rule keeps its level.

Why nobody would catch this in practice

The three rules above are shown failing because they were simulated against a known truth. It is worth asking what an experimenter running one of them once would see, because the answer is nothing.

A single run produces a sample size, a set of block means, an estimate and an interval. Every one of those is a perfectly ordinary number. The interval is not obviously narrow — it is 8% narrow on average, which is well inside the run-to-run variation of an interval on a dozen degrees of freedom. The estimate is not obviously low. The sample size is, if anything, reassuringly small.

Nor would a replication help much. Two experimenters running the same broken rule would report intervals that overlap, and a meta-analysis of twenty of them would show a spread of point estimates slightly larger than the reported standard errors imply — which is the same signature as mild heterogeneity, is the commonest observation in any literature, and would be attributed to anything before it was attributed to the block-sizing rule.

That is the general shape of a failure that lives in a procedure rather than in a number. It does not produce an outlier, it produces a whole distribution that is slightly wrong, and the only way to see it is to simulate the procedure against a truth. Which is available, costs a few seconds, and requires somebody to have written the schedule down as a function first.

The two degrees of freedom are a partition of one total. Each point is one run of one schedule. Every point sits on a line of slope −1, because (b − 1) + (N − b) = N − 1 exactly: every observation after the first gives a degree of freedom to the rule's spread estimate or to the interval, and never to both. That identity is the whole reason there is no schedule that improves both ends of the block-size dial. The lines are the runs that ended with the same N, and a schedule moves along one of them rather than off it — what it can decide is only when in a run each degree of freedom is spent.
Fig. 8 The honest construction’s own accounting, which is what a simulation of a proposed schedule should be checked against before its coverage is.

What a protocol has to record

The practical consequence is short and it is about writing rather than statistics.

The schedule has to be stated as a function. “Blocks were adapted as the trial progressed” is not a schedule; it is a description of one, and it is consistent with every rule in this essay including the three that do not cover. The function’s arguments are the part that matters.

Its arguments have to be listed. A schedule that takes the current target and the block count is exact. One that takes the running mean is not. Nothing downstream can tell them apart.

And the interval has to be the one the construction specifies. The weighted between-block sum of squares over b − 1, not the ordinary t interval on all N observations, which is right there and is narrower and does not cover.

None of that is onerous, and all three are things a protocol would have to contain anyway for the trial to be reproducible. The reason to insist is that all three failures above are invisible in the output: the interval looks like an interval, the width looks like a width, and the coverage is a property of a procedure nobody wrote down.

No block size is best at both things the procedure claims. Two claims and one dial. The honest interval's half-width falls as the blocks get smaller, because the interval's degrees of freedom are the number of blocks: 0.2602 at blocks of two against 0.2933 at blocks of sixteen. The fixed-width claim — that the mean is within 0.25 of the truth — gets more reliable as they get larger, because the sample size is less variable: 92.40% against 95.00%. Both are computed from the same runs, and the second is reproduced to within a tenth of a point by E[2Φ(d√N/σ) − 1], which needs the sample-size distribution and nothing else. The schedules sit at the bottom left: as narrow as the smallest fixed block and as few observations, with the rule's spread estimate on half as many degrees of freedom again.
Fig. 9 Where the honest rules sit, for comparison: every one of them exact, and the trade between them a matter of a few observations and a few thousandths of width.

What is claimed here, and what is not

This essay takes the boundary of what a schedule may read, and the claims are the coverage and the between-block estimate for three offending rules, the observation that the two schedules fail in opposite directions, and the account of why a stopping rule that reads the interval fails by an order more than a schedule that reads the means.

What stays out and is named as a decision: any repair for the offending rules — a randomisation test over the schedule’s own reference set would restore a level and is a different construction with a different cost; the size of the violation as a function of how strongly the rule reads the means, which is a dial nobody turned because the rules here were written to be natural rather than extreme; and rules that read the means of other arms in a comparative trial, which is a larger question and is where this construction would go next.

The boundary against the field that built the blinded rule is the direction of the argument. That the contrasts and the means are independent, and that an interval built from the stopping estimate covers far below its level, are established there. What is new is that the independence licenses far more adaptivity than that field used, and that the licence has a sharp edge.

The checks, and the refusals that make them mean something

Both of this field’s refusals are in this essay, and they are stated as refusals rather than as findings because a construction whose validity rests on what a rule is allowed to read has to be shown failing when the rule reads more.

The first is the schedule that chooses the next block size from the between-block spread. It is required to cover below its nominal level by more than two standard errors, and its between-block estimate is required to be below 0.92σ² — the second half being the mechanism, which is measured far more precisely than the first and is what a new schedule should be checked on.

The second is the rule that stops when the interval it is going to report is narrow enough. It is required to cover below 90%, and it is required to spend fewer observations than the honest rule while doing so — because a refusal that only showed the failure would miss the reason nobody notices it.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

BlindingBlockingConfidence intervalCoverageDegrees of freedomFixed-width intervalIndependenceMonte CarloNuisance parameterOptional stoppingSample sizeSelection effectSequential designStopping rule