What a schedule actually buys
Worth reading first: What the exactness buys · The shortest interval is the one that misses.
The reason to write a schedule was an argument that sounds right: the rule’s spread estimate matters early, when it decides how long the experiment will be, and the interval’s degrees of freedom matter late, when they decide how wide it is. So take large blocks first and small blocks last, and get both.
The argument is wrong in the form it was stated, and it is worth being exact about where.
There are not two ends to take. The interval’s degrees of freedom and the rule’s add to N − 1 on every run, so a schedule that increases one decreases the other. It does not matter when they are spent; the arithmetic is a division sum.
What is left after that is not nothing, and this essay is what it is.
Matched against the fixed size with the same number of blocks
The fair comparison is against a fixed block size with as many blocks, because the number of blocks is what sets the width per observation and a schedule with fewer blocks would be being compared with the wrong thing.
The front-loading schedule — sixteen until the target is in range, then twos — ends with 14.6 interval degrees of freedom. The nearest fixed size is four, at 13.7. Against each other:
| schedule | blocks of 4 | |
|---|---|---|
| observations | 58.36 | 60.73 |
| half-width | 0.2722 | 0.2647 |
| half-width × √N | 2.0798 | 2.0627 |
| fixed-width coverage | 92.53% | 93.80% |
| the rule’s degrees of freedom | 42.8 | 46.0 |
The third row is the one that settles it. Divide the width by the square root of the sample size and the two rules are within 0.8% of each other — they are the same point on the same curve. The schedule is not narrower per observation and it is not wider; it is the same rule, run for 2.4 fewer observations.
So the schedule’s whole contribution is 3.9% of the sample size, and it pays for it in width at exactly the rate 1/√N says it should.
Where the four per cent comes from
The overshoot, and nothing else.
A run stops at a multiple of its own block sizes, so it lands past its own target by about half a block: 1.48, 2.17, 3.11, 4.68 and 8.46 observations as the fixed size goes 2, 3, 5, 8 and 16.
Every schedule in this field ends in twos and lands where twos land: 1.62, 1.32 and 1.37. And it does that while having spent most of the run inside blocks four and eight times larger, which is why its rule has 42.8 degrees of freedom where blocks of two leave 32.5.
That is the genuine version of the original argument, and it is much narrower than the version that was hoped for. A schedule does not get more of both degrees of freedom. It gets the overshoot of a small block with the rule degrees of freedom of a large one, which is a real thing and is worth about half a block of observations.
The third quantity, which is where the argument nearly worked
There is one more column, and it is the one closest to the original instinct.
The rule’s spread estimate decides where the run stops, so the more degrees of freedom it has the less variable the sample size is. Across the fixed sizes the standard deviation of N falls from 16.23 at blocks of two to 13.39 at blocks of eight. The front-loading schedule sits at 15.11 with 58.36 observations — fewer observations than blocks of two, and a less variable stopping point.
Same observations, same width per observation, steadier sample size. That is a real improvement over the smallest fixed block and it is the only column in which the schedule dominates rather than trades.
It is also small, and it is worth saying why it does not show up in the fixed-width coverage. That coverage is E[2Φ(d√N/σ) − 1], which is dominated by E[N] and only second-order in the spread. The schedule has both fewer observations and a steadier stopping point, and the two push in opposite directions: 92.53% for the schedule against 92.40% for blocks of two, a difference well inside the counting error.
So the steadiness is real, measurable in sd(N), and worth almost nothing in the claim it was supposed to improve.
Why the halving schedule is the worst of the three
The three schedules are not equally good and the ordering is instructive.
Halve the distance to go produces very few blocks — 6.2 interval degrees of freedom — because each step covers half the remaining gap and the gap closes geometrically. Its width per observation is 2.3424, the worst of any rule in the field including blocks of sixteen at 2.4132. It spends 59.16 observations, which is not fewer than blocks of two.
A quarter of the distance produces 11.2 and a width per observation of 2.1409.
Sixteen, then two produces 14.6 and 2.0798, and spends the fewest observations of the three.
The pattern is that a schedule which spends a fraction of the remaining distance concentrates its blocks at the start, where the gap is large, and therefore ends up with few blocks in total. A schedule which spends a fixed large amount and then switches to twos gets the same rule degrees of freedom out of the early part of the run and many more blocks out of the late part.
The crude schedule beats both clever ones, and the reason is the arithmetic of the identity again: what matters is the split between the two degrees of freedom, and a fraction-of-the-gap rule chooses a worse split while looking more adaptive.
Where the overshoot is worth more than four per cent
Four per cent of sixty observations is two and a half, which is a small enough number that it is worth asking where it becomes a large one. The answer is available from the arithmetic without a further measurement.
The overshoot is about half a block. The sample size is about z²σ²/d². So the overshoot as a share of the experiment is roughly m/(2N), which grows as the requirement loosens and shrinks as it tightens. At a requirement of a quarter and blocks of sixteen the overshoot is 8.46 observations out of 67.7, which is 12.5%. At a requirement of a half, where an oracle needs 15.4 observations, the same block of sixteen is the entire experiment and the run overshoots by most of its own length.
So a schedule earns most in exactly the situation where a fixed block size is least defensible: a short run where the block is a large fraction of it. And it earns least where the experiment is long, which is where nobody would have worried.
That gives the field’s one piece of practical advice a shape. A schedule is worth writing when the experiment is expected to be short relative to the block size the spread estimate needs — which is the situation an experimenter is in whenever σ is poorly known and observations are costly, and is the situation the whole fixed-width apparatus exists for.
There is a second-order effect in the same direction. A short run has few blocks, so the interval’s degrees of freedom are scarce, so the t multiplier is large; and a schedule that ends in twos adds blocks exactly where they are most valuable. That is not a separate gain — it is still the same partition — but it does mean that the observations a schedule saves are saved at the point where each one is worth the most.
The curve every rule lands on is the t multiplier
“The schedules land on the fixed sizes’ curve” is the essay’s central claim and the curve has an equation, which makes the claim checkable rather than visual.
The width per observation should be proportional to the interval’s own t multiplier, since the width is and the rest divides out. Against the interval degrees of freedom each rule reports — 14.6 for the front-loading schedule, 11.2 for the quartering one, 6.2 for the halving one — the t multipliers are 2.14, 2.20 and 2.43, in ratios of 1.000, 1.028 and 1.136.
The counted widths per observation are 2.0798, 2.1409 and 2.3424, in ratios of 1.000, 1.029 and 1.126.
Agreement to within one per cent across a range from six degrees of freedom to fifteen. So the curve is not a curve fitted to the measurements: it is , and every rule in this field — schedule or fixed block, clever or crude — sits on it because the width per observation is a function of the block count and of nothing else.
That closes the field’s argument with an equation rather than a comparison. A rule’s width is decided by b, its stopping precision by N − b, and the two add to N − 1, so the whole design space is one number and the only question is where to put it. A schedule cannot move off the curve because there is nothing off the curve to move to.
The overshoot is half a block plus six tenths
The overshoot is described as about half a block, and the five readings say it is half a block plus a constant.
Subtracting m/2 from 1.48, 2.17, 3.11, 4.68 and 8.46 at block sizes 2, 3, 5, 8 and 16 leaves 0.48, 0.67, 0.61, 0.68 and 0.46 — a constant of about 0.6 observations, flat across an eightfold range of block size.
The half-block is the granularity: a run stops at a boundary and on average is half a block past its target when it does. The extra six tenths is the target itself moving, since σ̂ is re-estimated after every block and a run whose estimate rises finds its target has receded under it.
So the honest expression is overshoot ≈ m/2 + 0.6, and it says what a schedule can and cannot reach. Ending in twos takes the first term to 1 and leaves the second, which is why the three schedules land at 1.62, 1.32 and 1.37 rather than at zero — and why the four per cent a schedule saves is bounded below by a constant no arrangement of blocks removes.
What this says about the original instinct
The instinct — spend the rule’s degrees of freedom early and the interval’s late — is not wrong about what would be desirable. It is wrong about what is available, and the failure is worth generalising because it recurs.
A dial that trades two quantities looks like an optimisation problem until somebody checks whether the two quantities are independent. Here they are not: they are two parts of one total, and the whole apparent optimisation is a relabelling of a single choice. The schedule’s freedom is real but it lives somewhere else — in the granularity of the stopping decision, not in the accounting.
The honest summary a practitioner can use is three lines. Use a schedule if observations are expensive, because it saves about half a block of them. Choose the crude form, large blocks then twos, rather than a fraction of the remaining gap. And do not expect the interval to get narrower, because it will not: the width per observation is the number of blocks, and the number of blocks is half of a fixed budget however it is arranged.
What would have to change for the answer to be different
If the identity is what caps the gain, the useful question is what a construction would have to give up to escape it, and there are exactly three places to push.
Use the observations twice. The whole reason for the partition is that the interval refuses to use anything the stopping rule looked at. An interval built from all N observations has N − 1 degrees of freedom for both purposes and is narrower at every block size — and it does not cover, because the data at the stopping time was selected on its own spread. That is the trade the construction exists to refuse, and taking it back is not an improvement but a different procedure with a different guarantee.
Estimate σ from something else. If a pilot, a historical series or a parallel arm supplies a spread estimate that the stopping rule can use, the rule’s degrees of freedom stop competing with the interval’s and every block can be small. That is Stein’s two-stage rule in one direction and a genuine external prior in the other, and both are outside this construction because both need something the experiment does not contain.
Or stop on something other than the sample size. The requirement here is “N large enough that z²σ̂²/d² is met”, which is why N is a stopping time and why its variability costs anything. A rule that stopped on a quantity with no random horizon would not have the concavity problem at all — and would not be a fixed-width procedure.
None of the three is available cheaply, and listing them is the useful output of a negative result: the gain is capped by the partition, and the partition is what the exactness is bought with.
A negative result, reported as one
It is worth saying plainly that this is not the outcome the field was written to produce. The move was named as an obvious improvement, the construction generalised cleanly, every schedule is exact, and the gain is four per cent of the sample size and a steadier stopping point.
Two things make it worth having anyway.
The identity. Knowing that the two degrees of freedom are a partition of N − 1 rules out a whole family of proposals at once — anything of the form “spend the budget more cleverly” — and it is the kind of statement that only shows up when somebody tries.
And the direction. The measurement says where a real gain would have to come from: not from rearranging the blocks but from changing what the interval is built from, or from a stopping rule that uses the observations more than once without inheriting their selection. Both of those are outside this construction, and neither is available for free.
What is claimed here, and what is not
This essay takes what a schedule is worth, and the claims are the matched comparison against the fixed size with the same number of blocks, the overshoot as the source of the difference, the ordering of the three schedules, and the reduction in the sample size’s variability that does not show up in the coverage it was supposed to improve.
What stays out and is named as a decision: schedules tuned to the requirement rather than fixed across it, which would help at the loose end where the overshoot is a larger share and are not measured; a schedule chosen by optimising over a parametrised family, which is a search and would need its own selection correction; and any comparison against rules outside this construction, since the exactness is what is being held constant and a rule that gives it up is answering a different question.
The boundary against the previous essay is what is being measured. That the width is the number of blocks and the fixed-width claim is the sample-size distribution is established there. What is new is the matched comparison, which is the only way to see that a schedule moves along the curve rather than off it.
The checks, and the refusals that make them mean something
One claim is gated in this field’s library, in three parts. Every schedule is required to be within a few per cent of the fixed size with the same number of blocks, on width per observation — which is the statement that the schedules are on the curve and not above it, and is what makes this a negative result rather than a positive one. The front-loading schedule is required to spend no more observations than the smallest fixed block while giving the rule at least a fifth more degrees of freedom. And its sample size is required to be the less variable of the two, which is the only column in which it dominates.
The refusals for this field are in the next essay, and both are schedules that read what they may not. The relevant one here is the rule that stops on the interval it is about to report: it spends fewer observations than any rule in this essay, which is exactly what a schedule is being asked to do, and it buys them out of the coverage instead of out of the overshoot.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A ratio that changes between blocks — both name blinding, blocking, coverage, degrees of freedom, efficiency, fixed-width interval, nuisance parameter
- A width promised for a difference — both name blinding, coverage, degrees of freedom, fixed-width interval, stopping rule, variance
- Blinded, and still exact — both name blinding, coverage, degrees of freedom, efficiency, fixed-width interval, nuisance parameter
- The bias that lands in the slope — both name blinding, blocking, coverage, degrees of freedom, fixed-width interval, nuisance parameter
- Weights that need only a ratio — both name blocking, coverage, degrees of freedom, efficiency, fixed-width interval, nuisance parameter
- What a two-arm rule may not pool — both name blinding, coverage, degrees of freedom, fixed-width interval, sample size, stopping rule
Named objects
A flat tag is an object no other essay names yet.
BlindingBlockingCoverageDegrees of freedomEfficiencyFixed-width intervalMonte CarloNuisance parameterOvershootPrecisionSample sizeSequential designStopping ruleVariance