The block size as a schedule

A block size that changes

The blinded rule's exactness never needed the blocks to be the same size. Letting the size be chosen from the contrasts as the run goes on leaves the coverage exactly where it was — and runs straight into an identity that says what a schedule can and cannot buy.

Worth reading first: When the looking happens · The shortest interval is the one that misses.

An earlier field builds a fixed-width interval whose coverage is exact, and the construction is one sentence. Observations arrive in blocks. The rule stops on the within-block contrasts; the interval is built from the block means; the two are independent whatever the rule does, so the interval is a t interval on the number of blocks minus one and its coverage is what it says it is.

That field measures the construction at five block sizes, finds the size to be a dial between three costs with a different best value at each requirement, and names the obvious next move without making it: stop choosing one size. Large blocks early, when the spread estimate is poor and the target is far away and cannot be overshot; small blocks near the crossing, where a block is the granularity of the answer.

This essay makes that move. The construction survives, and then it meets an identity.

One experiment, with the blocks getting smaller as the target comes into range. A single run at a requirement of 0.25, with the block sizes 5, 5, 11, 25, 11, 8, 3, 2 and a total of 70 observations in 8 blocks. The rule stops when the observations in hand reach z²σ̂²/d², with σ̂² pooled from the within-block contrasts — an estimate that moves as the run goes on, so the target moves too. Early blocks are large because the target is far away and cannot be overshot; late ones are small because a block is the granularity of the answer. The interval afterwards is built from the 8 block means and from nothing the rule looked at, and it has 7 degrees of freedom against the rule's 62.
Fig. 1 One experiment with blocks of 5, 5, 11, 25, 11, 8, 3 and 2. The blocks grow while the target is far away and shrink as it comes into range.

Why the argument survives, in full

The exactness claim has to be re-derived rather than asserted, because the original is stated for equal blocks and the weighting changes when they are not.

Let block b have m_b observations and mean ȳ_b. Then √m_b (ȳ_b − μ) is σ times a standard normal, independently across blocks, whatever the sizes are. The overall mean

xˉ=bmbyˉbN\bar x = \frac{\sum_b m_b \bar y_b}{N}

is the ordinary sample mean of all N observations, and its variance is σ²/N exactly. And

SB2=bmb(yˉbxˉ)2b1S_B^2 = \frac{\sum_b m_b (\bar y_b - \bar x)^2}{b - 1}

is σ² times a χ² on b − 1 degrees of freedom, independent of x̄ — the ordinary weighted least squares decomposition, with the fixed-block case as the special case where all the weights are equal.

So xˉ±tb1SB2/N\bar{x} \pm t_{b-1}\sqrt{S_B^2/N} is exact conditional on the sizes and the number of blocks. And the sizes and the number of blocks may be anything at all, provided they are functions of the contrasts — which are independent of every block mean, so conditioning on them changes nothing about the block means’ distribution.

That is the whole generalisation. The sizes may be adaptive, they may depend on the pooled spread estimate, they may depend on each other, and there need be no pattern to them. What they may not depend on is any block mean, which is the refusal at the end of this field.

A schedule is exact for the same reason a fixed size is. Coverage of the interval built from the block means, over 1500 runs at a requirement of 0.25. The block means are independent of every within-block contrast whatever the sizes are, so a rule that stops on the contrasts and reports the weighted mean with Σm_b(ȳ_b − x̄)²/(b − 1) behind it is exact conditional on the sizes — and the sizes may then be chosen from the contrasts, adaptively, without touching the argument. Each bar is within 2.3% of the level it claims. Nothing here is an approximation that improves with the sample size.
Fig. 2 Three fixed sizes and three schedules, at a requirement of a quarter. Every bar is within a couple of standard errors of the level it claims.

Counted rather than argued

A theorem about a construction this fiddly earns a count, and the count is the first table of this field.

At a requirement of 0.25, where an oracle who knew σ would need 61.5 observations, blocks of 2, 5 and 16 cover 94.67%, 94.80% and 95.33%, and the halving, quartering and front-loading schedules cover 94.47%, 94.20% and 94.53%. The counting error is 0.56 points.

At a looser requirement of 0.5, where the oracle needs 15.4 observations and there are only a handful of blocks in the whole run, the same six cover 94.60%, 95.07%, 95.13%, 94.53%, 94.67% and 94.60%.

Nothing about those numbers improves with the sample size, because nothing about them is asymptotic. At fifteen observations and at sixty the interval is a t interval on the realised number of blocks, and the only reason the counted figure is not exactly 95% is that a simulation of three thousand runs does not resolve better than half a point.

Three schedules, and what they look like from inside

The three used throughout this field are stated in one line each.

Halve the distance to go. The next block is half the gap between the observations in hand and the current target, floored at two. The blocks grow while the estimate of σ is rising and shrink as the target is approached.

A quarter of the distance to go. The same with a gentler step, which produces more blocks and smaller ones.

Sixteen, then two. A fixed large block until the target is within reach, then twos. The crudest of the three and, as it turns out, the most useful.

A single run makes the difference concrete. On the same stream of observations, the halving schedule takes blocks of 5, 5, 11, 25, 11, 8, 3, 2 — 70 observations in 8 blocks, leaving the rule 62 degrees of freedom and the interval 7, and reporting a half-width of 0.2796. The front-loading schedule takes 76 observations in 21 blocks, leaving the rule 55 and the interval 20, and reports 0.1935. A fixed size of four takes 74 in 18, leaving 56 and 17, and reports 0.2085.

One experiment, with the blocks getting smaller as the target comes into range. A single run at a requirement of 0.25, with the block sizes 5, 5, 2, 2, 2, 16, 16, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2 and a total of 76 observations in 21 blocks. The rule stops when the observations in hand reach z²σ̂²/d², with σ̂² pooled from the within-block contrasts — an estimate that moves as the run goes on, so the target moves too. Early blocks are large because the target is far away and cannot be overshot; late ones are small because a block is the granularity of the answer. The interval afterwards is built from the 21 block means and from nothing the rule looked at, and it has 20 degrees of freedom against the rule's 55.
Fig. 3 The front-loading schedule on the same stream. Its sizes are not monotone, because the target moves as the spread estimate improves and a run can find itself far from a target it was close to.

The non-monotonicity in that second run is worth a sentence, because it looks like a bug. The target is z²σ̂²/d² and σ̂ is re-estimated after every block, so a run that has been unlucky in its early contrasts can watch its target recede — and a schedule reading the current gap will then take a large block again after several small ones. That is the schedule doing exactly what it was told; it is only surprising if the target is imagined as fixed.

The identity

Here is where the field stops being an optimisation and becomes an accounting problem.

The rule’s spread estimate is pooled from the within-block contrasts and has N − b degrees of freedom. The interval is a t interval on the block means and has b − 1. And

(b1)+(Nb)=N1,(b - 1) + (N - b) = N - 1,

exactly, on every run of every schedule. Every observation after the first contributes a degree of freedom to one of the two estimates and never to both.

The two degrees of freedom are a partition of one total. Each point is one run of one schedule. Every point sits on a line of slope −1, because (b − 1) + (N − b) = N − 1 exactly: every observation after the first gives a degree of freedom to the rule's spread estimate or to the interval, and never to both. That identity is the whole reason there is no schedule that improves both ends of the block-size dial. The lines are the runs that ended with the same N, and a schedule moves along one of them rather than off it — what it can decide is only when in a run each degree of freedom is spent.
Fig. 4 Every run of every schedule, plotted by its two degrees of freedom. All of them lie on lines of slope −1, because the two numbers are a partition of N − 1.

That is not an approximation and it is not a tendency. It is arithmetic, and it settles the shape of everything that follows: a schedule cannot make more of both. What it can decide is only when in a run each degree of freedom is spent — and whether that matters at all is the subject of the next two essays.

The instinct the earlier field expressed — that large blocks early and small blocks late would take both ends of the trade — is therefore wrong in the form it was stated, and interestingly wrong. There is no both-ends to take. There is a fixed budget of N − 1 and a decision about how to split it, plus one genuinely free quantity that the identity does not constrain and that turns out to be where the whole gain lives.

What the identity does not cover

Three things are not fixed by N − 1, and it is worth naming them now because they are what the field has left to measure.

The sample size itself. N is not a constant across rules; it is a stopping time, and a rule with a better spread estimate stops in a different place. So the budget being partitioned is itself a function of the partition, which is the loop that makes this a design problem rather than a division sum.

The overshoot. A run stops at a multiple of its own block sizes and cannot land between them, so a rule with large blocks lands past its own target. That is a pure loss and it is the only quantity in this field a schedule reduces without paying for it anywhere.

And the variability of N. Two rules can have the same mean sample size and different spreads, and the fixed-width claim — that the mean is within d of the truth — is a concave function of N, so a more variable sample size is worse at the same average.

Those three are the whole of what a schedule can move, and the next essay measures each of them.

A schedule is exact for the same reason a fixed size isCoverage of the interval built from the block means, over 1500 runs at a requirement of 0.5. The block means are independent of every within-block contrast whatever the sizes are, so a rule that stops on the contrasts and reports the weighted mean with Σm_b(ȳ_b − x̄)²/(b − 1) behind it is exact conditional on the sizes — and the sizes may then be chosen from the contrasts, adaptively, without touching the argument. Each bar is within 2.3% of the level it claims. Nothing here is an approximation that improves with the sample size.blocks of 294.60%blocks of 595.07%blocks of 1695.13%halve the distance to go94.53%a quarter of the distance94.67%sixteen, then two94.60%how the blocks were sized1500 runs, requirement 0.5, nominal 95%±1.13% on each bar
Fig. 5 The same six rules at a much looser requirement, where an entire run is a handful of blocks. Drag the requirement: the exactness does not depend on there being many.

What one run’s half-widths are and are not evidence of

The three runs on one stream report half-widths of 0.2796, 0.1935 and 0.2085, and the temptation is to read the front-loading schedule as 31% better than the halving one. Decomposing the ratio says how much of that is the schedule.

The half-width is tb1SB/Nt_{b-1}\,S_B/\sqrt{N}, so the comparison is three factors. The t multiplier accounts for a factor of t₇/t₂₀ = 2.365/2.086 = 1.134. The sample size accounts for √(76/70) = 1.042 in the other direction. And the between-block spread estimate accounts for the rest — 0.9891 against 0.8087, a factor of 1.223.

On a logarithmic reading that is 34% of the difference from the t and 55% from S_B, which is an estimate of σ on seven degrees of freedom in the first run and twenty in the second. Two thirds of the gap between those two half-widths is the noise in a spread estimated from seven numbers.

So a single run cannot separate a schedule from its own luck, and the three half-widths above are an illustration of what the construction produces rather than a comparison between schedules. The part that is genuinely structural is the t multiplier, which is a deterministic function of the number of blocks and is the smaller of the two terms.

The identity says the best split is a ratio

The partition of N − 1 between the two estimates has a consequence for how many blocks a schedule should aim at, and it follows from the identity plus one fact about how each cost behaves.

Both halves are paid in the currency of a degrees-of-freedom shortfall. The interval’s premium over a known-σ width is tb1/zt_{b-1}/z, which is close to 1 + 1.3/(b − 1) for b − 1 past about ten — it gives 1.13 at ten degrees of freedom against a true 1.137, and 1.031 at forty against 1.031. The rule’s cost is whatever a spread estimated on N − b degrees of freedom does to where the run stops, and it is of the same shape: some constant over N − b.

Minimising 1.3/(b − 1) + c/(N − b) over b gives

(N − b)/(b − 1) = √(c/1.3)

with no N in it. The best partition is a fixed ratio of the two degrees of freedom, so the optimal number of blocks grows in proportion to the run’s length rather than settling at some number, and a schedule that aims at a fixed b is wrong at every length but one.

That is the strongest thing the identity supports on its own, and it stops short of a recommendation because c is not measured here — it is what the next essay’s frontier is about, and the ratio it implies is the one number a schedule needs. What can be said without it is the shape: the answer is a proportion of the run, not a count of blocks, and the front-loading schedule’s twenty-one blocks in seventy-six observations is a ratio of about 2.6 rather than a number to carry to another requirement.

What a practitioner would actually run

The construction is fiddlier to describe than to implement, and it is worth stating as a procedure so that the arithmetic above does not make it sound like a research object.

  1. Fix the half-width d and the level. Compute z.
  2. Take two starting blocks of a few observations each. Pool the within-block sums of squares into σ̂² and record the block means separately.
  3. Compute the target z²σ̂²/d². If the observations in hand reach it, stop.
  4. Otherwise ask the schedule how large the next block should be — using the target, the count, and nothing else — take that many observations, update σ̂² from their contrasts, add their mean to the list, and return to step 3.
  5. Report xˉ±tb1SB2/N\bar{x} \pm t_{b-1}\sqrt{S_B^2/N}, with S_B² the weighted between-block sum of squares over b − 1.

The only line that differs from the fixed-block version is the fourth, and the only line anybody gets wrong is the fifth: the temptation is to report the ordinary t interval on all N observations, which is right there, is narrower, and does not cover.

There is one implementation detail with a consequence. The first block must have at least two observations, or there is no contrast and no spread estimate; and there must be at least two blocks before the interval exists, or b − 1 is zero. Both are satisfied by starting with two blocks rather than one, which is why every run in this field begins that way rather than with a single warm-up block.

What the schedule is allowed to know, precisely

The condition is “a function of the contrasts”, and it is worth spelling out what that admits and what it excludes, because two things that sound identical fall on opposite sides.

Admitted: the pooled within-block estimate of σ², the number of blocks so far, the sizes already used, the number of observations in hand, the current target z²σ̂²/d², any external quantity fixed before the trial, and any randomisation the experimenter cares to perform. Every schedule in this field uses the first four.

Excluded: any block mean, the overall mean, the between-block sum of squares, the interval’s current half-width, and anything computed from them. The exclusion is not “avoid these because they are risky”; it is that the argument above conditions on the sizes, and conditioning on a function of the block means changes their distribution.

The uncomfortable case is the current target, which is admitted, and it is worth seeing why. The target is z²σ̂²/d² and σ̂² is pooled from contrasts only, so the target is a function of the contrasts and carries no information about the means. A schedule may read it freely — every schedule here does — and the stopping decision itself is a comparison of N with it, which is why the whole construction works.

The case that catches people is the half-width. It looks like a natural thing to steer on: keep going until the interval is narrow enough. It is built from the block means, it is excluded, and it is the most expensive mistake in this field.

A schedule is exact for the same reason a fixed size is. Coverage of the interval built from the block means, over 1500 runs at a requirement of 0.35. The block means are independent of every within-block contrast whatever the sizes are, so a rule that stops on the contrasts and reports the weighted mean with Σm_b(ȳ_b − x̄)²/(b − 1) behind it is exact conditional on the sizes — and the sizes may then be chosen from the contrasts, adaptively, without touching the argument. Each bar is within 2.3% of the level it claims. Nothing here is an approximation that improves with the sample size.
Fig. 6 The same six rules at a middle requirement. The exactness is a property of what the schedule reads, not of how tight the requirement is.

Why this construction rather than the obvious one

A reader who has not met the blinded rule will want to know why the interval is not simply built from all N observations in the ordinary way, which would use every degree of freedom for both purposes and make the identity above irrelevant.

Because the ordinary interval is not exact. The rule stopped when the pooled spread estimate said enough observations were in hand, so the data at the stopping time is selected on its own spread; an interval built from the same observations inherits that selection and covers below its nominal level. The field that built it measures the ordinary t interval at the same stopping times covering well below 95%, and an interval built from the estimate that stopped the rule covering far below it.

The separation is what buys exactness, and the identity is the bill. That is an unusually clean statement of a trade: the construction is exact precisely because it refuses to use each observation twice, and the cost of the refusal is that each observation counts once.

What is claimed here, and what is not

This essay takes the exactness of a varying block size, and the claims are the weighted decomposition, the counted coverage of six rules at two requirements, and the conservation identity holding on every run.

What stays out and is named as a decision: schedules that depend on anything other than the contrasts and the counts — a schedule reading an external cost, say, or the calendar — which are fine by the same argument and are not measured; non-normal observations, for which the χ² and the t are approximations and the exactness claim becomes an asymptotic one; and unequal σ across blocks, which breaks the weighting and is a different construction entirely.

The boundary against the field that built the blinded rule is the size. That the block means and the within-block contrasts are independent, that an interval built from the stopping estimate does not cover, and what the fixed-block dial costs at five settings, are established there. What is new is that the size may be a sequence, and that the sequence is spending a fixed budget.

The checks, and the refusals that make them mean something

Two claims are gated in this field’s library. Every schedule is required to cover at its nominal level at two requirements — six rules, twelve measurements, none more than three standard errors away — which is the theorem, counted. And the two degrees of freedom are required to add to N − 1 on every run of every schedule, which is checked exactly rather than on average: a version of the interval that used the wrong weights would still cover approximately and would fail this immediately.

The refusals for this field are the two in the last essay, and both are schedules that read the block means. The one that bears on this essay is the version that shrinks the block whenever the between-block spread runs above what the contrasts say — a direct attempt to hold down the quantity the interval is built from, which succeeds, and takes the coverage with it.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

BlindingBlockingChi-squareConfidence intervalCoverageDegrees of freedomFixed-width intervalIndependenceMonte CarloNuisance parameterSample sizeSequential designStopping ruleWeighted least squares