The block size as a schedule

Two degrees of freedom, one total

The block size is a dial, and the two things a fixed-width procedure claims move in opposite directions along it. Divide the width by the square root of the sample size and one of them turns out to depend on the number of blocks and on nothing else.

Worth reading first: The shortest interval is the one that misses · When the looking happens.

A fixed-width procedure makes two claims and they are not the same claim.

The interval covers. The reported interval contains the mean 95% of the time — the exactness the construction was built for.

And the width was as promised. The experiment was run to make the mean known to within d, and whether it did is a question about |x̄ − μ| ≤ d, which is a statement about the sample size rather than about the interval.

The block size moves both, in opposite directions, and this essay is what each of them turns out to depend on.

No block size is best at both things the procedure claims. Two claims and one dial. The honest interval's half-width falls as the blocks get smaller, because the interval's degrees of freedom are the number of blocks: 0.2602 at blocks of two against 0.2933 at blocks of sixteen. The fixed-width claim — that the mean is within 0.25 of the truth — gets more reliable as they get larger, because the sample size is less variable: 92.40% against 95.00%. Both are computed from the same runs, and the second is reproduced to within a tenth of a point by E[2Φ(d√N/σ) − 1], which needs the sample-size distribution and nothing else. The schedules sit at the bottom left: as narrow as the smallest fixed block and as few observations, with the rule's spread estimate on half as many degrees of freedom again.
Fig. 1 The two claims against each other, with each fixed block size labelled. No point on the curve is best at both, and the schedules sit at its lower-left end.

The dial

At a requirement of a quarter, where an oracle who knew σ would need 61.5 observations:

blocks of half-width fixed-width coverage observations
2 0.2602 92.40% 59.03
3 0.2650 93.07% 59.88
4 0.2647 93.80% 60.73
5 0.2706 94.13% 61.42
8 0.2780 94.53% 63.41
16 0.2933 95.00% 67.71

Small blocks give a narrower interval and a worse fixed-width claim. Large blocks give the reverse. Both differences are worth having: the width moves by 12.7% across the range and the fixed-width coverage by 2.6 points, against a counting error of 0.56.

Everything in this essay is about why each column moves.

The width is the number of blocks, once the sample size is divided out

The half-width is tb1SB2/Nt_{b-1}\sqrt{S_B^2/N}, and N differs between rules because the stopping time does. So the comparison to make is not the width but the width per unit of √N, which strips the sample size out and leaves whatever else is going on.

What is left is a function of the interval’s degrees of freedom and of nothing else:

interval’s degrees of freedom half-width × √N
25.5 1.9995
17.6 2.0504
13.7 2.0627
11.3 2.1208
9.7 2.1529
7.7 2.2137
4.6 2.4132

That is the t multiplier and the spread of a χ² on b − 1, and it explains the width column entirely. A rule with more blocks has a smaller multiplier and a better-estimated spread, and both push the same way.

The width, once the sample size is divided out, is the number of blocks. Multiply each rule's half-width by the square root of the observations it spent and everything else falls away: what is left is a t multiplier on b − 1 degrees of freedom and the spread of a χ² on the same. The schedules sit on the same curve as the fixed sizes at the same number of blocks — halve the distance to go at 6.2, a quarter of the distance to go at 11.2, sixteen, then two at 14.6 — which is what the conservation identity requires and is why a schedule cannot buy a narrower interval for free. Everything a schedule does to the width, it does by changing how many blocks there were.
Fig. 2 Every rule, with the sample size divided out. The schedules land on the same curve as the fixed sizes at the same number of blocks, which is what the conservation identity requires.

The schedules sit on that curve rather than above it. The front-loading schedule has 14.6 interval degrees of freedom and a width-per-√N of 2.0798, where a fixed size of four has 13.7 and 2.0627. They are the same point on the same curve, reached differently.

That is the first half of the field’s answer, and it is a negative one. A schedule cannot buy a narrower interval per observation. The width is set by the number of blocks, the number of blocks is half of a fixed budget, and no arrangement of when the blocks happen changes the arithmetic.

The fixed-width claim is the sample-size distribution, and there is a formula

The other column has an equally clean reduction, and it comes with a second route.

If the stopping time carries no information about the mean — which is a theorem here rather than an assumption, since the rule reads only the contrasts — then

Pr(xˉμd)=E ⁣[2Φ ⁣(dNσ)1],\Pr(|\bar x - \mu| \le d) = \mathbb{E}\!\left[\,2\Phi\!\left(\frac{d\sqrt{N}}{\sigma}\right) - 1\,\right],

an expectation over the sample-size distribution alone. Nothing about the interval, the blocks or the rule appears in it except through N.

Computed that way from each rule’s own realised sample sizes, the coverages are 93.13%, 93.60%, 93.85%, 94.06%, 94.58% and 95.33% against the counted 92.40%, 93.07%, 93.80%, 94.13%, 94.53% and 95.00%. Every pair agrees to within a point, and the two routes share nothing: one counts how often the mean landed inside d, the other evaluates a normal integral at the realised sample sizes and never looks at a mean.

So the second column is a statement about where N lands, and the block size affects it through two channels. Larger blocks overshoot the target more, so E[N] is larger — 67.7 against 59.0 across the range. And larger blocks give the rule’s spread estimate more degrees of freedom, so N is less variable: its standard deviation falls from 16.23 to 13.39 as the blocks go from two to eight.

The expression is concave in N, so both channels push the same way and a rule with more observations and a steadier stopping point does better on the fixed-width claim. What it costs is exactly the extra observations.

The overshoot is the last block size and nothing else. A run stops at a multiple of its own block sizes and cannot land between them, so it ends past its own target by about half a block. Fixed sizes overshoot by 1.5, 2.2, 3.1, 4.7, 8.5 observations as the size goes 2, 3, 5, 8, 16. Every schedule here ends in blocks of two and every one of them lands where blocks of two land — 1.62, 1.32, 1.37 against 1.48 — while having spent most of the run inside blocks four and eight times larger. That is the one thing on this page a schedule genuinely takes from both ends.
Fig. 3 The overshoot: how far past its own target each rule lands. It is set by the last block size and it is the whole of the difference in E[N].

The curve is the t multiplier, shrunk by the chi-square bias

The width-per-√N column is attributed to the t multiplier and the spread of a chi-square, and dividing by the first says how much of it is the second.

Against t on the seven degree-of-freedom counts — 2.060, 2.109, 2.156, 2.201, 2.240, 2.326 and 2.706 — the counted values give ratios of 0.971, 0.972, 0.957, 0.964, 0.961, 0.952 and 0.892.

Six of the seven sit at 0.96 within a hundredth, and the seventh, at four and a half degrees of freedom, falls to 0.89. That drift is the second term: E[S_B] is below σ by about 1/(4ν), which is 1.0% at twenty-five degrees of freedom and 5.4% at four and a half — a predicted fall of 0.955 in the ratio against a counted 0.919.

So the column is tb1t_{b-1} times a constant times a chi-square bias, and the constant is the same for every rule. That is a stronger statement than the essay’s, which attributes the column to two effects without separating them: the t multiplier does nearly all of the work, and the spread’s own bias contributes only at the very smallest block counts, where it adds a further eight per cent.

The price list for a point of coverage

The dial’s two columns are usually read as a trade, and the observation column says it is not quite one.

Across the range the width rises 12.7%, the fixed-width coverage rises 2.60 points, and the sample size rises 14.7%. So a point of coverage costs about 4.9% of width and about 5.7% more observations — both, not one or the other.

The block-size dial does not trade the interval against the claim. It buys the claim with width and with observations at once, which is the honest description and is worse than a trade: an experimenter moving to larger blocks pays twice and a reader seeing only the coverage column sees neither payment.

That also says where a schedule’s four per cent fits. A schedule saves about four per cent of the observations and leaves the width per observation where it was, so on this price list it buys about 0.7 points of coverage-equivalent for nothing — which is small, real, and the only quantity in the field that is not being paid for out of one of the other two.

The closed form’s residuals belong in the same reading. It over-states coverage by 0.73, 0.53, 0.05, −0.07, 0.05 and 0.33 points against a counting error of 0.56, so five of the six are inside one standard error and the two largest are at the two smallest block sizes — where N is most variable and a concave function’s expectation is hardest to evaluate from a finite sample of stopping times. Suggestive, and not established on six points.

Where the closed form’s remaining point comes from

The two routes agree to within a point everywhere, and the residual is systematic rather than noise: the formula is above the count at every fixed size, by 0.73, 0.54, 0.05, −0.07, 0.05 and 0.33 points. It is worth saying what that is, because a systematic gap between two routes is either a real effect or an error in one of them.

The formula assumes the stopping time is independent of the mean. It is — the rule reads contrasts only, and the contrasts are independent of the block means. But the fixed-width event is about x̄, the mean of all N observations, and x̄ is not a function of the block means alone once N is random: the observations inside the last block are in x̄, and how many of them there are was decided by contrasts that involve those same observations.

The effect is small, it is a within-block object, and it is largest where blocks are large relative to the run — which is not quite what the residual column shows, so most of what is there is counting error at three thousand runs. The honest statement is that the two routes agree to within the resolution of the comparison and that a systematic difference of a few tenths of a point would not be detectable here.

That is the difference between this and the field’s other two-route check. The degrees of freedom adding to N − 1 is exact and is checked exactly; this one is a numerical agreement with a stated tolerance, and the tolerance is the simulation’s own.

The overshoot is the last block size and nothing else. A run stops at a multiple of its own block sizes and cannot land between them, so it ends past its own target by about half a block. Fixed sizes overshoot by 1.5, 2.1, 3.2, 4.5, 8.4 observations as the size goes 2, 3, 5, 8, 16. Every schedule here ends in blocks of two and every one of them lands where blocks of two land — 1.58, 1.48, 1.46 against 1.51 — while having spent most of the run inside blocks four and eight times larger. That is the one thing on this page a schedule genuinely takes from both ends.
Fig. 4 The overshoot at a looser requirement, where a block of sixteen is a quarter of the whole run and the loss is proportionally much larger.

Why the dial cannot be beaten by rearranging it

The conservation identity — (b − 1) + (N − b) = N − 1 on every run — now has a consequence that can be stated in the columns above.

The width per observation depends on b − 1. The variability of the sample size depends on N − b. They add to N − 1. So moving along the dial trades one for the other at a fixed exchange rate, and a schedule that spends its rule’s degrees of freedom early and its interval’s late is still spending the same total.

The only quantity in the whole field that the identity does not touch is the overshoot, and it is worth seeing it isolated: 1.48, 2.17, 3.11, 4.68 and 8.46 observations past the target as the fixed block size goes 2, 3, 5, 8, 16. That is about half a block, as it must be, and it is a pure loss — observations spent for nothing because the run cannot stop between blocks.

The schedules overshoot by 1.62, 1.32 and 1.37: they end in twos, so they land where twos land, having spent most of the run in blocks four and eight times larger. That is the one thing on this page a schedule takes from both ends, and the next essay is about how much it is worth.

Why not simply use more observations

The obvious objection to the whole dial is that both columns can be improved at once by tightening d and then reporting the interval that comes out, or by ignoring the rule and taking a hundred observations. Both are true and neither is the problem being solved.

A fixed-width procedure exists because observations are expensive and somebody wants the smallest number that will do. Its whole content is the stopping rule, and the stopping rule is where the σ that nobody knows enters. Every column above is a cost of not knowing σ, measured against an oracle who does and would need 61.5 observations: the best rule here spends 59.0 and the worst 67.7, and the one that spends fewest is the one whose fixed-width claim is worst.

Read that way the table is a menu rather than a ranking. An experimenter for whom observations are the binding constraint takes small blocks and accepts a fixed-width claim at 92.4%. One who has promised a width takes large blocks and pays eight observations for it. The interval covers exactly either way, which is the part the construction guarantees and the part that does not depend on the choice.

No block size is best at both things the procedure claims. Two claims and one dial. The honest interval's half-width falls as the blocks get smaller, because the interval's degrees of freedom are the number of blocks: 0.3838 at blocks of two against 0.4685 at blocks of sixteen. The fixed-width claim — that the mean is within 0.35 of the truth — gets more reliable as they get larger, because the sample size is less variable: 92.60% against 95.20%. Both are computed from the same runs, and the second is reproduced to within a tenth of a point by E[2Φ(d√N/σ) − 1], which needs the sample-size distribution and nothing else. The schedules sit at the bottom left: as narrow as the smallest fixed block and as few observations, with the rule's spread estimate on half as many degrees of freedom again.
Fig. 5 The same trade at a looser requirement, where runs are shorter and every rule has fewer blocks. The curve moves and its shape does not.

The oracle, and what the whole apparatus costs

Everything here is measured against a comparison worth stating on its own: an experimenter who knew σ would take 61.5 observations, report an interval of exactly ±d, and be right 95% of the time by construction. That is the benchmark, and the gap to it is the price of not knowing one number.

The price has three parts and they are separable.

Observations. The rules here spend between 59.0 and 67.7, so the best of them spends fewer than the oracle and the worst spends 10% more. Spending fewer is not a free lunch: a rule with small blocks has a noisy spread estimate and stops early about as often as it stops late, and the early stops are what pull the mean below the oracle’s number.

Width. The oracle’s interval is exactly d = 0.25. The rules report 0.2602 to 0.2933, so even the narrowest is 4% wider — the price of estimating σ, paid as a t multiplier on the interval’s own degrees of freedom.

And the fixed-width claim. The oracle is at 95% exactly. The rules are at 92.40% to 95.00%, and the shortfall is concavity: a random sample size around the right mean covers less than a fixed sample size at that mean.

The third of those is the one that is easiest to forget, because it is not visible in anything the experiment reports. The interval covers exactly, the width is written on the page, and the statement that has quietly failed is the one in the protocol — the promise to determine the mean to within a quarter — which nothing in the output contradicts.

A schedule is exact for the same reason a fixed size is. Coverage of the interval built from the block means, over 1500 runs at a requirement of 0.25. The block means are independent of every within-block contrast whatever the sizes are, so a rule that stops on the contrasts and reports the weighted mean with Σm_b(ȳ_b − x̄)²/(b − 1) behind it is exact conditional on the sizes — and the sizes may then be chosen from the contrasts, adaptively, without touching the argument. Each bar is within 2.3% of the level it claims. Nothing here is an approximation that improves with the sample size.
Fig. 6 The claim that does hold, at every setting of the dial, which is the one the construction was built to guarantee.

What moves when the requirement moves

The dial’s setting is not universal, and the direction it shifts is worth having.

At a tighter requirement every rule has more observations, so every rule has more blocks and more within-block contrasts, and both degrees of freedom are larger. The t multiplier flattens out towards z and the width column compresses; the sample-size variability falls relative to its mean and the fixed-width column compresses too. The dial is at its least consequential where the experiment is largest, which is the usual shape for a small-sample correction.

At a looser requirement the opposite: a run may be five blocks long, the interval has four degrees of freedom, and the multiplier is 2.78 rather than 1.96. There the choice of block size is most of the design.

So the practical reading is that the block size matters in inverse proportion to how much data there is, and an experimenter who is going to take six hundred observations may take them in any blocks they like.

The inversion is not gentle. Between the tight requirement and the loose one the number of blocks a run contains changes by a factor of four, and the t multiplier that sets the width moves with the inverse of that. A rule that is a sensible compromise at sixty observations — blocks of four or five, say, giving the interval a dozen degrees of freedom and the rule fifty — becomes a rule with three blocks and two degrees of freedom at fifteen observations, where the multiplier is above four and the interval is three times wider than the requirement it was built to meet. The dial does not merely matter more in a small experiment; the settings that were reasonable in a large one are actively wrong in a small one, and there is nothing in the procedure that flags it.

What is claimed here, and what is not

This essay takes what each end of the block-size dial is a function of, and the claims are the two columns across seven fixed sizes, the reduction of the width to the interval’s degrees of freedom once √N is divided out, the closed-form route to the fixed-width coverage agreeing to within a point everywhere, and the overshoot’s dependence on the last block size.

What stays out and is named as a decision: any loss function that weighs observations against width explicitly, which would turn the menu into a single answer and would need a price nobody here has; the choice of level, which moves z and therefore the target and therefore everything, and is held at 95% throughout; and the behaviour at requirements loose enough that a run is two or three blocks, where the t multiplier is large enough that the interval is not a useful object whatever its coverage.

The boundary against the field that built the blinded rule is which claim is being discussed. That the interval covers exactly, and that the fixed-width claim’s shortfall is Jensen’s inequality on a random sample size rather than a dependence between the stopping time and the mean, are established there, together with the exactness the varying size inherits. What is new is that both are functions of one partitioned budget.

The checks, and the refusals that make them mean something

Two claims are gated in this field’s library. The honest interval is required to widen with every increase in the block size and the fixed-width claim’s coverage to rise with it by more than two points across the range — the two directions of the dial, which would fail together if the mechanism were being described backwards. And the closed-form coverage is required to agree with the counted one to within a point for every fixed size, which is the second route and shares no arithmetic with the first: one is a count of how often a mean landed inside d, the other a normal integral evaluated at sample sizes.

The refusal for this field that bears on this essay is the rule that stops when the interval it is about to report is narrow enough. It optimises the width column directly, it spends fewer observations than any rule here, and its interval does not cover — which is the shape of every shortcut across this particular trade.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

BlockingChi-squareConfidence intervalCoverageDegrees of freedomFixed-width intervalJensens inequalityMonte CarloNuisance parameterOvershootPrecisionSample sizeSequential designStopping ruleVariance