Two degrees of freedom, one total
Worth reading first: The shortest interval is the one that misses · When the looking happens.
A fixed-width procedure makes two claims and they are not the same claim.
The interval covers. The reported interval contains the mean 95% of the time — the exactness the construction was built for.
And the width was as promised. The experiment was run to make the mean known to within d, and whether it did is a question about |x̄ − μ| ≤ d, which is a statement about the sample size rather than about the interval.
The block size moves both, in opposite directions, and this essay is what each of them turns out to depend on.
The dial
At a requirement of a quarter, where an oracle who knew σ would need 61.5 observations:
| blocks of | half-width | fixed-width coverage | observations |
|---|---|---|---|
| 2 | 0.2602 | 92.40% | 59.03 |
| 3 | 0.2650 | 93.07% | 59.88 |
| 4 | 0.2647 | 93.80% | 60.73 |
| 5 | 0.2706 | 94.13% | 61.42 |
| 8 | 0.2780 | 94.53% | 63.41 |
| 16 | 0.2933 | 95.00% | 67.71 |
Small blocks give a narrower interval and a worse fixed-width claim. Large blocks give the reverse. Both differences are worth having: the width moves by 12.7% across the range and the fixed-width coverage by 2.6 points, against a counting error of 0.56.
Everything in this essay is about why each column moves.
The width is the number of blocks, once the sample size is divided out
The half-width is , and N differs between rules because the stopping time does. So the comparison to make is not the width but the width per unit of √N, which strips the sample size out and leaves whatever else is going on.
What is left is a function of the interval’s degrees of freedom and of nothing else:
| interval’s degrees of freedom | half-width × √N |
|---|---|
| 25.5 | 1.9995 |
| 17.6 | 2.0504 |
| 13.7 | 2.0627 |
| 11.3 | 2.1208 |
| 9.7 | 2.1529 |
| 7.7 | 2.2137 |
| 4.6 | 2.4132 |
That is the t multiplier and the spread of a χ² on b − 1, and it explains the width column entirely. A rule with more blocks has a smaller multiplier and a better-estimated spread, and both push the same way.
The schedules sit on that curve rather than above it. The front-loading schedule has 14.6 interval degrees of freedom and a width-per-√N of 2.0798, where a fixed size of four has 13.7 and 2.0627. They are the same point on the same curve, reached differently.
That is the first half of the field’s answer, and it is a negative one. A schedule cannot buy a narrower interval per observation. The width is set by the number of blocks, the number of blocks is half of a fixed budget, and no arrangement of when the blocks happen changes the arithmetic.
The fixed-width claim is the sample-size distribution, and there is a formula
The other column has an equally clean reduction, and it comes with a second route.
If the stopping time carries no information about the mean — which is a theorem here rather than an assumption, since the rule reads only the contrasts — then
an expectation over the sample-size distribution alone. Nothing about the interval, the blocks or the rule appears in it except through N.
Computed that way from each rule’s own realised sample sizes, the coverages are 93.13%, 93.60%, 93.85%, 94.06%, 94.58% and 95.33% against the counted 92.40%, 93.07%, 93.80%, 94.13%, 94.53% and 95.00%. Every pair agrees to within a point, and the two routes share nothing: one counts how often the mean landed inside d, the other evaluates a normal integral at the realised sample sizes and never looks at a mean.
So the second column is a statement about where N lands, and the block size affects it through two channels. Larger blocks overshoot the target more, so E[N] is larger — 67.7 against 59.0 across the range. And larger blocks give the rule’s spread estimate more degrees of freedom, so N is less variable: its standard deviation falls from 16.23 to 13.39 as the blocks go from two to eight.
The expression is concave in N, so both channels push the same way and a rule with more observations and a steadier stopping point does better on the fixed-width claim. What it costs is exactly the extra observations.
The curve is the t multiplier, shrunk by the chi-square bias
The width-per-√N column is attributed to the t multiplier and the spread of a chi-square, and dividing by the first says how much of it is the second.
Against t on the seven degree-of-freedom counts — 2.060, 2.109, 2.156, 2.201, 2.240, 2.326 and 2.706 — the counted values give ratios of 0.971, 0.972, 0.957, 0.964, 0.961, 0.952 and 0.892.
Six of the seven sit at 0.96 within a hundredth, and the seventh, at four and a half degrees of freedom, falls to 0.89. That drift is the second term: E[S_B] is below σ by about 1/(4ν), which is 1.0% at twenty-five degrees of freedom and 5.4% at four and a half — a predicted fall of 0.955 in the ratio against a counted 0.919.
So the column is times a constant times a chi-square bias, and the constant is the same for every rule. That is a stronger statement than the essay’s, which attributes the column to two effects without separating them: the t multiplier does nearly all of the work, and the spread’s own bias contributes only at the very smallest block counts, where it adds a further eight per cent.
The price list for a point of coverage
The dial’s two columns are usually read as a trade, and the observation column says it is not quite one.
Across the range the width rises 12.7%, the fixed-width coverage rises 2.60 points, and the sample size rises 14.7%. So a point of coverage costs about 4.9% of width and about 5.7% more observations — both, not one or the other.
The block-size dial does not trade the interval against the claim. It buys the claim with width and with observations at once, which is the honest description and is worse than a trade: an experimenter moving to larger blocks pays twice and a reader seeing only the coverage column sees neither payment.
That also says where a schedule’s four per cent fits. A schedule saves about four per cent of the observations and leaves the width per observation where it was, so on this price list it buys about 0.7 points of coverage-equivalent for nothing — which is small, real, and the only quantity in the field that is not being paid for out of one of the other two.
The closed form’s residuals belong in the same reading. It over-states coverage by 0.73, 0.53, 0.05, −0.07, 0.05 and 0.33 points against a counting error of 0.56, so five of the six are inside one standard error and the two largest are at the two smallest block sizes — where N is most variable and a concave function’s expectation is hardest to evaluate from a finite sample of stopping times. Suggestive, and not established on six points.
Where the closed form’s remaining point comes from
The two routes agree to within a point everywhere, and the residual is systematic rather than noise: the formula is above the count at every fixed size, by 0.73, 0.54, 0.05, −0.07, 0.05 and 0.33 points. It is worth saying what that is, because a systematic gap between two routes is either a real effect or an error in one of them.
The formula assumes the stopping time is independent of the mean. It is — the rule reads contrasts only, and the contrasts are independent of the block means. But the fixed-width event is about x̄, the mean of all N observations, and x̄ is not a function of the block means alone once N is random: the observations inside the last block are in x̄, and how many of them there are was decided by contrasts that involve those same observations.
The effect is small, it is a within-block object, and it is largest where blocks are large relative to the run — which is not quite what the residual column shows, so most of what is there is counting error at three thousand runs. The honest statement is that the two routes agree to within the resolution of the comparison and that a systematic difference of a few tenths of a point would not be detectable here.
That is the difference between this and the field’s other two-route check. The degrees of freedom adding to N − 1 is exact and is checked exactly; this one is a numerical agreement with a stated tolerance, and the tolerance is the simulation’s own.
Why the dial cannot be beaten by rearranging it
The conservation identity — (b − 1) + (N − b) = N − 1 on every run — now has a consequence that can be stated in the columns above.
The width per observation depends on b − 1. The variability of the sample size depends on N − b. They add to N − 1. So moving along the dial trades one for the other at a fixed exchange rate, and a schedule that spends its rule’s degrees of freedom early and its interval’s late is still spending the same total.
The only quantity in the whole field that the identity does not touch is the overshoot, and it is worth seeing it isolated: 1.48, 2.17, 3.11, 4.68 and 8.46 observations past the target as the fixed block size goes 2, 3, 5, 8, 16. That is about half a block, as it must be, and it is a pure loss — observations spent for nothing because the run cannot stop between blocks.
The schedules overshoot by 1.62, 1.32 and 1.37: they end in twos, so they land where twos land, having spent most of the run in blocks four and eight times larger. That is the one thing on this page a schedule takes from both ends, and the next essay is about how much it is worth.
Why not simply use more observations
The obvious objection to the whole dial is that both columns can be improved at once by tightening d and then reporting the interval that comes out, or by ignoring the rule and taking a hundred observations. Both are true and neither is the problem being solved.
A fixed-width procedure exists because observations are expensive and somebody wants the smallest number that will do. Its whole content is the stopping rule, and the stopping rule is where the σ that nobody knows enters. Every column above is a cost of not knowing σ, measured against an oracle who does and would need 61.5 observations: the best rule here spends 59.0 and the worst 67.7, and the one that spends fewest is the one whose fixed-width claim is worst.
Read that way the table is a menu rather than a ranking. An experimenter for whom observations are the binding constraint takes small blocks and accepts a fixed-width claim at 92.4%. One who has promised a width takes large blocks and pays eight observations for it. The interval covers exactly either way, which is the part the construction guarantees and the part that does not depend on the choice.
The oracle, and what the whole apparatus costs
Everything here is measured against a comparison worth stating on its own: an experimenter who knew σ would take 61.5 observations, report an interval of exactly ±d, and be right 95% of the time by construction. That is the benchmark, and the gap to it is the price of not knowing one number.
The price has three parts and they are separable.
Observations. The rules here spend between 59.0 and 67.7, so the best of them spends fewer than the oracle and the worst spends 10% more. Spending fewer is not a free lunch: a rule with small blocks has a noisy spread estimate and stops early about as often as it stops late, and the early stops are what pull the mean below the oracle’s number.
Width. The oracle’s interval is exactly d = 0.25. The rules report 0.2602 to 0.2933, so even the narrowest is 4% wider — the price of estimating σ, paid as a t multiplier on the interval’s own degrees of freedom.
And the fixed-width claim. The oracle is at 95% exactly. The rules are at 92.40% to 95.00%, and the shortfall is concavity: a random sample size around the right mean covers less than a fixed sample size at that mean.
The third of those is the one that is easiest to forget, because it is not visible in anything the experiment reports. The interval covers exactly, the width is written on the page, and the statement that has quietly failed is the one in the protocol — the promise to determine the mean to within a quarter — which nothing in the output contradicts.
What moves when the requirement moves
The dial’s setting is not universal, and the direction it shifts is worth having.
At a tighter requirement every rule has more observations, so every rule has more blocks and more within-block contrasts, and both degrees of freedom are larger. The t multiplier flattens out towards z and the width column compresses; the sample-size variability falls relative to its mean and the fixed-width column compresses too. The dial is at its least consequential where the experiment is largest, which is the usual shape for a small-sample correction.
At a looser requirement the opposite: a run may be five blocks long, the interval has four degrees of freedom, and the multiplier is 2.78 rather than 1.96. There the choice of block size is most of the design.
So the practical reading is that the block size matters in inverse proportion to how much data there is, and an experimenter who is going to take six hundred observations may take them in any blocks they like.
The inversion is not gentle. Between the tight requirement and the loose one the number of blocks a run contains changes by a factor of four, and the t multiplier that sets the width moves with the inverse of that. A rule that is a sensible compromise at sixty observations — blocks of four or five, say, giving the interval a dozen degrees of freedom and the rule fifty — becomes a rule with three blocks and two degrees of freedom at fifteen observations, where the multiplier is above four and the interval is three times wider than the requirement it was built to meet. The dial does not merely matter more in a small experiment; the settings that were reasonable in a large one are actively wrong in a small one, and there is nothing in the procedure that flags it.
What is claimed here, and what is not
This essay takes what each end of the block-size dial is a function of, and the claims are the two columns across seven fixed sizes, the reduction of the width to the interval’s degrees of freedom once √N is divided out, the closed-form route to the fixed-width coverage agreeing to within a point everywhere, and the overshoot’s dependence on the last block size.
What stays out and is named as a decision: any loss function that weighs observations against width explicitly, which would turn the menu into a single answer and would need a price nobody here has; the choice of level, which moves z and therefore the target and therefore everything, and is held at 95% throughout; and the behaviour at requirements loose enough that a run is two or three blocks, where the t multiplier is large enough that the interval is not a useful object whatever its coverage.
The boundary against the field that built the blinded rule is which claim is being discussed. That the interval covers exactly, and that the fixed-width claim’s shortfall is Jensen’s inequality on a random sample size rather than a dependence between the stopping time and the mean, are established there, together with the exactness the varying size inherits. What is new is that both are functions of one partitioned budget.
The checks, and the refusals that make them mean something
Two claims are gated in this field’s library. The honest interval is required to widen with every increase in the block size and the fixed-width claim’s coverage to rise with it by more than two points across the range — the two directions of the dial, which would fail together if the mechanism were being described backwards. And the closed-form coverage is required to agree with the counted one to within a point for every fixed size, which is the second route and shares no arithmetic with the first: one is a count of how often a mean landed inside d, the other a normal integral evaluated at sample sizes.
The refusal for this field that bears on this essay is the rule that stops when the interval it is about to report is narrow enough. It optimises the width column directly, it spends fewer observations than any rule here, and its interval does not cover — which is the shape of every shortcut across this particular trade.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- Stopping when it is precise enough — both name coverage, fixed-width interval, monte carlo, precision, sample size, sequential design, stopping rule
- The interval after a stop it chose — both name confidence interval, coverage, fixed-width interval, monte carlo, sequential design, stopping rule
- What a two-arm rule may not pool — both name confidence interval, coverage, degrees of freedom, fixed-width interval, sample size, stopping rule
- A ratio that changes between blocks — both name blocking, coverage, degrees of freedom, fixed-width interval, nuisance parameter
- The bias that lands in the slope — both name blocking, coverage, degrees of freedom, fixed-width interval, nuisance parameter
- Weights that need only a ratio — both name blocking, coverage, degrees of freedom, fixed-width interval, nuisance parameter
Named objects
A flat tag is an object no other essay names yet.
BlockingChi-squareConfidence intervalCoverageDegrees of freedomFixed-width intervalJensens inequalityMonte CarloNuisance parameterOvershootPrecisionSample sizeSequential designStopping ruleVariance