What the procedure may not read

What the blindfold costs

The exactly-covering rule pays for it in the width of the interval, and the block size is a dial between two costs that run in opposite directions. And on an interval whose width was fixed in advance, the same repair buys nothing at all.

Worth reading first: What the exactness buys · When the looking happens.

A rule that stops on the within-block contrasts and reports an interval built from the block means covers exactly its nominal level, at every block size, under any stopping rule of that form. Nothing in this subject is free, so the question is what was paid and where it went.

It went into the width. And the accounting has a dial in it, which is the more interesting half.

The price, measured

At a requirement of d = 0.4 with a first look after five observations, the median half-width of the interval each rule reports:

rule coverage median half-width observations
the purely sequential rule 91.72% 0.4169 20.8
blinded, blocks of 2 94.97% 0.4746 20.3
blinded, blocks of 3 94.80% 0.4716 22.6
blinded, blocks of 5 95.08% 0.5186 24.6
blinded, blocks of 10 95.30% 0.7174 28.3
Stein’s two stages 95.38% 0.3979 49.8

The exact rule at its best block size reports an interval 13% wider than the rule that does not cover. That is the whole of the price at this setting, and it is a fair swap: an interval 13% wider that means what it says, against one that is short precisely on the runs where it should not be.

What the exactness costs, and the dial it is bought with. The median half-width of the interval each rule reports, at a requirement of 0.4 and a first look after 5 observations. The flat line is the interval a practitioner writes at the purely sequential rule's stopping time, which covers 91.72% rather than 95%. The curve is the blinded rule, which covers its nominal level at every block size: it reads b − 1 degrees of freedom where the other reads n − 1, and pays for the exactness in width. The best block size is 3, at 0.4712. Larger blocks give the stopping rule a better estimate and the interval a worse one, and the two costs go opposite ways, which is what puts the minimum in the middle.
Fig. 1 Median half-width against block size, with the sequential rule’s interval as a flat line. The curve has a minimum, and the minimum is not at either end.

Stein’s row is the one worth staring at. It covers exactly, and its interval is narrower than anything else here — because it spends 49.8 observations to get it, more than twice what knowing σ would have needed. Exactness is bought either with observations or with width, and the two rules that have it have made opposite choices: Stein buys a narrow interval with a long experiment, the blinded rule buys a short experiment with a wide interval, and the purely sequential rule declines to buy at all and does not have the thing.

The width the oracle would have reported

There is a natural reference point for all six rows and it is worth putting in, because “13% wider than the rule that does not cover” invites the question wider than what, really.

The oracle experiment — the one run by somebody who knows σ — takes n* observations and reports a half-width of exactly 0.400, because n* is defined as the count at which zσ/√n equals the requirement. Against that:

  • the purely sequential rule reports 0.4169, 4% wide, on 20.8 observations, and covers 91.7%;
  • the blinded rule at blocks of three reports 0.4716, 18% wide, on 22.6 observations, and covers its level;
  • Stein’s rule reports 0.3979, marginally narrow, on 49.8 observations, and covers its level.

Read that way the three rules are three points on one curve and the curve is not mysterious: an interval that has to be honest about an estimated σ is wider than one computed from the truth, and how much wider depends on how many degrees of freedom are left over after the stopping rule has taken its share. Stein’s is narrow because it takes so many observations that its first-stage estimate, poor as it is, is being applied to a large sample.

Both adaptations at once, and which of them costs. 260 experiments on the two-parameter model with a true K of 3 and a guess of 1, each asked for a standard error of 8% of the estimate. Fixing the design at the guess and letting the rule stop takes 57.6 runs; redrawing the design at the current estimate after every batch takes 46.1, a saving of 20%. The third row is the control: the same adaptive design, run for a number of runs fixed in advance and rounded down to a whole batch, so it never spends more than the stopping rule does on average. It covers 94.2% against the stopping rule's 93.5%, on 46.1 − 44.0 = 2.1 fewer runs, and its interval is 6.6% wider. Choosing where the runs go is free; choosing how many is not.
Fig. 2 The same trade inside an experiment that also chooses where its runs go, from the field that measured it: precision demanded against runs spent, with the interval afterwards as the thing that pays.

The dial, and why the best block is in the middle

The block size splits the sample between the two jobs, and the split is a genuine trade rather than a tuning parameter with a monotone answer.

A larger block gives the rule a better estimate. With blocks of m, the stopping rule’s estimate of σ² has b(m − 1) degrees of freedom — half the sample at m = 2, nine tenths at m = 10 — so a large block stops closer to the right moment and the sample size is less variable. Measured: the standard deviation of the number of observations falls from 10.3 at blocks of two to 6.9 at blocks of ten.

A larger block gives the interval a worse one. The interval reads b − 1 degrees of freedom, and b is the number of blocks: at twenty observations that is nine at m = 2 and one at m = 10. The t multiplier explodes accordingly, and with it the width.

And a larger block overshoots. The rule can only stop at a block boundary, so a coarse check sails past the crossing: the observations spent rise from 20.3 at m = 2 to 28.3 at m = 10, a 39% increase, and every one of those observations is spent because the rule was not allowed to look between blocks.

Two of those three costs rise with m and one falls, so there is a best block size and it is small: 3 at this requirement and 2 at the gentler one. Which is a slightly disappointing answer to a question with a nice shape — the minimum is shallow, 0.4716 against 0.4746, and any small block will do. What the shape is really saying is that the interval’s degrees of freedom are the binding constraint, and everything that helps the rule is bought from the wrong pocket.

What the exactness costs, and the dial it is bought withThe median half-width of the interval each rule reports, at a requirement of 0.25 and a first look after 10 observations. The flat line is the interval a practitioner writes at the purely sequential rule's stopping time, which covers 94.80% rather than 95%. The curve is the blinded rule, which covers its nominal level at every block size: it reads b − 1 degrees of freedom where the other reads n − 1, and pays for the exactness in width. The best block size is 2, at 0.2626. Larger blocks give the stopping rule a better estimate and the interval a worse one, and the two costs go opposite ways, which is what puts the minimum in the middle.00.1000.2000.300235810observations per blockmedian half-width of the interval reportedthe rule that does not cover: 94.8%two stages, at 1.32× the observations2,500 runs per block size, d = 0.25narrowest at blocks of 2
Fig. 3 The same curve at a looser requirement, where the experiment is two and a half times longer. Drag it: the whole curve flattens, because with sixty observations the interval has enough degrees of freedom at any block size and only the overshoot is left.

There is a fourth cost hidden in the same dial and it is the one an experimenter would notice first. A rule that only looks every ten observations is a rule that cannot stop at 21, and in a laboratory where each observation is a day, the difference between checking every two and checking every ten is not a statistical question. The measurements above price it: eight extra observations, on an experiment of about twenty-four, for the convenience of looking five times less often. That is the same trade as the width one and it is made in the same units, which is worth knowing before it is made by accident.

Where the cost goes as the experiment grows

The width penalty is a small-sample phenomenon and it disappears at the rate a t multiplier disappears.

At d = 0.4 the experiment is about 24 observations, the interval reads nine degrees of freedom at blocks of two, and the t multiplier is 2.26 against a normal’s 1.96 — a 15% penalty, which is very nearly the whole of the 13% measured. At d = 0.25 the experiment is about 61 observations, the interval reads 27, the multiplier is 2.05, and the measured penalty is 4%: 0.2638 against the sequential rule’s 0.2536.

So the honest summary of the price is not a number but a rate. The blindfold costs a t multiplier on half the sample’s degrees of freedom, and that is 15% at twenty observations, 4% at sixty, and under 2% at two hundred. It is expensive exactly where every other adaptation is expensive and cheap everywhere else.

Where the observations go. The distribution of the number of observations each rule spends, against the 61.5 that knowing σ would have required. The sequential rule spends 59.8 on average and covers 94.8%. The blinded rule spends 57.2 and covers 94.8%. Stein's two-stage rule also covers exactly and spends 81.3, which is 1.32 times the oracle count, with a long right tail: it commits to a sample size on the strength of 10 observations and cannot revise it downwards.
Fig. 4 The three sample-size distributions at the looser requirement. The blinded rule’s is wider than the sequential rule’s, which is the second cost of the split: the rule stops on half the information, so it stops at a more variable moment.
Exact coverage, at every block size. Coverage of the interval each rule reports, at a nominal 95%, over 2,500 runs each with a standard error of 0.44 points. The blinded rule stops on the within-block contrasts and reports an interval built from the block means, and those two are independent whatever the rule does — so the interval is an ordinary t interval on b − 1 degrees of freedom and its coverage is exact. It is exact at every block size drawn. The interval a practitioner writes at the purely sequential rule's stopping time covers 91.72%, and Stein's two-stage rule is exact for the same reason as the blinded rule and spends 2.10 times the observations to be so. The bars are truncated at 86% so the differences can be seen.
Fig. 5 The coverages that go with those widths. Every row in the width table above has to be read against this one, because a narrow interval that covers 91.7% is not a better interval — it is a different promise.

And the interval whose width was fixed in advance

There is a second interval in this problem, and the accounting for it is completely different. It is worth a section because it is the sharpest correction this field makes to its own predecessor’s framing.

The fixed-width problem asks for an interval x̄ ± d with d decided before the experiment starts, and the rule’s job is to take enough observations that such an interval covers. Nothing about the interval is estimated, so the only thing that decides coverage is how many observations there were:

coverage=E[2Φ ⁣(dNσ)1],\text{coverage} = \mathbb{E}\left[2\Phi\!\left(\frac{d\sqrt{N}}{\sigma}\right) - 1\right],

which holds exactly whenever N carries no information about the mean.

That expression can be computed from each rule’s own distribution of sample sizes, without simulating a single interval, and compared with the coverage counted directly. If the dependence between a rule’s stopping time and its mean matters, the two must disagree.

A fixed-width interval knows only how many observations it got. Coverage of the interval x̄ ± 0.4, counted, with the mark on each bar showing what E[2Φ(d√N/σ) − 1] gives for that rule's own distribution of sample sizes — the coverage it would have if the stopping time carried no information about the mean at all. The two agree to within 0.38 of a point for every rule here, including the sequential one. So the shortfall of a fixed-width sequential rule is not its dependence: it is the concavity of that expression in a random N, and blinding the rule makes it worse rather than better by making N more variable. A repair has to be matched to the defect it is for.
Fig. 6 Counted coverage of a fixed-width interval, with the closed form computed from each rule’s own sample-size distribution marked on the bar. Five rules, and the two routes agree on all of them.

They do not disagree. Counted against computed: 90.05% and 89.70% for the purely sequential rule, 95.90% and 95.48% for Stein’s, 87.85% and 88.14% for the blinded rule at blocks of two, 93.52% and 93.43% at blocks of five, 96.00% and 95.92% at blocks of ten. Every gap is under half a point, which is the Monte Carlo noise of the comparison.

So for an interval whose width was fixed in advance, the dependence is worth nothing and the shape of the sample-size distribution is worth everything. The sequential rule’s 90% is not a symptom of its stopping time knowing about its mean: it is Jensen’s inequality. Coverage is a concave function of N near the target, so a random N — however honestly obtained — averages to less coverage than the same expected N delivered exactly.

And then the sting. Blinding the rule makes that worse: at blocks of two the fixed-width coverage falls from 90.05% to 87.85%, because the rule now stops on half the degrees of freedom and its N is more variable. At blocks of ten it rises to 96.00%, for the same reason in reverse — a better estimate, a tighter N, and an overshoot that pushes the average count above the oracle’s.

The repair has to be matched to the defect. The same construction that makes an estimated-width interval exact makes a fixed-width interval worse, and the two facts are not in tension: they are about different intervals with different failure modes, and the only way to tell them apart was to compute the second route and see that it left nothing over.

A fixed-width interval knows only how many observations it got. Coverage of the interval x̄ ± 0.2, counted, with the mark on each bar showing what E[2Φ(d√N/σ) − 1] gives for that rule's own distribution of sample sizes — the coverage it would have if the stopping time carried no information about the mean at all. The two agree to within 0.61 of a point for every rule here, including the sequential one. So the shortfall of a fixed-width sequential rule is not its dependence: it is the concavity of that expression in a random N, and blinding the rule makes it worse rather than better by making N more variable. A repair has to be matched to the defect it is for.
Fig. 7 The same decomposition at a tighter requirement. Drag it: the agreement between the count and the closed form survives every setting, which is what makes it a check rather than a coincidence at one point.

One more consequence is worth drawing out, because it is the kind of thing that gets asserted rather than measured. It is often said that a sequential rule “spends” its optional stopping and that the price is a wider interval. On the fixed-width problem that sentence has no content at all: the interval’s width is a constant, the rule cannot make it wider or narrower, and everything that happens to coverage happens through the count. The price of optional stopping is paid in a different currency depending on which of the two intervals is being reported — width in one case, coverage in the other — and no single sentence covers both.

The fixed-width coverage, split into a mean and a variance

The closed form E[2Φ(d√N/σ) − 1] is used above as a check and it will also decompose, which says how much of each rule’s shortfall is stopping too early and how much is stopping at a variable time.

Expanding it about the mean sample size gives g(E[N]) + ½·g″(E[N])·Var(N), and both terms are computable from the two numbers each rule already reports. For the blinded rule at blocks of two, with E[N] = 20.3 and a standard deviation of 10.3, the first term is 92.9% and the second is −3.9 points — a predicted 89.0% against a counted 87.85%, which is as close as a two-term expansion has any right to be at that much variance.

So of the seven points the blinded rule’s fixed-width interval gives up against its nominal level, about two are the mean sample size falling short of the oracle’s twenty-four, and about four are the spread of N. The variability is the larger term, and it is the one no amount of recentring the rule would remove.

That also prices the sting at the end of the essay. Blinding raises the standard deviation of N from the sequential rule’s 8.5 to 10.3, so it raises Var(N) by 33.85 — worth ½·g″·33.85 = 1.2 points of coverage — and it lowers the mean count from 20.8 to 20.3, worth a further 0.3. Predicted drop 1.5 points, counted 2.2. The mechanism is the whole of it to within the accuracy of a quadratic, and it is entirely the second moment of a stopping time.

The expansion also says where it stops working. Stein’s rule has a standard deviation of 32.7 on a mean of 49.8, so the quadratic term is worth two points and the higher ones are not negligible: the expansion predicts 97.5% against a counted 95.90%. A rule whose sample size is a scaled chi-square on four degrees of freedom is not well described by two moments, which is the same statement the two-stage rule’s 5th-to-95th range of 9 to 111 makes in a different currency.

Read together, the two halves give the fixed-width problem a complete accounting with no simulation in it. Coverage is a concave function of N; a rule is judged by where it puts E[N] and how tightly; and the whole of the difference between five rules is two numbers each, which is why the closed form agrees with every one of them and why the dependence the estimated-width interval is destroyed by contributes nothing here at all.

What a practitioner should take from the pair

Three rules, and the choice between them is not a matter of taste once the question is stated precisely.

If the interval’s width is to be estimated at the end — the ordinary case, where somebody reports an estimate and a standard error — the blinded rule is the one that covers, and it costs a t multiplier on half the degrees of freedom. Nothing else on offer here is exact at a comparable number of observations.

If the width is fixed in advance and the only question is how many observations to buy, then nothing about blinding helps, the variability of N is the enemy, and the two-stage rule’s answer — overshoot deliberately, accept 2.07 times the oracle count — is buying exactly the right thing.

If the rule is being written into a protocol somebody else will follow, the block construction has a property the other two do not: it is impossible to get wrong by looking too often. A protocol that says “compute the interval from the block means” cannot be violated by an over-eager analyst who peeks, because peeking at contrasts is exactly what the rule is made of. The two-stage rule, by contrast, is exact only if nobody re-estimates the sample size a second time, which is a discipline rather than a theorem.

And if the experiment is large, all of this is a second-order argument about a few per cent, and the reason to prefer the exact rule is not efficiency but that its guarantee is a theorem rather than an asymptotic hope.

Where the observations go. The distribution of the number of observations each rule spends, against the 24.0 that knowing σ would have required. The sequential rule spends 20.6 on average and covers 91.7%. The blinded rule spends 20.3 and covers 95.0%. Stein's two-stage rule also covers exactly and spends 50.3, which is 2.10 times the oracle count, with a long right tail: it commits to a sample size on the strength of 5 observations and cannot revise it downwards.
Fig. 8 And where the observations go at the tighter requirement, which is where the three rules differ most. The long right tail belongs to the rule that fixed its sample size on five observations.

What is claimed here, and what is not

This essay takes the cost of the exactly-covering rule, and the claims are two: the price is a t multiplier on the interval’s own degrees of freedom, with a best block size in the middle; and a fixed-width interval’s coverage is a function of the sample-size distribution alone, so the same repair does not apply to it.

What stays out and is named as a decision: block sizes chosen adaptively — larger blocks early when the estimate is poor and smaller ones near the crossing, which would buy both ends of the trade and is a rule this essay does not build; non-normal data, where the second route above stops being a closed form; and any attempt to optimise the split formally, since the objective would be an expected width over a random stopping time and the answer would be a number at one setting rather than the rate the measurements give.

The boundary against the previous field is the diagnosis. That the sequential rule’s spread estimate at stopping is 0.83 of the truth, and that its estimated-width interval is short because of it, is established there; what this essay adds is that its fixed-width interval is short for an unrelated reason, and that the two were being explained by the same sentence.

The checks, and the refusals that make them mean something

Three claims are gated in this field’s library. The exactness comparison is required to hold in the form that makes it a trade: the two-stage rule at more than 1.6 times the oracle count and the blinded rule at under 1.05, with the blinded interval wider than the interval that does not cover, and wider still at large blocks — four statements about one comparison, any of which failing would mean the trade had been described wrongly. The observations spent are required to rise monotonically with the block size, which is the overshoot. And the fixed-width decomposition is required to agree with its closed form for every rule, to within four standard errors, with the blinded rule at blocks of two required to cover worse than the sequential rule — the one measurement in this field that says a repair made something worse, stated so that it would fail if it stopped being true.

The refusals for this half belong to the essay that built the construction, and one of them is worth naming again here because it is the cheapest possible version of this essay’s mistake: an interval built from the estimate that stopped the rule covers 85.1%. Everything above is an argument about a few percentage points of width. That is ten points of coverage, available to anybody who reaches for the number in hand at the moment the rule fires.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

BlindingClosed formCoverageDegrees of freedomEfficiencyFixed-width intervalInterval widthJensens inequalityNuisance parameterSample sizeSequential analysisStein two-stageStopping ruleT interval