What the blindfold costs
Worth reading first: What the exactness buys · When the looking happens.
A rule that stops on the within-block contrasts and reports an interval built from the block means covers exactly its nominal level, at every block size, under any stopping rule of that form. Nothing in this subject is free, so the question is what was paid and where it went.
It went into the width. And the accounting has a dial in it, which is the more interesting half.
The price, measured
At a requirement of d = 0.4 with a first look after five observations, the median half-width of the interval each rule reports:
| rule | coverage | median half-width | observations |
|---|---|---|---|
| the purely sequential rule | 91.72% | 0.4169 | 20.8 |
| blinded, blocks of 2 | 94.97% | 0.4746 | 20.3 |
| blinded, blocks of 3 | 94.80% | 0.4716 | 22.6 |
| blinded, blocks of 5 | 95.08% | 0.5186 | 24.6 |
| blinded, blocks of 10 | 95.30% | 0.7174 | 28.3 |
| Stein’s two stages | 95.38% | 0.3979 | 49.8 |
The exact rule at its best block size reports an interval 13% wider than the rule that does not cover. That is the whole of the price at this setting, and it is a fair swap: an interval 13% wider that means what it says, against one that is short precisely on the runs where it should not be.
Stein’s row is the one worth staring at. It covers exactly, and its interval is narrower than anything else here — because it spends 49.8 observations to get it, more than twice what knowing σ would have needed. Exactness is bought either with observations or with width, and the two rules that have it have made opposite choices: Stein buys a narrow interval with a long experiment, the blinded rule buys a short experiment with a wide interval, and the purely sequential rule declines to buy at all and does not have the thing.
The width the oracle would have reported
There is a natural reference point for all six rows and it is worth putting in, because “13% wider than the rule that does not cover” invites the question wider than what, really.
The oracle experiment — the one run by somebody who knows σ — takes n* observations and reports a half-width of exactly 0.400, because n* is defined as the count at which zσ/√n equals the requirement. Against that:
- the purely sequential rule reports 0.4169, 4% wide, on 20.8 observations, and covers 91.7%;
- the blinded rule at blocks of three reports 0.4716, 18% wide, on 22.6 observations, and covers its level;
- Stein’s rule reports 0.3979, marginally narrow, on 49.8 observations, and covers its level.
Read that way the three rules are three points on one curve and the curve is not mysterious: an interval that has to be honest about an estimated σ is wider than one computed from the truth, and how much wider depends on how many degrees of freedom are left over after the stopping rule has taken its share. Stein’s is narrow because it takes so many observations that its first-stage estimate, poor as it is, is being applied to a large sample.
The dial, and why the best block is in the middle
The block size splits the sample between the two jobs, and the split is a genuine trade rather than a tuning parameter with a monotone answer.
A larger block gives the rule a better estimate. With blocks of m, the stopping rule’s estimate of σ² has b(m − 1) degrees of freedom — half the sample at m = 2, nine tenths at m = 10 — so a large block stops closer to the right moment and the sample size is less variable. Measured: the standard deviation of the number of observations falls from 10.3 at blocks of two to 6.9 at blocks of ten.
A larger block gives the interval a worse one. The interval reads b − 1 degrees of freedom, and b is the number of blocks: at twenty observations that is nine at m = 2 and one at m = 10. The t multiplier explodes accordingly, and with it the width.
And a larger block overshoots. The rule can only stop at a block boundary, so a coarse check sails past the crossing: the observations spent rise from 20.3 at m = 2 to 28.3 at m = 10, a 39% increase, and every one of those observations is spent because the rule was not allowed to look between blocks.
Two of those three costs rise with m and one falls, so there is a best block size and it is small: 3 at this requirement and 2 at the gentler one. Which is a slightly disappointing answer to a question with a nice shape — the minimum is shallow, 0.4716 against 0.4746, and any small block will do. What the shape is really saying is that the interval’s degrees of freedom are the binding constraint, and everything that helps the rule is bought from the wrong pocket.
There is a fourth cost hidden in the same dial and it is the one an experimenter would notice first. A rule that only looks every ten observations is a rule that cannot stop at 21, and in a laboratory where each observation is a day, the difference between checking every two and checking every ten is not a statistical question. The measurements above price it: eight extra observations, on an experiment of about twenty-four, for the convenience of looking five times less often. That is the same trade as the width one and it is made in the same units, which is worth knowing before it is made by accident.
Where the cost goes as the experiment grows
The width penalty is a small-sample phenomenon and it disappears at the rate a t multiplier disappears.
At d = 0.4 the experiment is about 24 observations, the interval reads nine degrees of freedom at blocks of two, and the t multiplier is 2.26 against a normal’s 1.96 — a 15% penalty, which is very nearly the whole of the 13% measured. At d = 0.25 the experiment is about 61 observations, the interval reads 27, the multiplier is 2.05, and the measured penalty is 4%: 0.2638 against the sequential rule’s 0.2536.
So the honest summary of the price is not a number but a rate. The blindfold costs a t multiplier on half the sample’s degrees of freedom, and that is 15% at twenty observations, 4% at sixty, and under 2% at two hundred. It is expensive exactly where every other adaptation is expensive and cheap everywhere else.
And the interval whose width was fixed in advance
There is a second interval in this problem, and the accounting for it is completely different. It is worth a section because it is the sharpest correction this field makes to its own predecessor’s framing.
The fixed-width problem asks for an interval x̄ ± d with d decided before the experiment starts, and the rule’s job is to take enough observations that such an interval covers. Nothing about the interval is estimated, so the only thing that decides coverage is how many observations there were:
which holds exactly whenever N carries no information about the mean.
That expression can be computed from each rule’s own distribution of sample sizes, without simulating a single interval, and compared with the coverage counted directly. If the dependence between a rule’s stopping time and its mean matters, the two must disagree.
They do not disagree. Counted against computed: 90.05% and 89.70% for the purely sequential rule, 95.90% and 95.48% for Stein’s, 87.85% and 88.14% for the blinded rule at blocks of two, 93.52% and 93.43% at blocks of five, 96.00% and 95.92% at blocks of ten. Every gap is under half a point, which is the Monte Carlo noise of the comparison.
So for an interval whose width was fixed in advance, the dependence is worth nothing and the shape of the sample-size distribution is worth everything. The sequential rule’s 90% is not a symptom of its stopping time knowing about its mean: it is Jensen’s inequality. Coverage is a concave function of N near the target, so a random N — however honestly obtained — averages to less coverage than the same expected N delivered exactly.
And then the sting. Blinding the rule makes that worse: at blocks of two the fixed-width coverage falls from 90.05% to 87.85%, because the rule now stops on half the degrees of freedom and its N is more variable. At blocks of ten it rises to 96.00%, for the same reason in reverse — a better estimate, a tighter N, and an overshoot that pushes the average count above the oracle’s.
The repair has to be matched to the defect. The same construction that makes an estimated-width interval exact makes a fixed-width interval worse, and the two facts are not in tension: they are about different intervals with different failure modes, and the only way to tell them apart was to compute the second route and see that it left nothing over.
One more consequence is worth drawing out, because it is the kind of thing that gets asserted rather than measured. It is often said that a sequential rule “spends” its optional stopping and that the price is a wider interval. On the fixed-width problem that sentence has no content at all: the interval’s width is a constant, the rule cannot make it wider or narrower, and everything that happens to coverage happens through the count. The price of optional stopping is paid in a different currency depending on which of the two intervals is being reported — width in one case, coverage in the other — and no single sentence covers both.
The fixed-width coverage, split into a mean and a variance
The closed form E[2Φ(d√N/σ) − 1] is used above as a check and it will also decompose, which says how much of each rule’s shortfall is stopping too early and how much is stopping at a variable time.
Expanding it about the mean sample size gives g(E[N]) + ½·g″(E[N])·Var(N), and both terms are computable from the two numbers each rule already reports. For the blinded rule at blocks of two, with E[N] = 20.3 and a standard deviation of 10.3, the first term is 92.9% and the second is −3.9 points — a predicted 89.0% against a counted 87.85%, which is as close as a two-term expansion has any right to be at that much variance.
So of the seven points the blinded rule’s fixed-width interval gives up against its nominal level, about two are the mean sample size falling short of the oracle’s twenty-four, and about four are the spread of N. The variability is the larger term, and it is the one no amount of recentring the rule would remove.
That also prices the sting at the end of the essay. Blinding raises the standard deviation of N from the sequential rule’s 8.5 to 10.3, so it raises Var(N) by 33.85 — worth ½·g″·33.85 = 1.2 points of coverage — and it lowers the mean count from 20.8 to 20.3, worth a further 0.3. Predicted drop 1.5 points, counted 2.2. The mechanism is the whole of it to within the accuracy of a quadratic, and it is entirely the second moment of a stopping time.
The expansion also says where it stops working. Stein’s rule has a standard deviation of 32.7 on a mean of 49.8, so the quadratic term is worth two points and the higher ones are not negligible: the expansion predicts 97.5% against a counted 95.90%. A rule whose sample size is a scaled chi-square on four degrees of freedom is not well described by two moments, which is the same statement the two-stage rule’s 5th-to-95th range of 9 to 111 makes in a different currency.
Read together, the two halves give the fixed-width problem a complete accounting with no simulation in it. Coverage is a concave function of N; a rule is judged by where it puts E[N] and how tightly; and the whole of the difference between five rules is two numbers each, which is why the closed form agrees with every one of them and why the dependence the estimated-width interval is destroyed by contributes nothing here at all.
What a practitioner should take from the pair
Three rules, and the choice between them is not a matter of taste once the question is stated precisely.
If the interval’s width is to be estimated at the end — the ordinary case, where somebody reports an estimate and a standard error — the blinded rule is the one that covers, and it costs a t multiplier on half the degrees of freedom. Nothing else on offer here is exact at a comparable number of observations.
If the width is fixed in advance and the only question is how many observations to buy, then nothing about blinding helps, the variability of N is the enemy, and the two-stage rule’s answer — overshoot deliberately, accept 2.07 times the oracle count — is buying exactly the right thing.
If the rule is being written into a protocol somebody else will follow, the block construction has a property the other two do not: it is impossible to get wrong by looking too often. A protocol that says “compute the interval from the block means” cannot be violated by an over-eager analyst who peeks, because peeking at contrasts is exactly what the rule is made of. The two-stage rule, by contrast, is exact only if nobody re-estimates the sample size a second time, which is a discipline rather than a theorem.
And if the experiment is large, all of this is a second-order argument about a few per cent, and the reason to prefer the exact rule is not efficiency but that its guarantee is a theorem rather than an asymptotic hope.
What is claimed here, and what is not
This essay takes the cost of the exactly-covering rule, and the claims are two: the price is a t multiplier on the interval’s own degrees of freedom, with a best block size in the middle; and a fixed-width interval’s coverage is a function of the sample-size distribution alone, so the same repair does not apply to it.
What stays out and is named as a decision: block sizes chosen adaptively — larger blocks early when the estimate is poor and smaller ones near the crossing, which would buy both ends of the trade and is a rule this essay does not build; non-normal data, where the second route above stops being a closed form; and any attempt to optimise the split formally, since the objective would be an expected width over a random stopping time and the answer would be a number at one setting rather than the rate the measurements give.
The boundary against the previous field is the diagnosis. That the sequential rule’s spread estimate at stopping is 0.83 of the truth, and that its estimated-width interval is short because of it, is established there; what this essay adds is that its fixed-width interval is short for an unrelated reason, and that the two were being explained by the same sentence.
The checks, and the refusals that make them mean something
Three claims are gated in this field’s library. The exactness comparison is required to hold in the form that makes it a trade: the two-stage rule at more than 1.6 times the oracle count and the blinded rule at under 1.05, with the blinded interval wider than the interval that does not cover, and wider still at large blocks — four statements about one comparison, any of which failing would mean the trade had been described wrongly. The observations spent are required to rise monotonically with the block size, which is the overshoot. And the fixed-width decomposition is required to agree with its closed form for every rule, to within four standard errors, with the blinded rule at blocks of two required to cover worse than the sequential rule — the one measurement in this field that says a repair made something worse, stated so that it would fail if it stopped being true.
The refusals for this half belong to the essay that built the construction, and one of them is worth naming again here because it is the cheapest possible version of this essay’s mistake: an interval built from the estimate that stopped the rule covers 85.1%. Everything above is an argument about a few percentage points of width. That is ten points of coverage, available to anybody who reaches for the number in hand at the moment the rule fires.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A ratio that changes between blocks — both name blinding, coverage, degrees of freedom, efficiency, fixed-width interval, interval width, nuisance parameter
- Blinded, and still exact — both name blinding, coverage, degrees of freedom, efficiency, fixed-width interval, interval width, nuisance parameter
- A width the trial has to stop for — both name blinding, coverage, fixed-width interval, interval width, sequential analysis, stopping rule
- Stopping on the arms — both name blinding, coverage, fixed-width interval, interval width, sequential analysis, stopping rule
- The bias that lands in the slope — both name blinding, coverage, degrees of freedom, fixed-width interval, interval width, nuisance parameter
- The condition that cannot be dropped — both name blinding, coverage, degrees of freedom, efficiency, fixed-width interval, interval width
Named objects
A flat tag is an object no other essay names yet.
BlindingClosed formCoverageDegrees of freedomEfficiencyFixed-width intervalInterval widthJensens inequalityNuisance parameterSample sizeSequential analysisStein two-stageStopping ruleT interval