The interval after a stop it chose
Worth reading first: When the looking happens · What the 95% refers to.
The purely sequential rule stops the first time its estimated precision is good enough, and its interval covers 90.2% where it claims 95%. That is the number the previous essay ends on, and by itself it is only a fact. What is worth having is which of the several things the rule does is responsible for it, because the candidates suggest completely different repairs.
Three are available. The sample size is random. The sample size is too small on average — 20.9 observations where knowing σ would demand 24.0. And the sample size is a function of the same numbers the interval is built from. Each of the three would be repaired differently and only one of them can be repaired at all.
Four intervals at one stopping time
The way to separate them is to change one thing at a time and count. Every experiment below uses the identical stopping rule, so N has the same distribution in all four rows; what changes is what the interval is built from.
The rule’s own interval — x̄ₙ ± 0.4 — covers 90.2%, which is the same shape of failure as twenty intervals and one miss counted at a rule rather than at a formula.
A t interval on the same data, x̄ₙ ± t·sₙ/√N, covers 91.9%. Replacing the fixed half-width with one estimated from the sample repairs about a fifth of the shortfall and no more.
The fixed half-width on a fresh sample of the same random size covers 89.7% — still short, and by about the same amount. So a half-width of 0.4 is genuinely too narrow for the N this rule chooses, which is the second of the three candidates and is worth its own line: the rule stops early on average, and an interval built for the sample size it stops at is short even on data the rule never touched.
A t interval on a fresh sample of the same random size covers 95.4%.
That last row is the one that decides the question. Keep the rule, keep its random and systematically-small sample size, and change nothing except whose data the interval is built from, and the coverage is what it claims. The randomness of N costs nothing. The dependence between N and the sample costs everything the estimated width does not.
The mechanism, one level down
The dependence has a direction and it is visible in a single number: the estimated variance at the moment of stopping averages 0.8315 of the truth.
The rule approaches its boundary from below, taking one observation at a time and stopping the first time n ≥ z²s²ₙ/d². Among all the samples that reach n without having stopped, the ones that stop there are the ones whose s happens to be small — so the sample the interval is built from is a sample selected for having a low spread, and its s underestimates σ by nearly a fifth.
An interval whose width is proportional to that s is too narrow by nearly a tenth, and the shortfall in coverage follows. This is the winner’s curse with a stopping rule in place of a selection: a quantity conditioned on having crossed a threshold is not distributed like the quantity, and the estimate that triggered the decision is biased by the fact that it triggered it.
The four rows are a factorial, and it has an interaction
The four intervals differ in two respects and not one, and laying them out that way says something the list of rows does not. One factor is whose data the interval is built from — the sample the rule stopped on, or a fresh one of the same random size. The other is where the half-width comes from — the fixed 0.4, or t·sₙ/√N estimated from whatever data the row uses. Four rows, two factors, every combination present: this is a complete two-by-two design and it can be read as one.
Estimating the width instead of fixing it is worth 1.7 points on the rule’s own sample, from 90.2% to 91.9%. The same change on a fresh sample is worth 5.7, from 89.7% to 95.4%. Going the other way: using fresh data instead of the rule’s own costs 0.5 points when the width is fixed, and buys 3.5 when the width is estimated. Both readings give the same interaction, 4.0 points, and its size is the whole finding — it is larger than either main effect measured on the rule’s own data.
Neither factor has an effect that can be quoted on its own, and that is why the three candidate explanations at the top of this essay pointed in such different directions. A reader who changes one thing at a time, in the arrangement the rule actually delivers, learns that estimating the width buys 1.7 points and that freshening the data buys nothing at all — and concludes, reasonably and wrongly, that neither repair is worth pursuing. Both are worth pursuing, and only together. One factor at a time is the general statement of that failure, and this is an unusually clean instance of it: the interaction is not a subtlety at the edge of the noise, it is the largest of the three effects.
The mechanism behind the interaction is stated in the paragraphs above and can now be read off the arithmetic. A fixed half-width does not know what the sample’s spread was, so replacing the sample changes almost nothing; an estimated half-width is computed from a spread that the stopping rule selected, so replacing the sample removes the selection and the width recovers. The two factors interact because only one of them can carry the defect the other one introduces.
Nine percent of a width, and why one multiplier cannot return it
The variance at the stopping moment averages 0.8315 of the truth, which puts the spread at √0.8315 = 0.912 of σ. An interval whose half-width is proportional to that spread is 8.8% too narrow, and restoring it would take a multiplier of 1/0.912 — about 1.10 on the width.
That multiplier is the obvious repair and it is the third of the three the ranking above refuses, for a reason the numbers make concrete. The bias is a selection effect, and selection effects are not constant across the thing selected on: among samples that stop early, the spread had to be unusually small to reach the boundary that soon, and among samples that run long it barely had to be small at all. 0.8315 is an average over a mixture whose components are biased by different amounts, so a single multiplier over-corrects the long runs and under-corrects the short ones, leaving an interval that is wrong in both directions instead of one direction.
The first-stage evidence says the same thing from outside. At n₀ = 20 the sequential rule already covers 94.9% at d = 0.25, so a fixed 10% inflation applied there would push it past its claim and buy width nobody needs. A multiplier calibrated at one first stage is not a repair at another, which is the difference between a correction and a construction — and the two-stage rule is a construction: its width is exact for every σ, every d and every n₀, because the quantity it is built from was never allowed to see the boundary.
Three candidates, three different repairs
It is worth spelling out what each of the three explanations would have implied, because they are not variations on one story and the measurement above chose between them.
If the trouble were the randomness of N, the repair would be a correction that accounts for a random sample size — an adjusted critical value, a conditional argument, something that prices the variability of N. That family of repairs would be aimed at nothing: the fresh-sample row has exactly the same random N and covers 95.4%.
If the trouble were that N is too small on average, the repair would be to inflate it — stop at 1.1 times the boundary, or add a fixed number of observations at the end. That family aims at something real, worth 5.3 points on its own by the fresh-sample fixed-width row, and it would leave the rest.
If the trouble is the dependence, the repair has to break it, and the only ways to break it are to compute the width from data the rule did not use or to stop for a reason the interval does not depend on. The two-stage rule does the second, and there is no version of the sequential rule that does either.
The measurement says the third is most of it and the second is the rest, in the proportions the four rows give. That is a more useful answer than the raw shortfall, because it says which repairs cannot work rather than only how large the problem is.
The two-stage rule as the control
The previous essay’s other rule is the control that makes all of this a measurement rather than an argument. It has a random sample size, from the same family of requirements, and its interval covers 95.9% at the same setting.
The difference is one line of construction: its width is computed from s₀, the first stage’s spread, and its stopping point is computed from s₀ as well — so the quantity that decides when to stop and the quantity that sets the width are the same number, and neither of them is the mean the interval is about. √N(x̄ₙ − μ)/s₀ is a t on n₀ − 1 degrees of freedom whatever N turns out to be.
So the repair for a fixed-width interval is known and is the two-stage rule, at 2.01 times the observations. What is not available is a repair that keeps the sequential rule’s sample size and fixes its interval, because the thing that makes its sample size good — that it uses every observation to decide — is the thing that makes its interval bad.
Inside a design, where both halves adapt
The same question inside the field’s non-linear model has one extra complication and one extra control. The design can also read the data, and the previous field established that this costs nothing; so an experiment that adapts both where its runs go and how many there are can have the stopping half charged separately.
The stopping rule takes 46.2 runs on average and its interval covers 93.0%. The fixed-length control takes 44.0 runs — two fewer — and covers 94.2%, with an interval 6.3% wider.
Fewer runs, better coverage, wider interval. The stopping rule was buying its narrow interval by stopping when the estimate said it could, which is the same selection as before wearing a design’s clothes. And the control being shorter is what makes the comparison conservative in the right direction: if the fixed-length experiment covered better because it had more data, the number would prove nothing.
What is left over, and what it is not
The non-linear case has a second shortfall in it that has nothing to do with stopping, and the robust field already isolated it: a Wald interval after a design fixed in advance is already short of 95% at small n, because the quadratic approximation behind it is poor for a non-linear model. The profile interval, computed from the same residual sums of squares with no derivative in it, covers what it claims under both designs.
That is why the fixed-length control matters more here than the raw coverage does. 93.0% against 94.2% is the stopping rule’s contribution; the remaining gap to 95% belongs to the approximation and would be there in an experiment with no rule in it at all. Reporting 93.0% as the cost of adaptive stopping would be charging the stopping rule for the model’s curvature.
Why this is not the sequential field’s problem
Two fields on this site already measure what happens when a rule reads the data, and neither of them covers this.
The sequential field is about repeated testing: five looks at a true null reject 14% of the time at a nominal 5%, and the repair is to spend the error rate across the looks with a boundary. There is one hypothesis and many opportunities to reject it.
The adaptive field is about re-estimating a sample size: the repair is to stay blind to the arms while doing it, and with that the error rate is held exactly.
Here there is no hypothesis and nothing to be blind to. The rule reads the spread, which is the quantity the interval’s width is made of and the quantity the analysis cannot proceed without. A boundary cannot help — there is no multiplicity to spend — and blinding cannot help, because the thing that would have to be hidden is the thing being estimated.
The number does not shrink with the requirement
One reassurance is available in the literature and it is worth measuring rather than repeating: the sequential rule’s shortfall is asymptotic in the required width, vanishing as d → 0, because the sample size grows and the relative bias in the spread estimate falls with it.
It is true, and the rate is slow. At a first stage of five the coverage runs 93.7%, 90.8% and 93.9% as the requirement tightens from 0.7 through 0.4 to 0.25 — not monotone, because two things are moving at once, and nowhere near 95%. At a first stage of twenty it runs 99.9%, 95.1% and 94.9%, which is at the claim for a different reason: with twenty observations already taken, most experiments at the looser widths never get to exercise the rule at all.
So the asymptotic result describes a limit the experiments anybody runs are not in. What actually controls the shortfall is the first stage — how many observations the spread is estimated from before the rule starts acting on it — and that is a quantity the experimenter chooses rather than one the requirement imposes.
What to do about it
Three things are available and it is worth ranking them, because the temptation is to reach for the one that does not work.
Use the two-stage rule if the guarantee matters more than the observations. It is exact for every σ, and its overspend falls from 2.01× to 1.14× if the first stage can be twenty rather than five.
Report the shortfall if the sequential rule is what the budget allows. It is a stated quantity at a stated setting — 90.2% at n₀ = 5 and d = 0.4, 94.9% at n₀ = 20 and d = 0.25 — rather than an unknown, and an interval described accurately is not a defective one.
Do not repair the width from the same data. The t interval at the stopping time is the obvious fix and it recovers 1.7 points of the 4.8 that are missing, which is enough to look like it worked.
What is claimed here, and what is not
This essay claims the cost of data-dependent stopping to an interval, decomposed: that the randomness of the sample size is not the problem, that the dependence between the stopping time and the spread estimate is, and that a fixed-length control on no more runs covers better and produces a wider interval.
What stays out: corrected intervals for sequential stopping, which exist and would be a fourth rule rather than a repair to this one; stopping rules that read a quantity other than the one the interval is built from, where the whole argument here does not apply and which is what the adaptive field’s blinded re-estimation is; and the asymptotic result that the shortfall vanishes as the required width goes to zero, which is true and is not what an experiment at a stated width is entitled to.
The checks, and the refusal that makes them mean something
Two claims are gated in this field’s library. The sequential rule’s own interval is required to fall short of its claim by more than three standard errors while the same rule’s sample size with a fresh sample of that size is required to cover — the pair, because either alone would be consistent with several explanations. And the spread at the stopping moment is required to be far below the truth, which is the mechanism rather than the symptom.
The refusal is the interval itself: a fixed-width rule whose stopping criterion and whose width come from the same numbers is refused, on the standard that a stated coverage has to be the coverage of the whole procedure with its stopping rule included. The check requires that failure, and requires the fresh-sample control to pass at the same time, because a refusal that fired on both would be detecting something else.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A schedule that reads the mean — both name confidence interval, coverage, fixed-width interval, monte carlo, selection effect, sequential design, stopping rule
- A block size that changes — both name confidence interval, coverage, fixed-width interval, monte carlo, sequential design, stopping rule
- Two degrees of freedom, one total — both name confidence interval, coverage, fixed-width interval, monte carlo, sequential design, stopping rule
- The trials that stopped early — both name coverage, fixed-width interval, monte carlo, random sample size, stopping rule
- What a schedule actually buys — both name coverage, fixed-width interval, monte carlo, sequential design, stopping rule
- A simulation that stops when it looks settled — both name confidence interval, coverage, monte carlo, stopping rule
Named objects
A flat tag is an object no other essay names yet.
Confidence intervalCoverageExperimental designFixed-width intervalMichaelis–MentenMonte CarloProfile likelihoodRandom sample sizeSelection effectSequential designStopping ruleTwo-stage designVariance estimate