A design that assumes less

What a design chosen from the data costs

Two fields on this site measured what happens when a rule reads the data, and the error rate broke both times. A design that reads the data to decide where to put its runs breaks nothing — and the control that proves it also finds what the real shortfall is.

Worth reading first: The design that needs the answer · When the looking happens.

Two fields on this site have measured what happens when a rule looks at the data before deciding what to do next, and both times something broke.

The stopping-rule field: testing five times at the nominal level rejects a true null 14% of the time, and the p-value’s meaning depends on a sampling plan that appears nowhere in the data. The adaptive field: an estimate from an arm chosen for being ahead is ahead by more than it should be, and the unbiased estimate is the one that throws away the data the choice was made on.

The design in the previous essay reads the data — the whole point of it is that the second stage is placed at an estimate computed from the first — and reports an interval afterwards computed as though the settings had been fixed in advance. On the evidence of those two fields, that interval should not cover what it claims.

It does. Nearly.

The measurement, with the control that makes it one

What the interval covers, after a design that read the data800 experiments of 12 runs, all at the same truth. The first pair is a design fixed in advance; the second is one whose settings were chosen from the first stage's own outcomes. If choosing the design from the data broke the inference, the second pair would sit below the first, and it does not — 91.3% against 92.6%. What does move the coverage is the shape of the interval rather than the design: the Wald interval assumes the estimate is normal around its own standard error and falls short under both designs, and the profile interval, computed from the same residual sums of squares with no derivative in it, covers 94.3% and 94.9%.fixed at the guess — Wald92.6%fixed at the guess — profile94.3%chosen from the data — Wald91.3%chosen from the data — profile94.9%95%, which all four claim800 experiments of 12 runs, guess K = 1, truth 3the design is not what the shortfall is about
Fig. 1 Sixteen hundred experiments of twelve runs. The first pair is a design fixed before the experiment; the second is one whose settings were chosen from the first stage’s own responses. Two interval constructions on each.

At twelve runs the nominal 95% Wald interval covers 92.6% after a design fixed in advance and 91.3% after a design chosen from the data. The difference is 1.3 points, on eight hundred experiments per row, where a single row’s standard error is about 0.9 points.

The control is what makes that number readable. The fixed design’s interval does not cover 95% either. Whatever is wrong is wrong at both, and the adaptivity is contributing a point or so on top of a shortfall of two and a half that has nothing to do with it.

Both of those quantities are worth having, and reporting only the second — the number a study actually produces — would have made a data-chosen design look responsible for a defect it mostly inherited.

Where the shortfall actually comes from

A Wald interval is the estimate plus or minus a multiple of its standard error, and the standard error comes from the curvature of the fitted surface at the estimate. That construction assumes the surface is a parabola. For a non-linear model it is not, and the further the model is from linear in its own parameters the worse the assumption is.

The check available is another interval built from the same fit without that assumption. A profile interval is the set of parameter values whose best fit is not significantly worse than the best fit overall — an inversion of a residual sum of squares, with no derivative in it anywhere.

What the interval covers, after a design that read the data. 800 experiments of 24 runs, all at the same truth. The first pair is a design fixed in advance; the second is one whose settings were chosen from the first stage's own outcomes. If choosing the design from the data broke the inference, the second pair would sit below the first, and it does not — 94.5% against 93.5%. What does move the coverage is the shape of the interval rather than the design: the Wald interval assumes the estimate is normal around its own standard error and falls short under both designs, and the profile interval, computed from the same residual sums of squares with no derivative in it, covers 95.1% and 94.1%.
Fig. 2 The same four rows on twice as many runs: Wald at 93.5% and 94.5%, profile at 95.1% and 94.1%. The gap between the two constructions has closed to about the noise of eight hundred experiments, which is what a small-sample effect looks like halfway to being gone.

At twelve runs the profile interval covers 94.3% after the fixed design and 94.9% after the data-chosen one, against the Wald interval’s 92.6% and 91.3%. Both designs recover, so the shortfall belongs to the interval’s shape rather than to how the settings were arrived at. The diagnosis is the profile interval fixing it for the design that never adapted.

That is a satisfying answer and it should be stated with the caveat that comes with it: the profile interval is not exact either, it is better here at these sample sizes, and it costs a search over the parameter rather than one derivative.

Reading the two shortfalls apart

It is worth being explicit about how the two components are separated, because it is the only inferential move in the essay and everything rests on it.

Four cells are measured: two designs — fixed in advance, chosen from the data — crossed with two interval constructions — Wald, profile. Comparing down a column isolates the design, since the interval construction is held constant. Comparing across a row isolates the construction, since the design is held constant. The two comparisons are made on the same experiments with the same seeds, so nothing in the table is looking at different data from anything else in it.

Down the Wald column at twelve runs: 92.6% to 91.3%, so the design costs about a point. Across the fixed-design row: 92.6% to 94.3%, so the construction is worth about two. The two effects are roughly additive and the larger of them is the one nobody would have suspected from the way the question is usually asked.

What the interval covers, after a design that read the data. 800 experiments of 12 runs, all at the same truth. The first pair is a design fixed in advance; the second is one whose settings were chosen from the first stage's own outcomes. If choosing the design from the data broke the inference, the second pair would sit below the first, and it does not — 91.1% against 89.9%. What does move the coverage is the shape of the interval rather than the design: the Wald interval assumes the estimate is normal around its own standard error and falls short under both designs, and the profile interval, computed from the same residual sums of squares with no derivative in it, covers 94.5% and 93.5%.
Fig. 3 The same four cells started from a twelvefold wrong guess rather than a threefold one. The ordering reverses: the Wald interval covers 89.9% after the fixed design and 91.1% after the data-chosen one, because a design stuck at a badly wrong guess estimates the parameter poorly and a poorly estimated parameter is where the quadratic approximation is worst.

Why the adaptivity is nearly free, and where it is not

The reason is a version of the argument the stopping-rule field makes, arriving at the opposite conclusion, and the difference between the two cases is worth pinning down.

A sequential design chooses its next settings from the responses already observed. Those responses are in the likelihood, and the choice adds a factor to it that depends only on data already conditioned on — so the likelihood for the whole experiment factorises into the part that is about the parameter and a part that is about the design decision and contains no parameter. Anything computed from the likelihood alone therefore does not know the design was adaptive, which is the likelihood principle’s own statement.

Frequentist properties are not determined by the likelihood, so this does not settle the question — which is exactly why it was measured rather than argued. What the measurement says is that at these sizes the effect is a point of coverage and the approximation is worth two and a half.

The contrast with the stopping-rule field is that a stopping rule makes the sample size random and correlated with the estimate: stopping early when the data look extreme selects on the very statistic being reported. A design rule here changes where the runs go and not how many, and it selects on an estimate of K in a way that is nearly orthogonal to the error in the final estimate. Nothing is being chosen for looking good.

That reversal is worth a sentence of its own, because it is the strongest form the essay’s claim takes. At a badly wrong guess, choosing the design from the data improves the interval’s coverage. Not because adaptivity is good for coverage — it is not, and at a mild guess it costs a point — but because the dominant term in the shortfall is the curvature of a fitted surface, and a better-placed experiment has a better-behaved surface. The profile intervals at the same setting are 94.5% and 93.5%, so the construction still explains most of it under both designs.

The interval is narrower as well as honest

There is a second half to the accounting that a coverage figure alone hides.

At twelve runs the data-chosen design’s interval is 1.995 wide on average and the fixed design’s is 2.161 — eight percent narrower, at coverage that is a point lower. That is the design doing its job: a better-placed experiment estimates the parameter more precisely, and the interval reflects it.

Combining the two, the data-chosen design produces intervals that are shorter and very slightly less reliable. Whether that is a good trade is the same question every efficiency gain poses, and the numbers to weigh are on the page rather than in a general argument.

Updating between runs, against never updating at all. A threefold wrong guess — the design is built for K = 1 and the truth is 3 — and three ways of spending the runs. The flat line is the local design at the guess, 81.1% efficient at any size, because running the wrong design more times does not make it a better design. The two-stage design spends 40% of its runs there and the rest at the estimate, and lands near 93% throughout. The fully sequential design updates after every 2 runs and improves with size, from 87.8% at 12 runs to 99.1% at 160 — crossing the two-stage design where its early estimates stop being noise.
Fig. 4 And the efficiency that narrowness comes from, from the previous essay: the data-chosen design reaching 96.5% of the oracle at forty runs where the fixed design stays at 81.1%.

What the estimate itself does

Coverage is a property of an interval, and an interval has a centre. The centre is worth checking separately, because the adaptive field’s failure was a bias in a point estimate rather than a defect in an interval, and a bias would show up here as a coverage that is fine on average and asymmetric.

At twelve runs, on the same experiments the coverage above was counted on, the root mean squared error of K̂ is 0.566 after the fixed design at the guess and 0.540 after the fully sequential one, against 0.486 for the design that knew the answer. The data-chosen design’s estimate is better than the fixed design’s and short of the oracle’s, which is the ordering the efficiency numbers imply and is the third place the same fact has been measured.

There is no sign of the adaptive field’s problem here and there is a structural reason not to expect one: nothing in the procedure is chosen for being large. The second stage is placed where information about K is cheapest given the current estimate, and that placement is a function of K̂ rather than of whether K̂ came out flattering. An estimator whose subsequent data are collected at settings determined by its own value is not the same thing as an estimator selected for its value.

Updating between runs, against never updating at all. A threefold wrong guess — the design is built for K = 0.25 and the truth is 3 — and three ways of spending the runs. The flat line is the local design at the guess, 34.6% efficient at any size, because running the wrong design more times does not make it a better design. The two-stage design spends 40% of its runs there and the rest at the estimate, and lands near 79% throughout. The fully sequential design updates after every 2 runs and improves with size, from 63.2% at 12 runs to 97.7% at 160 — crossing the two-stage design where its early estimates stop being noise.
Fig. 5 The efficiency behind all of it at a twelvefold wrong guess: 34.6% for the design that never learns, against 90.4% for the one that does. Everything this essay measures is the inferential price of that gap, and the price is about a point of coverage at twelve runs and nothing at forty-eight.

What happens as the experiment grows

What the interval covers, after a design that read the data. 800 experiments of 48 runs, all at the same truth. The first pair is a design fixed in advance; the second is one whose settings were chosen from the first stage's own outcomes. If choosing the design from the data broke the inference, the second pair would sit below the first, and it does not — 95.3% against 94.9%. What does move the coverage is the shape of the interval rather than the design: the Wald interval assumes the estimate is normal around its own standard error and falls short under both designs, and the profile interval, computed from the same residual sums of squares with no derivative in it, covers 95.4% and 94.8%.
Fig. 6 Forty-eight runs, where all four rows are within half a point of 95% and of each other. Everything measured in this essay is a small-sample effect, and both of its components — the approximation and the adaptivity — have gone by the time the experiment is the size of an ordinary one.

At forty-eight runs the Wald intervals cover 94.9% after the fixed design and 95.3% after the data-chosen one, and the profile intervals 95.4% and 94.8%. The four constructions have converged on each other and on their claim, the ordering between them has stopped being stable, and nothing in the essay’s argument is visible any more.

That is worth stating in the same breath as everything above. The failure being measured here is small and it is a small-sample failure, which is the honest summary and also the reason the measurements were made at twelve runs rather than at forty: a defect that is invisible at the sizes anybody uses is a defect nobody needs to be warned about, and a defect that is visible only below them should be reported with the size attached.

Five cells, and the sign reverses in every one

The essay measures the adaptivity’s contribution five times — at twelve, twenty-four and forty-eight runs from a threefold wrong guess, and at twelve and twenty-four from a twelvefold one — and reading the five contrasts together says more than any of them says alone.

Taking the data-chosen design’s coverage minus the fixed design’s, the Wald contrasts are −1.3, +1.0, +0.4, +1.2 and +2.3 points. The profile contrasts on the same experiments are +0.6, −1.0, −0.6, −1.0 and −2.0.

Every pair has opposite signs, and in four of the five the two magnitudes agree to within four tenths of a point. Summing each pair gives −0.7, 0.0, −0.2, +0.2 and +0.3 — all within seven tenths of zero, on rows whose own standard error is about eight tenths.

So what the adaptivity moves is not coverage; it is the gap between the two constructions. Averaged over the pair, the design’s contribution is 0.1 points across all five settings, which is a far stronger statement than “about a point at twelve runs” and is the one the recommendation not to adjust actually rests on. A correction sized to the Wald column would be correcting in the wrong direction for the profile column, on the same experiments.

The mechanism the reversal points at is the one the essay diagnoses from the other side. A better-placed experiment has a better-conditioned fitted surface, so the quadratic approximation the Wald interval makes is less wrong — which lifts the Wald coverage — while the profile interval, which never made that approximation, has nothing to gain and gives back the same amount through the estimate being slightly more variable. The two constructions are reading opposite halves of one change in curvature.

That also explains the reversal the essay flags at a twelvefold wrong guess without needing it as a special case: there the curvature improvement is largest, so the Wald column gains most (+1.2 and +2.3) and the profile column loses most (−1.0 and −2.0). It is the same effect at a larger amplitude rather than a different phenomenon.

The interval narrows faster than the estimate improves

The width and the estimate are both reported at twelve runs and dividing one by the other says where the small negative in the Wald column comes from.

The data-chosen design’s interval is 1.995 wide against the fixed design’s 2.161, a ratio of 0.923. Its root mean squared error is 0.540 against 0.566, a ratio of 0.954. The interval shrank by 7.7% and the estimate improved by 4.6%, so the interval is about 3.2% narrower than the improvement warrants — and an interval 3.2% too narrow covers about 94.2% instead of 95%, which is eight tenths of a point.

The counted Wald contrast at that cell is 1.3 points, so the width accounts for most of it and the rest is noise on eight hundred experiments. The adaptivity’s cost, where it has one, is not a distortion of the interval’s centre — it is the interval keeping pace with a precision gain slightly faster than the gain arrives. That is a rounding of the design’s own success rather than a defect in the inference, and it is why it does not survive being averaged with the profile column.

What to report after a design that read the data

The measurements support a short list, and it is shorter than the one the neighbouring fields require.

Report the design that was run, not the design that was planned. A sequential experiment’s settings are data, they differ between replications of the same protocol, and a reader cannot check an efficiency claim without them.

Use an interval that does not assume a parabola. At the sizes where any of this matters, the profile interval is worth about two points of coverage and the adaptivity about one, so the larger repair is the one nobody was asking about.

Do not adjust for the adaptivity. There is nothing to adjust: the effect is a point at twelve runs, it is not consistently signed across guesses, and a correction sized to it would be noise dressed as rigour.

Say what the rule read. That is the sentence that separates this from the two fields where things broke, and it is one line: the design was revised from an estimate of the parameter, and no response was used to decide anything about how the results would be analysed or when the experiment would end.

What the interval covers, after a design that read the data. 800 experiments of 24 runs, all at the same truth. The first pair is a design fixed in advance; the second is one whose settings were chosen from the first stage's own outcomes. If choosing the design from the data broke the inference, the second pair would sit below the first, and it does not — 94.4% against 92.1%. What does move the coverage is the shape of the interval rather than the design: the Wald interval assumes the estimate is normal around its own standard error and falls short under both designs, and the profile interval, computed from the same residual sums of squares with no derivative in it, covers 95.8% and 93.8%.
Fig. 7 Twenty-four runs from a twelvefold wrong guess — the case where the design has the most to do and the sample is still small. Wald at 92.1% and 94.4%, profile at 95.8% and 93.8%: the spread across the four cells is under four points and the reversal seen at twelve runs is still there, with the data-chosen design’s Wald interval covering better than the fixed design’s.

What a reader should take from three fields at once

The three fields that read data before deciding have now been measured against each other, and the pattern is legible.

Reading the outcome to allocate treatment breaks the estimate and the error rate, because the rule’s objective and the analysis’s objective are different and the rule selects on the quantity being reported.

Reading the covariates to allocate treatment breaks nothing and makes the unadjusted analysis conservative — 0.6% of true nulls rejected where 5% is claimed — because the rule removes variation the analysis is not told about.

Reading the responses to place the runs costs about a point of coverage at twelve runs and nothing at forty-eight, because the design’s objective is precision about the same parameter the analysis reports.

What separates them is not how much data the rule reads but what it selects on. A rule that selects on the thing being estimated is dangerous; a rule that selects on where information is cheapest is not.

How long to go on guessing. 40 runs in total, a guess of K = 1 against a truth of 3, and the first stage varied from 5% of the runs to 70%. The optimum is interior and it is early: 97.1% at 10%, which is 4 runs — barely more than the four a two-parameter model needs to be fitted at all. Spending less leaves the second stage designed at an estimate with nothing in it; spending more leaves most of the experiment at a setting already known to be wrong, and by 70% the whole procedure has given back 9.8 points of what it bought.
Fig. 8 The design decision behind every experiment measured here — how much of it to spend before revising. Nothing in this essay’s coverage numbers moves detectably with that choice, which is worth knowing before optimising it.

What is claimed, and what is not

The claim is the inferential cost of a design chosen from the data, measured with a fixed design at the same number of runs as the control and with two interval constructions to separate the approximation from the adaptivity.

What stays out: exact conditional inference after a sequential design, which would condition on the realised design and is the correct answer where one is needed; bootstrap intervals, which are the practical answer and cost a refit per resample per experiment; and adaptive designs whose stopping is also data-dependent, where the stopping-rule field’s results apply in full and none of the reassurance here transfers.

The boundary against the stopping-rule field is the one this essay is mostly about: that field owns what a rule for when to stop does to a p-value, and this owns what a rule for where to look does to an interval. They share the shape of the question and have different answers, and the difference is in what is being selected on rather than in how much the rule knows.

The checks, and the refusal

Four claims are gated. The Wald interval after a design fixed in advance must already fall short of 95%, which is the control and the whole basis of the diagnosis. The data-chosen design’s coverage must be within three points of the fixed design’s — the claim that adaptivity is nearly free, stated as a bound rather than as a reassurance. The profile interval must cover more than the Wald interval under both designs, which is what attributes the shortfall to the approximation. And the estimate after the data-chosen design must be better than the one the fixed design produces, so that the trade being described has both of its sides measured.

The refusal for this field is the unstandardised maximin search two essays back, and this essay adds no second one — because the honest finding here is a null result, and a refusal that manufactured a failure to match the neighbouring fields would be the opposite of what the machinery is for.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Adaptive designConfidence intervalCoverageDegrees of freedomLeast squaresMichaelis–MentenMonte CarloThe non-linear modelOptional stoppingSelection biasTwo-stage designWald interval