The run that confirms it
Worth reading first: Walking up the gradient.
Every response-surface text ends the same way. Fit the surface, find the optimum, and then run a confirmation experiment at the setting found — because a recommendation that has not been checked is an extrapolation.
The advice is right and the reason usually given for it is not. The confirmation run is not there to guard against a mistake, an arithmetic slip or a model that failed to fit. It is there because the prediction it checks is systematically too high — not on unlucky studies, on average, by an amount the design knows in advance and nobody computes.
Where the gap comes from
The mechanism is one sentence and it has nothing to do with the model being wrong.
The recommended setting is , the point where the fitted surface is highest. The fitted surface is the true surface plus estimation error, and that error varies from setting to setting. So is not merely where the truth is highest — it is preferentially where the error happens to be positive.
The fitted height at that point is therefore the true height there plus an error that was selected for being large. Its expectation is above the truth, and the amount is the selection effect: the expected maximum of a random field over the region the search ranged across.
That is the same object as the winner’s curse in the adaptive field — the arm chosen for being ahead is ahead by more than it should be — arriving in a setting where nothing was chosen between arms and the selection is over a continuum.
Two gaps, and they are different quantities
The figure has three levels and it is worth being clear which pair each gap is between.
The prediction against the truth at the same setting is the selection effect. It is a property of how the setting was chosen, it would be there even if the recommended setting happened to be the true optimum, and it is what a confirmation run measures.
The truth at the recommended setting against the true best is the cost of being in the wrong place. It is a property of how accurate the fit’s direction was, and a confirmation run does not measure it — the run reports the response at the setting, and nothing in the experiment says what was available elsewhere.
At twice the noise the first is 0.858 and the second is 0.527. At four times the noise they are 2.098 and 2.380, and they have crossed.
So an experimenter who runs a confirmation and finds it below the prediction has learned about the first gap and nothing about the second, which is the one that decides whether the recommendation was worth acting on.
The scale is the prediction’s own standard error
The size of the selection effect has a natural unit and using it makes the pattern legible.
The fit reports a standard error for its prediction at the recommended setting — it is with the model vector there, which the design supplies. Measured in those units the gap runs 0.24, 0.45, 0.72, 0.83, 0.82 across the five noise levels.
That is a quantity between a quarter and five sixths of a standard error, and it does not grow without limit. It cannot: the selection effect is bounded by the spread of the random field over the region, and the standard error grows with the same , so the ratio approaches a constant.
Three quarters of a standard error is the number to carry. A fit that predicts at the setting it recommends is, on average, predicting about 0.9 too much — and a confirmation run that comes in that far below is behaving exactly as it should rather than contradicting anything.
The comparison between the two rules is the one useful design consequence here. The careful rule lands at better settings and is more over-optimistic about them, because it maximises over a larger set — the whole region rather than one point per fit — so the maximum of the error field it selects on is larger.
Choosing better makes the prediction worse. That is the general form of the selection effect and it is worth stating separately from the arithmetic: any procedure that searches harder for the best reported value reports a more inflated one, whatever the searching is over.
The flattening is worth one more sentence because it is not obvious. The selection effect is the expected maximum of the prediction error over the set the search ranged across, and that maximum grows linearly in — every realisation of the error field scales together. The standard error grows linearly in too. So the ratio is asymptotically constant and the only reason it moves at all across the sweep is that the set the search ranges over changes with the noise: at low noise the recommended setting is nearly always the same place, and at high noise it wanders, which is a larger set to take a maximum over.
That reading also says where the ratio would be larger. A design whose prediction variance is very uneven across the region gives the search more room to find a favourable error, so the selection effect is larger relative to the standard error at the chosen point. A rotatable design, whose precision depends only on distance, is the arrangement that gives the search least to exploit — which is a property this field chose rotatability for other reasons and gets here as a side effect.
What a confirmation run can and cannot settle
Three questions get asked of a confirmation run and it answers one of them.
Is the response at this setting what was predicted? Answered directly, and the answer is usually no, by about three quarters of a standard error. A confirmation run that matches the prediction is mildly surprising.
Is this setting better than where the process was? Answered, if the previous operating point’s response is known — which it usually is, because the experiment was centred there. This is the question the run is actually for.
Is this the best setting available? Not answered at all. The run measures one point. The response elsewhere is not observed, the fit’s opinion about it is the thing being checked, and a single observation at one setting carries no information about a surface.
The third is the question the recommendation was about, and it is the one a confirmation run cannot touch. What would answer it is another experiment centred at the new setting, which is what the steepest-ascent discipline does as a matter of course and what response-surface practice does only when the confirmation run disappoints.
What to predict instead
The selection effect is estimable before the confirmation run, which means the prediction can be corrected rather than merely disbelieved.
The gap is about three quarters of the prediction’s standard error at this design and these noise levels, and the standard error is computed by the software already. Subtracting three quarters of it from the predicted response gives a number a confirmation run would match on average.
That is a crude correction and it is better than none, and the honest caveats are two. The factor is not a constant — it runs from 0.24 to 0.83 across the noise levels here, rising as the fit gets noisier and flattening — so a single multiplier is an approximation. And it is specific to this design and this family of truths; a different design ranges its search over a different field.
What generalises is the sign and the scale. The prediction at a chosen optimum is too high, by a fraction of a standard error rather than by a multiple of one, and an experimenter expecting a confirmation run to land on the prediction is expecting the wrong thing.
The same effect, four fields over
The selection effect has now been measured in five places and the arithmetic has been the same every time, so the general statement is worth extracting.
A quantity is estimated at many candidates, the best is chosen, and the estimate at the chosen one is reported. An arm chosen for being ahead is ahead by more than it should be. The smallest of twenty p-values is not a p-value. A model chosen by a criterion fits its own selection sample better than it fits anything else. A setting chosen for being the fitted optimum has an inflated fitted height.
The size is always the expected maximum of a noise field over the set searched, which means it grows with the number of effective candidates and with the noise, and is bounded by the field’s own spread. And the repair is always one of three: condition on the choice, estimate from data the choice was not made on, or shrink by the amount the selection is worth.
What is unusual here is only that the candidate set is a continuum rather than a list, so nobody experiences it as a choice. There is no moment at which twenty options were narrowed to one; there is a formula that returns a point. The selection is inside the arithmetic, and that is why response surfaces are the one place the effect is missing from the standard account.
Why the interval does not cover it either
The fit draws an interval around its prediction, and the natural next question is whether the truth at the recommended setting falls inside it.
It falls below the interval’s lower end on 7.0% of studies at twice the noise — against the 2.5% a correctly centred 95% interval would give — and on 7.8% at three times. The asymmetry is the selection effect again: the interval is centred on an inflated prediction, so it misses low far more often than it misses high.
Seven per cent is not a catastrophe and the direction is the whole point. A 95% prediction interval that missed symmetrically would be a nuisance; one that misses low three times as often as it should is telling an experimenter something systematic, and what it is telling them is that the centre is wrong rather than that the width is.
What it costs to ignore
Putting the two gaps together gives the number an experimenter is actually exposed to, which neither alone reports.
At twice the noise the fit predicts 62.679, the truth at the recommended setting is 61.821, and the best available is 62.348. So the recommendation is 0.527 below the best available and the expectation set for it is 0.858 above what it delivers. An experimenter who acts on the recommendation and expects the prediction is disappointed by 0.858, of which 0.331 is the setting being wrong and 0.527… which does not add, because the two are measured against different baselines and the decomposition is the one from the third section.
Stated without arithmetic: the process improves less than the experiment said it would, for two reasons that are easy to confuse, and only one of them is visible from the confirmation run. A team that attributes the whole disappointment to the setting being wrong will run another experiment somewhere else; a team that attributes it all to the prediction being inflated will keep the setting. Both are partly right at twice the noise and the first is right at four times, where the setting’s own shortfall has overtaken the selection effect.
The two gaps cross, and where they cross depends on the noise — which is the reason to measure both rather than to adopt a rule about which matters.
There is a cheap diagnostic that separates them and it needs one extra run. Measure the response at the centre of the design as well as at the recommended setting. The centre’s response is known from the experiment already, so comparing the two says how much the process actually improved, and that number is free of the selection effect entirely: the centre was not chosen for anything. An improvement that is real but smaller than promised is the signature of an inflated prediction; no improvement at all is the signature of a setting in the wrong place.
What is claimed here, and what is not
The claim is what a confirmation run at a fitted optimum finds: that the fit predicts more than is there at every noise level, by 0.858 at twice the noise and 2.098 at four times; that measured in the prediction’s own standard error the gap runs from 0.24 to 0.83 and flattens rather than growing; that the rule which searches over the whole region rather than one point per fit lands at better settings and is more over-optimistic about them; and that the truth at the recommended setting falls below the fit’s own 95% interval on 7.0% of studies rather than 2.5%.
Every number is two and a half to four thousand studies with the same thirteen-run composite design, on a truth whose optimum is inside the region.
What stays out: a correction derived rather than measured, which would need the expected maximum of the prediction-error field over the region and is the kind of quantity the extremes field computes; the confirmation run’s own noise, which this essay holds separate by comparing against the truth rather than against a simulated observation; and sequential confirmation, where a disappointing run triggers another experiment and the selection compounds across stages.
The prediction interval quoted here is the fit’s ordinary one and takes the estimated coefficients as known, so its shortfall is partly the same parameter-uncertainty gap the forecast field measures. The asymmetry is what this essay’s claim rests on and it is not explained by that.
Still open: an optimum chosen and then reported
The inflation is measured above and not removed. What would remove it is what adaptive designs already use: an estimate that conditions on the choice, or one computed from data the choice was not made on.
The second is the cheaper and the one this design could support. Splitting the runs — fitting the surface on some and estimating the height at the chosen setting from the rest — gives an unbiased prediction at the cost of a noisier surface, and whether that trade comes out positive at thirteen runs is a measurement nobody here has made. It is the same trade a criterion that predicts a hold-out makes when it decides whether to spend data on being checked, arriving in a field with no hold-out tradition at all.
The check, and the refusal
Three claims are gated. That the fit predicts more at the setting it chose than is there, at every noise level — an inequality rather than a value, because the value depends on the truth and the inequality does not. That the gap grows with the noise by more than a factor of three across the sweep, which is what makes it a selection effect rather than a constant offset. And that the response actually obtained falls away as the noise grows, which is the second gap moving and is what stops the first being read as the whole story.
The refusal is the one that keeps the two gaps apart: the measurement is taken against the truth at the chosen setting rather than against a simulated confirmation observation. A comparison against a noisy observation would have the right expectation and would let the selection effect hide inside the observation’s own error, which is exactly how the effect stays invisible in practice — a single confirmation run that comes in low is indistinguishable from bad luck, and only the average over many studies shows that low is where it comes in.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A design is a number — both name central composite design, experimental design, prediction variance, response-surface
- The theorem that says when to stop — both name central composite design, experimental design, prediction variance, response-surface
- Four letters and two camps — both name central composite design, experimental design, prediction variance
- Guessing one arm in three — both name experimental design, monte carlo, selection bias
- The design that refuses the corners — both name central composite design, experimental design, prediction variance
- The effect a stopped trial reports — both name monte carlo, selection bias, the winner's curse
Named objects
A flat tag is an object no other essay names yet.
Central composite designExperimental designForecast intervalMonte CarloPrediction varianceResponse-surfaceRidge analysisSelection biasStationary pointThe winner's curse