The surface between the corners

The run that confirms it

The setting a response-surface analysis recommends was chosen because the fitted surface was highest there, so the height the fit predicts at it is a maximum over a random field. At twice the noise the fit predicts 0.858 more than is there — 0.72 of the prediction's own standard error — and the gap is not noise, it is the selection.

Worth reading first: Walking up the gradient.

Every response-surface text ends the same way. Fit the surface, find the optimum, and then run a confirmation experiment at the setting found — because a recommendation that has not been checked is an extrapolation.

The advice is right and the reason usually given for it is not. The confirmation run is not there to guard against a mistake, an arithmetic slip or a model that failed to fit. It is there because the prediction it checks is systematically too high — not on unlucky studies, on average, by an amount the design knows in advance and nobody computes.

What a confirmation run at the chosen setting would find. The true optimum is worth 62.348. At σ = 2 the fit predicts 62.679 at the setting it recommends and the truth there is 61.821 — a gap of 0.858, which is 0.72 of the prediction's own standard error. The setting itself gives up 0.527 against the best available.
Fig. 1 What the fit predicts at the setting it recommends, against what is actually there, at five noise levels. At twice the noise the prediction is 62.679 and the truth at that setting is 61.821 — a gap of 0.858, which is 0.72 of the prediction’s own standard error.

Where the gap comes from

The mechanism is one sentence and it has nothing to do with the model being wrong.

The recommended setting is x^\hat{\mathbf x}^{*}, the point where the fitted surface is highest. The fitted surface is the true surface plus estimation error, and that error varies from setting to setting. So x^\hat{\mathbf x}^{*} is not merely where the truth is highest — it is preferentially where the error happens to be positive.

The fitted height at that point is therefore the true height there plus an error that was selected for being large. Its expectation is above the truth, and the amount is the selection effect: the expected maximum of a random field over the region the search ranged across.

That is the same object as the winner’s curse in the adaptive field — the arm chosen for being ahead is ahead by more than it should be — arriving in a setting where nothing was chosen between arms and the selection is over a continuum.

The ridge, when the fitted optimum is outside the region. One fitted surface. Its stationary point is at a radius of 2.289 and the fit calls the shape a maximum. The ridge is the best setting at each radius, found by the Lagrange condition (B̂ − μI)x = −ĝ/2; the fitted response rises along it from 59.93 at the centre to 62.10 at the edge. The true optimum is at (0.4, 0.3).
Fig. 2 One study’s recommendation, from the ridge that leaves the region: the ridge running out to the region’s edge, with the point it ends at marked. The fit predicts 62.05 there. That number is the object this essay is about.

Two gaps, and they are different quantities

The figure has three levels and it is worth being clear which pair each gap is between.

The prediction against the truth at the same setting is the selection effect. It is a property of how the setting was chosen, it would be there even if the recommended setting happened to be the true optimum, and it is what a confirmation run measures.

The truth at the recommended setting against the true best is the cost of being in the wrong place. It is a property of how accurate the fit’s direction was, and a confirmation run does not measure it — the run reports the response at the setting, and nothing in the experiment says what was available elsewhere.

At twice the noise the first is 0.858 and the second is 0.527. At four times the noise they are 2.098 and 2.380, and they have crossed.

So an experimenter who runs a confirmation and finds it below the prediction has learned about the first gap and nothing about the second, which is the one that decides whether the recommendation was worth acting on.

What each rule gives up, against a true best of 60.73. The true optimum is worth 60.726. Reporting the fitted stationary point, pulled back onto the boundary where it left, gives up 0.493 at σ = 2; taking the maximum of the fitted surface over the region gives up 0.514. The two rules are within 4% of each other there and the ordering between them changes across the sweep.
Fig. 3 The second gap on its own: the response given up by being at the recommended setting rather than the best one, against the noise. The first gap and this one are both growing and they are measuring different things, which is why a confirmation run answers one and not the other.

The scale is the prediction’s own standard error

The size of the selection effect has a natural unit and using it makes the pattern legible.

The fit reports a standard error for its prediction at the recommended setting — it is σf(XX)1f\sigma\sqrt{\mathbf f'(\mathbf X'\mathbf X)^{-1}\mathbf f} with f\mathbf f the model vector there, which the design supplies. Measured in those units the gap runs 0.24, 0.45, 0.72, 0.83, 0.82 across the five noise levels.

That is a quantity between a quarter and five sixths of a standard error, and it does not grow without limit. It cannot: the selection effect is bounded by the spread of the random field over the region, and the standard error grows with the same σ\sigma, so the ratio approaches a constant.

Three quarters of a standard error is the number to carry. A fit that predicts 62.7±1.262.7 \pm 1.2 at the setting it recommends is, on average, predicting about 0.9 too much — and a confirmation run that comes in that far below is behaving exactly as it should rather than contradicting anything.

What a confirmation run at the chosen setting would find. The true optimum is worth 62.348. At σ = 2 the fit predicts 62.813 at the setting it recommends and the truth there is 61.972 — a gap of 0.841, which is 0.70 of the prediction's own standard error. The setting itself gives up 0.376 against the best available.
Fig. 4 The same measurement with the recommended setting chosen by the maximum of the fitted surface over the region rather than by the stationary point. The selection effect is slightly larger — 0.841 at twice the noise and 2.588 at four times — and the setting itself is better: it gives up 0.376 rather than 0.527.

The comparison between the two rules is the one useful design consequence here. The careful rule lands at better settings and is more over-optimistic about them, because it maximises over a larger set — the whole region rather than one point per fit — so the maximum of the error field it selects on is larger.

Choosing better makes the prediction worse. That is the general form of the selection effect and it is worth stating separately from the arithmetic: any procedure that searches harder for the best reported value reports a more inflated one, whatever the searching is over.

The flattening is worth one more sentence because it is not obvious. The selection effect is the expected maximum of the prediction error over the set the search ranged across, and that maximum grows linearly in σ\sigma — every realisation of the error field scales together. The standard error grows linearly in σ\sigma too. So the ratio is asymptotically constant and the only reason it moves at all across the sweep is that the set the search ranges over changes with the noise: at low noise the recommended setting is nearly always the same place, and at high noise it wanders, which is a larger set to take a maximum over.

That reading also says where the ratio would be larger. A design whose prediction variance is very uneven across the region gives the search more room to find a favourable error, so the selection effect is larger relative to the standard error at the chosen point. A rotatable design, whose precision depends only on distance, is the arrangement that gives the search least to exploit — which is a property this field chose rotatability for other reasons and gets here as a side effect.

What a confirmation run can and cannot settle

Three questions get asked of a confirmation run and it answers one of them.

Is the response at this setting what was predicted? Answered directly, and the answer is usually no, by about three quarters of a standard error. A confirmation run that matches the prediction is mildly surprising.

Is this setting better than where the process was? Answered, if the previous operating point’s response is known — which it usually is, because the experiment was centred there. This is the question the run is actually for.

Is this the best setting available? Not answered at all. The run measures one point. The response elsewhere is not observed, the fit’s opinion about it is the thing being checked, and a single observation at one setting carries no information about a surface.

The third is the question the recommendation was about, and it is the one a confirmation run cannot touch. What would answer it is another experiment centred at the new setting, which is what the steepest-ascent discipline does as a matter of course and what response-surface practice does only when the confirmation run disappoints.

What the fit calls the shape, against what it is. One eigenvalue held at −3 and the other swept from −2 to 2, so the truth is a maximum on the left and a saddle on the right and the change happens at exactly zero. At an eigenvalue of −0.25 — a genuine maximum — the fit reports a saddle on 40.1% of studies; at +0.25 — a genuine saddle — it reports a maximum on 34.6%. The standard error of a squared coefficient under this design is 0.7583, and the region of confusion is about that wide either side of zero.
Fig. 5 Why the setting wanders in the first place, from the sign the curvature has: at twice the noise the fitted shape is read correctly on barely more than half of studies near a ridge, and the studies that read it wrongly are the studies whose stationary point is furthest away.

What to predict instead

The selection effect is estimable before the confirmation run, which means the prediction can be corrected rather than merely disbelieved.

The gap is about three quarters of the prediction’s standard error at this design and these noise levels, and the standard error is computed by the software already. Subtracting three quarters of it from the predicted response gives a number a confirmation run would match on average.

That is a crude correction and it is better than none, and the honest caveats are two. The factor is not a constant — it runs from 0.24 to 0.83 across the noise levels here, rising as the fit gets noisier and flattening — so a single multiplier is an approximation. And it is specific to this design and this family of truths; a different design ranges its search over a different field.

What generalises is the sign and the scale. The prediction at a chosen optimum is too high, by a fraction of a standard error rather than by a multiple of one, and an experimenter expecting a confirmation run to land on the prediction is expecting the wrong thing.

400 fits of the same surface, σ = 2. Each point is the stationary point of one fitted quadratic, from one central composite design run on the same true surface — which has a maximum at (0.91, 0.35), marked. 2% of the fits are saddles rather than maxima, so their stationary point is not an optimum of anything, and 26% land outside the region the design explored. The median distance from the centre is 1.07 against a true 0.98.
Fig. 6 Where the recommended setting itself lands, which is the other half of the accounting. The selection effect is about the height reported at the chosen point; this is about the point, and the two combine into the response an experimenter actually obtains.

The same effect, four fields over

The selection effect has now been measured in five places and the arithmetic has been the same every time, so the general statement is worth extracting.

A quantity is estimated at many candidates, the best is chosen, and the estimate at the chosen one is reported. An arm chosen for being ahead is ahead by more than it should be. The smallest of twenty p-values is not a p-value. A model chosen by a criterion fits its own selection sample better than it fits anything else. A setting chosen for being the fitted optimum has an inflated fitted height.

The size is always the expected maximum of a noise field over the set searched, which means it grows with the number of effective candidates and with the noise, and is bounded by the field’s own spread. And the repair is always one of three: condition on the choice, estimate from data the choice was not made on, or shrink by the amount the selection is worth.

What is unusual here is only that the candidate set is a continuum rather than a list, so nobody experiences it as a choice. There is no moment at which twenty options were narrowed to one; there is a formula that returns a point. The selection is inside the arithmetic, and that is why response surfaces are the one place the effect is missing from the standard account.

Why the interval does not cover it either

The fit draws an interval around its prediction, and the natural next question is whether the truth at the recommended setting falls inside it.

It falls below the interval’s lower end on 7.0% of studies at twice the noise — against the 2.5% a correctly centred 95% interval would give — and on 7.8% at three times. The asymmetry is the selection effect again: the interval is centred on an inflated prediction, so it misses low far more often than it misses high.

Seven per cent is not a catastrophe and the direction is the whole point. A 95% prediction interval that missed symmetrically would be a nuisance; one that misses low three times as often as it should is telling an experimenter something systematic, and what it is telling them is that the centre is wrong rather than that the width is.

What it costs to ignore

Putting the two gaps together gives the number an experimenter is actually exposed to, which neither alone reports.

At twice the noise the fit predicts 62.679, the truth at the recommended setting is 61.821, and the best available is 62.348. So the recommendation is 0.527 below the best available and the expectation set for it is 0.858 above what it delivers. An experimenter who acts on the recommendation and expects the prediction is disappointed by 0.858, of which 0.331 is the setting being wrong and 0.527… which does not add, because the two are measured against different baselines and the decomposition is the one from the third section.

Stated without arithmetic: the process improves less than the experiment said it would, for two reasons that are easy to confuse, and only one of them is visible from the confirmation run. A team that attributes the whole disappointment to the setting being wrong will run another experiment somewhere else; a team that attributes it all to the prediction being inflated will keep the setting. Both are partly right at twice the noise and the first is right at four times, where the setting’s own shortfall has overtaken the selection effect.

The two gaps cross, and where they cross depends on the noise — which is the reason to measure both rather than to adopt a rule about which matters.

There is a cheap diagnostic that separates them and it needs one extra run. Measure the response at the centre of the design as well as at the recommended setting. The centre’s response is known from the experiment already, so comparing the two says how much the process actually improved, and that number is free of the selection effect entirely: the centre was not chosen for anything. An improvement that is real but smaller than promised is the signature of an inflated prediction; no improvement at all is the signature of a setting in the wrong place.

What is claimed here, and what is not

The claim is what a confirmation run at a fitted optimum finds: that the fit predicts more than is there at every noise level, by 0.858 at twice the noise and 2.098 at four times; that measured in the prediction’s own standard error the gap runs from 0.24 to 0.83 and flattens rather than growing; that the rule which searches over the whole region rather than one point per fit lands at better settings and is more over-optimistic about them; and that the truth at the recommended setting falls below the fit’s own 95% interval on 7.0% of studies rather than 2.5%.

Every number is two and a half to four thousand studies with the same thirteen-run composite design, on a truth whose optimum is inside the region.

What stays out: a correction derived rather than measured, which would need the expected maximum of the prediction-error field over the region and is the kind of quantity the extremes field computes; the confirmation run’s own noise, which this essay holds separate by comparing against the truth rather than against a simulated observation; and sequential confirmation, where a disappointing run triggers another experiment and the selection compounds across stages.

The prediction interval quoted here is the fit’s ordinary one and takes the estimated coefficients as known, so its shortfall is partly the same parameter-uncertainty gap the forecast field measures. The asymmetry is what this essay’s claim rests on and it is not explained by that.

Still open: an optimum chosen and then reported

The inflation is measured above and not removed. What would remove it is what adaptive designs already use: an estimate that conditions on the choice, or one computed from data the choice was not made on.

The second is the cheaper and the one this design could support. Splitting the runs — fitting the surface on some and estimating the height at the chosen setting from the rest — gives an unbiased prediction at the cost of a noisier surface, and whether that trade comes out positive at thirteen runs is a measurement nobody here has made. It is the same trade a criterion that predicts a hold-out makes when it decides whether to spend data on being checked, arriving in a field with no hold-out tradition at all.

The check, and the refusal

Three claims are gated. That the fit predicts more at the setting it chose than is there, at every noise level — an inequality rather than a value, because the value depends on the truth and the inequality does not. That the gap grows with the noise by more than a factor of three across the sweep, which is what makes it a selection effect rather than a constant offset. And that the response actually obtained falls away as the noise grows, which is the second gap moving and is what stops the first being read as the whole story.

The refusal is the one that keeps the two gaps apart: the measurement is taken against the truth at the chosen setting rather than against a simulated confirmation observation. A comparison against a noisy observation would have the right expectation and would let the selection effect hide inside the observation’s own error, which is exactly how the effect stays invisible in practice — a single confirmation run that comes in low is indistinguishable from bad luck, and only the average over many studies shows that low is where it comes in.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Central composite designExperimental designForecast intervalMonte CarloPrediction varianceResponse-surfaceRidge analysisSelection biasStationary pointThe winner's curse