The surface between the corners

The height at the chosen setting

The height a response-surface fit predicts at the setting it recommends reads high, because the setting was chosen where the fit was highest. The obvious repair — choose on some runs, estimate on the rest — is impossible in the usual thirteen-run design, which has nine distinct settings where two separate fits need twelve. A parametric bootstrap of the optimism removes four fifths of it at no cost in runs, 0.884 down to 0.183 at twice the noise. With the design run twice, correcting the full fit beats splitting it: an error of 0.924 against 1.213, at a better setting.

Worth reading first: A design is a number · A break that was looked for.

The run that confirms it measured why a response-surface analysis should end with a confirmation experiment. The setting it recommends is where the fitted surface is highest, and the fitted surface is the true one plus an error that varies from setting to setting, so the recommended setting is preferentially where that error happens to be positive. The height the fit predicts there is a maximum over a random field, and on average it reads high — by about three quarters of the prediction’s own standard error at twice the noise. That essay measured the inflation and did not remove it. It named two ways to: an estimate that conditions on the choice, or an estimate from data the choice was not made on.

Both are measured here, on the same surface and the same thirteen-run central composite design: four factorial corners, four axial points at ±1.414 and five centre runs, fitting the six coefficients of a second-order surface in two factors. Every repair is scored by the bias of the height it reports at the chosen setting, against the true height there, and by its root mean squared error, over 2,000 studies at each of three noise levels.

Why thirteen runs cannot be split

The second repair is the natural one and it fails before any data exist. Splitting the runs means fitting the surface on one set to choose the setting and fitting it again on the other set to estimate the height there, and each fit needs the six coefficients to be estimable from its own runs.

The thirteen-run central composite design's nine distinct settings, and why its runs cannot be split into two fittable halves. Four factorial corners, four axial points at ±1.414 and five runs at the centre: 9 distinct settings. A second-order surface in two factors has six coefficients and needs six distinct settings, so two disjoint sets that each fit it need twelve. Of the 1716 ways to split the thirteen runs into two sets of at least six, 0 leave both able to fit the surface.
Fig. 1 The thirteen-run central composite’s settings: four corners, four axial points and five runs at the centre. Nine distinct settings, where one second-order fit needs six and two separate fits need twelve. None of the 1,716 ways to split the runs works.

A second-order surface in two factors has six coefficients, so it needs six distinct settings; replicating a setting adds precision but not a new equation. The composite design has nine distinct settings, because its five centre runs are one setting five times. Two disjoint sets that each fit the surface would need twelve. Visiting every way of dividing the thirteen runs into two sets of at least six confirms it: of 1,716 splits, 0 leave both sets able to fit the surface. The design was built to estimate one surface efficiently, and that efficiency is exactly what leaves nothing to spare for a second.

The same arithmetic holds for every composite design in common use, because the count of distinct settings grows more slowly than twice the count of coefficients. A second-order surface in k factors has (k+1)(k+2)/2(k+1)(k+2)/2 coefficients; a composite design on the full factorial has 2k+2k+12^k + 2k + 1 distinct settings. At three factors that is fifteen settings against the twenty two separate fits would need; at four, twenty-five against thirty; at five, on a half fraction, twenty-seven against forty-two. A three-level factorial does better at three factors — its twenty-seven settings are not ruled out by the count, though whether a particular division of them leaves two good designs is a separate question — and no better at two, where it has nine settings, as the composite does. Replicated centre runs, which every composite design carries for an estimate of pure error, add runs and no settings, so no number of them makes the design splittable.

So at thirteen runs the hold-out repair is not a trade to be weighed. It is unavailable, and the trade a criterion that predicts a hold-out weighs in regression — spend data on being checked — has to be bought with new runs or replaced by arithmetic.

Removing the optimism by bootstrapping it

The arithmetic is a parametric bootstrap of the selection itself. Take the fitted surface as if it were the truth, simulate the thirteen responses again from it with the fit’s own residual noise, refit, choose the setting again as the analysis did, and record how much the refitted surface’s height at its chosen setting exceeds the fitted surface’s height there. Two hundred such refits estimate the optimism a study of this design and this noise suffers, and subtracting it from the reported height corrects the report. It needs no extra runs and uses only what the analysis already has.

Seven ways to report the height at the setting a response-surface analysis recommends, at noise 2Bias and root mean squared error of the reported height about the true height at the chosen setting, over 2,000 studies with a central composite design of thirteen runs, or the same design run twice. 13 runs, the fit's own height: bias 0.884, error 1.537. 13 runs, optimism bootstrapped away: bias 0.183, error 1.396. One confirmation run alone: bias -0.005, error 2.005. Corrected fit and one confirmation run: bias 0.056, error 1.121. 26 runs, the fit's own height: bias 0.475, error 0.998. 26 runs, optimism bootstrapped away: bias 0.050, error 0.924. 26 runs split: choose on 13, estimate on 13: bias -0.014, error 1.213.0.00.51.01.52.013 runs, the fit's own height13 runs, optimism bootstrapped awayone confirmation run alonecorrected fit and one confirmation run26 runs, the fit's own height26 runs, optimism bootstrapped away26 runs split: choose on 13, estimate on 13upper bar: bias · thin bar: root mean squared error2,000 studies; 200 bootstrap refits eachthe fit's own height reads high
Fig. 2 Seven ways to report the height at the recommended setting, at twice the noise: the bias and the root mean squared error of each, about the true height there. The fit’s own height reads high at thirteen runs and still at twenty-six; the bootstrap removes most of it; a split of the twenty-six runs removes all of it at a cost in error.

At twice the noise the fit’s own height at its chosen setting reads 0.884 high — the earlier essay’s selection effect on these studies — with a root mean squared error of 1.537. The bootstrap-corrected height reads 0.183 high, with an error of 1.396: about four fifths of the optimism removed, and the error lower because the bias was the larger part of it. At unit noise the correction takes the bias from 0.259 to 0.026, nine tenths of it; at four times the noise from 2.150 to 0.913, a little under three fifths.

How much the reported height at the chosen setting reads high, by the noise, before and after the bootstrap correction. 13 runs, fit's height: 0.259, 0.884, 2.150. 13 runs, bootstrapped: 0.026, 0.183, 0.913. 26 runs, fit's height: 0.124, 0.475, 1.466. 26 runs, bootstrapped: 0.001, 0.050, 0.440 — at noise 1, 2 and 4.
Fig. 3 The bias of the reported height against the noise, before and after the bootstrap correction, at thirteen runs and at twenty-six. The correction removes most of the optimism at every noise level and less of it the noisier the study.

The correction leaves some of the optimism behind, and more of it the noisier the study, for a reason the bootstrap cannot avoid. It simulates from the fitted surface, whose optimum is itself a selected point and whose curvature is itself an estimate; at high noise the fitted surface is flatter or more sharply peaked than the truth in ways that change how much selection it would produce, and the bootstrap reproduces the selection of the fitted world rather than the true one. It is the same limit the estimate after the choice met in adaptive designs: a correction for selection is computed from a model of the selection, and the model is estimated from the data that were selected.

What the bootstrap knows about one study

The correction is an average, and it is worth seeing how much of an average. At twice the noise the bootstrap’s estimate of a study’s optimism is 0.701 on average, against the 0.884 the studies actually have, and it varies from study to study with a standard deviation of 0.353. A reader might hope that variation tracks the study’s own luck — that a study whose fitted peak happens to sit on an especially large error gets an especially large correction. It does not. The correlation between the bootstrap’s estimate and the optimism each study actually had is −0.24: slightly negative, and far from useful.

The reason is what each quantity depends on. A study’s actual optimism is the error at its chosen setting, which the study cannot see. The bootstrap’s estimate is a property of the fitted surface’s shape — how flat its peak is and how far its stationary point sits from the centre — and a study whose noise happened to raise one corner of the surface gets a surface that is sharper there, whose own selection effect, simulated, is smaller. So the correction moves every study down by roughly the same amount, removing the bias the studies share, and leaves each study’s own error untouched. That is why the error falls only from 1.537 to 1.396 while the bias falls by four fifths: the bias was the shared part of the error, and the bootstrap can only see what is shared.

It is the same reason two routes to every number insists on two independent ways of computing a quantity before trusting either: a correction that is an average is checked by averaging, and a per-study claim needs per-study information, which here only new runs supply.

What a confirmation run is worth

The textbook answer is to run the chosen setting again. A confirmation run’s response is an unbiased measurement of the true height there, and the question is what it costs.

What confirmation runs at the recommended setting are worth, alone and combined with the bias-corrected fit, at noise 2. Root mean squared error of the reported height. Confirmation runs alone: 2.005, 1.381, 0.984. Combined by precision with the bootstrap-corrected fit: 1.121, 0.952, 0.785 — at 1, 2, 4 runs. The fit's own height errs by 1.537 and the corrected fit by 1.396.
Fig. 4 The error of the reported height at twice the noise, against the number of confirmation runs at the chosen setting: the confirmation runs’ mean alone, and the confirmation runs combined with the bootstrap-corrected fit by their precisions.

One confirmation run alone is unbiased and noisier than the biased fit: its error is 2.005 against the fit’s 1.537, because a single run carries the full noise of one measurement and the fit averages thirteen. Two runs alone err by 1.381 and four by 0.984. The confirmation run’s value is that it is unbiased, and on its own a single one buys that at a cost in accuracy.

Combined with the corrected fit, weighting each by its precision, one confirmation run brings the error to 1.121 and the bias to 0.056: better than either source alone, because the fit supplies precision and the run supplies the unbiasedness the corrected fit still lacks. Two runs bring the error to 0.952 and four to 0.785. So the confirmation run’s best use is not as a replacement for the fit’s prediction but as a check that the combination absorbs, and the bootstrap correction is what makes the fit safe to combine with it.

Splitting a design run twice

The split becomes possible when the design is run twice. Twenty-six runs give two copies of the thirteen, and the split chooses the setting from one copy and estimates the height there from a fit to the other. It is unbiased: −0.014 at twice the noise, within its counting error of zero.

But it is not the best use of twenty-six runs. Fitting all of them together and reporting the fit’s own height leaves an optimism of 0.475 — a little over half the thirteen-run figure — with an error of 0.998. Bootstrapping that optimism away leaves a bias of 0.050 and an error of 0.924. The split’s error is 1.213: unbiased, and about a third worse than the corrected full fit, because each half fits the surface with half the data.

And the split pays twice. It chooses the setting with thirteen runs rather than twenty-six, so the setting it recommends is worse: the true height there falls short of the true optimum by 0.543 on average, the same as the thirteen-run design’s, against 0.251 when all twenty-six runs choose. A split of a doubled design buys an unbiased report of the height at a worse setting, estimated less precisely than the corrected full fit estimates the height at a better one.

What the extra runs buy in the setting

The doubled design’s runs can be spent on two things — choosing the setting and measuring the height there — and the measurements separate them. Used to choose, thirteen more runs halve the shortfall of the recommended setting at every noise level: from 0.125 to 0.067 at unit noise, from 0.543 to 0.251 at twice the noise, and from 2.263 to 1.284 at four times. Used to measure, the same thirteen runs give an unbiased height at the worse setting with an error of 1.213, where the full fit corrected by the bootstrap reports the height at the better setting with an error of 0.924.

The two uses are not equally valuable. An experimenter running a process at the recommended setting lives with the setting’s true height, and a better setting is worth having whether or not its height is reported exactly; an unbiased report of a worse setting’s height is worth less than a slightly biased report of a better one. The arithmetic and the purpose point the same way: every run should help choose, and the selection it causes should be removed by computation, with new runs placed at the chosen setting as the check rather than the estimate.

What the comparison says about splitting

The split is the cleanest repair conceptually, and the one the winner’s curse recommends in its simplest form: let one part of the data choose and another part measure. In a designed experiment its cost is unusually high, for two reasons that compound. A response surface needs every coefficient from every half, so splitting halves the information about each coefficient twice over — once for choosing and once for estimating — rather than spending a small hold-out on checking. And the design’s runs are not interchangeable observations but settings, so a split of a design that was efficient whole leaves halves that are inefficient or, at thirteen runs, unable to fit at all.

The bootstrap avoids both. It uses every run to choose and every run to estimate, and pays for the selection by computing it rather than by holding data back. Its failure is that it leaves a residual bias that grows with the noise. At the noise levels a response-surface study is usually run at, that residual is a fraction of the prediction’s standard error, and the corrected full fit is the most accurate report of the height measured here.

What a recommendation from a fitted surface should report

The height at the chosen setting, with its optimism removed or stated. The fit’s own height reads high by a predictable amount, and a report that gives it uncorrected is overstating the process’s best performance by about three quarters of a standard error at moderate noise.

How the correction was made. A parametric bootstrap of the selection, using the analysis’s own choice rule, removes four fifths of the optimism at twice the noise and costs nothing but computation. Its residual should be stated with it.

Confirmation runs, combined rather than substituted. A single confirmation run is too noisy to replace the fit and is the right complement to it: together with the corrected fit it gives the lowest error measured here for one extra run.

Not a split, unless the design was built for one. A thirteen-run composite cannot be split, and a design run twice is better used whole with a bootstrap correction than divided.

Counted, on what

Two thousand simulated studies at each noise level, each with its own thirteen responses from the surface the earlier essay used — 60+4x1+3x2−2x12−x1x2−3x2260 + 4x_1 + 3x_2 - 2x_1^2 - x_1x_2 - 3x_2^2 — and its own second set of thirteen for the doubled design, one stated seed per study. Each bootstrap correction uses two hundred refits drawn from the fitted surface with the fit’s residual standard deviation. The recommended setting is the fitted stationary point, pulled back onto the region’s edge at a radius of 1.414 when it falls outside, which is the rule the earlier essay called “stationary”. The confirmation runs are drawn at the chosen setting with the same noise. The split count is exhaustive over every division of the thirteen runs into sets of at least six. With two thousand studies the bias of each estimate has a standard error of about 0.03 at twice the noise, so the split’s small bias is within its counting error, and the corrected doubled fit’s 0.050 and the thirteen-run correction’s 0.183 are not.

Still open: a design that is built to be split

The comparison was between designs built to estimate one surface, and the split lost because the halves were not designs in their own right. A design could be built the other way: two smaller second-order designs, each able to fit the surface alone, run as the two halves of one experiment — for instance two composites with different axial distances, or a composite and a Box–Behnken, whose union is a better whole design than either copy of one.

Whether such a pair, split, beats the same number of runs used whole with a bootstrap correction — and whether the answer changes when the true optimum lies near the edge of the region, where the best setting is outside it and the selection effect is concentrated on a ridge — is the measurement this leaves. The same bootstrap makes it directly; what it needs is a choice of designs that each fit alone and fit well together, and the design that refuses the corners is one obvious candidate for a half, since it fits the surface in fifteen runs without visiting the corners a composite needs.

The other open question is the residual the bootstrap leaves at high noise, nearly a full unit at four times the noise. A second level of bootstrap — estimating the bootstrap’s own shortfall by repeating it inside each simulated study — is the standard way to remove a bootstrap correction’s bias, and it costs the square of the computation. Whether it removes the residual here, or whether the residual comes from the fitted surface’s shape being wrong in a way no amount of resampling from it can detect, is not measured.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

BootstrapCentral composite designConfirmation runExperimental designMean squared errorParametric bootstrapPrediction varianceResponse-surfaceSample splittingSelection effectStationary pointThe winner's curse