Concept

Central composite design — where it appears

A factorial arrangement with axial points and centre points added, so that a quadratic model in several factors can be fitted from a modest number of runs. It is a catalogue design, and comparing it against one chosen by optimising a criterion directly is the question the optimal-design field asks.

Named by 9 essays across 3 fields — each of them below, with the objects they name alongside it.

Where a D-optimal design puts its runs. The D-optimal measure over 121 candidate settings on a square region. It keeps 9 of them and discards the rest, and the 9 it keeps are the settings a catalogue would have offered without any of this arithmetic. What the search adds is the weights: 0.1458, 0.0962, 0.0802, which nine equal runs cannot express.

A design is a number

A standard design is taken from a catalogue and then measured. Turn the arithmetic round and a design becomes the answer to an optimisation — and over 121 candidate settings the search keeps nine of them, which are exactly the nine a catalogue would have offered, at weights nine equal runs cannot express.

optimality · Criterion
The variance touches p and never crosses it. d(x) = f(x)′M⁻¹f(x) along the diagonal of a square region, for the D-optimal measure. The line at 6 is the number of parameters in the model. Kiefer and Wolfowitz's theorem says a design is D-optimal exactly when the largest d anywhere in the region is p — not approximately, equals — so the optimal curve is tangent to that line at its support points and below it everywhere else. Here the largest value anywhere on a 41×41 grid is 6.000000000.

The theorem that says when to stop

A search that maximises the volume of the information has no way of knowing it has finished, because nothing tells it what the maximum is. Kiefer and Wolfowitz's equality does — a design is D-optimal exactly when the worst prediction anywhere in the region equals the number of parameters, which is 6.000000000059 here, gated at machine precision.

optimality · Equivalence
Four criteria on 5 designs, on a square region. Each design scored as an efficiency — its value over the best attainable — so four criteria in as many different units sit on one scale where 1 is the optimum. Rows are ordered by D. D picks 13-run exchange; A picks face-centred composite; G picks 13-run exchange; I picks face-centred composite. Every design has been scaled to just fit the region first, because a design run at settings the region does not contain is not a competitor on it. The disagreement is the point: the letter is a choice, and it is almost never reported as one.

Four letters and two camps

D, A, G and I are four ways of turning one matrix into one number, and they do not agree. The design that wins on D is the worst thing here on I. And the same two designs swap places entirely when the region changes from a square to a disc — on all four criteria at once.

optimality · Criterion
A central composite design, 13 runs. Adding 4 axial runs at ±√2 gives every factor three levels, which is the least that can estimate a squared term. The normal matrix now inverts, so each βᵢᵢ has an estimate of its own — and at exactly this axial distance the design is rotatable, which the next figure measures.

Three levels, and the ring where the design says the same thing

A central composite design puts its axial runs at ±α, and α is not a matter of taste. At F to the quarter the prediction variance depends only on how far a point is from the centre and not at all on which direction it lies in — a property with no simulation in it, exact or absent.

surface · Factorial
Where the maximum is, from 15 runs. One dataset, one fitted quadratic, and two answers to "where is the best setting". The delta method reports 0.80 ± 0.46, a finite interval it will report whatever the data does. Fieller's set is 0.49 to 1.76, because the curvature here has t = -4.04. The true optimum is at 0.75.

The optimum is a ratio, and its interval is sometimes the whole line

The best setting is −b₁/2b₂: a ratio of two estimates whose denominator is a curvature the design can often barely see. The delta method reports a finite interval every time and covers 68.8% where the curvature is weak; Fieller's set covers 95% and says so by being unbounded.

surface · Optimum
What the fit calls the shape, against what it is. One eigenvalue held at −3 and the other swept from −2 to 2, so the truth is a maximum on the left and a saddle on the right and the change happens at exactly zero. At an eigenvalue of −0.25 — a genuine maximum — the fit reports a saddle on 26.4% of studies; at +0.25 — a genuine saddle — it reports a maximum on 25.1%. The standard error of a squared coefficient under this design is 0.3791, and the region of confusion is about that wide either side of zero.

The sign the curvature has

A fitted surface reports a maximum, a minimum or a saddle, and the report is a comparison of two estimated eigenvalues against zero. At a true second eigenvalue of −0.25 the fit calls a genuine maximum a saddle on 26.4% of studies, and at +0.25 it calls a genuine saddle a maximum on 25.1%.

surface · Optimum
The Box–Behnken design in three factors: 15 runs, none at a corner. Twelve runs at the midpoints of the cube's edges and 3 at its centre. Every run holds one factor at zero, so no run puts all three factors at an extreme — which is what makes it runnable where a corner is not. The three panels are the design's coordinate projections, with repeated positions marked.

The design that refuses the corners

Box–Behnken runs three factors in fifteen runs and puts none of them at a corner, which is what makes it usable where a corner cannot be run. It predicts the corner 1.84 times worse than the seventeen-run design that goes there, and 1.31 times worse at the middle of a face, and all three numbers are matrix computations with no simulation in them.

design · Factorial
What one lost run costs a 16-run factorial fitting 11 coefficients. Every run is worth the same: dropping any one multiplies every coefficient's variance by 1.2000, which is 1 + 1/(N − p) with N = 16 and p = 11, and gives every pair of coefficients a correlation of 0.1667 where the complete design had exactly zero.

The run that did not happen

Lose one run from any orthogonal design and every coefficient's variance is multiplied by exactly 1 + 1/(N − p), and every pair of coefficients acquires a correlation of exactly 1/(N − p + 1) where there was none. The price is set by the design's spare capacity and by nothing else, and a saturated design cannot survive it at all.

design · Factorial
What a confirmation run at the chosen setting would find. The true optimum is worth 62.348. At σ = 2 the fit predicts 62.679 at the setting it recommends and the truth there is 61.821 — a gap of 0.858, which is 0.72 of the prediction's own standard error. The setting itself gives up 0.527 against the best available.

The run that confirms it

The setting a response-surface analysis recommends was chosen because the fitted surface was highest there, so the height the fit predicts at it is a maximum over a random field. At twice the noise the fit predicts 0.858 more than is there — 0.72 of the prediction's own standard error — and the gap is not noise, it is the selection.

surface · Optimum

Named alongside it

The objects these essays reach for when they reach for this one.

Experimental designPrediction varianceCurvatureFactorial designLeast squaresResponse-surfaceD-optimalityInformation matrixOptimal designRotatabilityStationary pointBox behnken design

All concepts