A design chosen rather than looked up

How many places a design goes

Carathéodory's bound puts an optimal design's support between six and twenty-one settings, and every design in this field that can fit the model visits exactly nine. The count is not a choice anybody makes, it decides how many degrees of freedom are left to check the model with, and the first spare setting costs six points of efficiency to get back.

Worth reading first: The theorem that says when to stop.

The theorem that says when to stop certifies a design by a condition on its prediction variance. What it does not say, and what the search returns without being asked, is how many distinct settings the certified design actually visits.

The answer here is nine, out of a candidate grid of a hundred and twenty-one. Nobody chose nine.

What visiting fewer settings costs, 6 parameters. Carathéodory's bound puts the support of an optimal measure between 6 and 21. 6 settings: D-efficiency 88.90%, G-efficiency 57.18%, 0 degrees of freedom for lack of fit; 7 settings: D-efficiency 94.54%, G-efficiency 61.22%, 1 degrees of freedom for lack of fit; 8 settings: D-efficiency 95.99%, G-efficiency 64.60%, 2 degrees of freedom for lack of fit; 9 settings: D-efficiency 97.40%, G-efficiency 82.76%, 3 degrees of freedom for lack of fit. The saturated design has none, and buying the first one costs about five points of efficiency to get back.
Fig. 1 What a design gives up by visiting fewer settings, from the saturated design at six up to the optimum’s own nine. Six settings reach 88.90% of the optimal determinant and 57.18% of its worst-case efficiency; nine reach 97.40% and 82.76%.

The bounds, and neither is chosen

Two numbers bracket the support of any optimal design measure and both come from outside the subject.

The lower bound is pp, the number of parameters. A design supported at fewer settings than it has parameters gives a singular information matrix — the model vectors at kk points span a space of dimension at most kk, and the matrix is a sum of kk rank-one terms, so its rank is at most kk. Every criterion in this field is a function of the inverse, and there is no inverse.

The upper bound is p(p+1)/2p(p+1)/2, which is Carathéodory’s and is a fact about convex hulls. The information matrix is symmetric p×pp\times p, so it lives in a space of dimension p(p+1)/2p(p+1)/2; any point in the convex hull of a set in a dd-dimensional space can be written using at most d+1d+1 of them, and the constraint that the weights sum to one removes one. For the full quadratic in two factors that is twenty-one.

So an optimal design in this field has between six and twenty-one support points. The search returns nine, and it returns nine for all three criteria.

Distinct settings, against a bound from 6 to 21. D-optimal measure: 9 distinct settings, 3 degrees of freedom for lack of fit; A-optimal measure: 9 distinct settings, 3 degrees of freedom for lack of fit; I-optimal measure: 9 distinct settings, 3 degrees of freedom for lack of fit; 3² factorial: 9 distinct settings, 3 degrees of freedom for lack of fit; face-centred composite: 9 distinct settings, 3 degrees of freedom for lack of fit; rotatable composite: 9 distinct settings, 3 degrees of freedom for lack of fit; factorial with centres: 5 distinct settings, -1 degrees of freedom for lack of fit. Carathéodory's bound allows anything from 6 to 21 and every design that can fit the model lands on the same number.
Fig. 2 Every design this field has scored, by how many distinct settings it visits. Three optimal measures and three catalogue designs all land on nine. The fourth catalogue design visits five, which is below pp, and cannot fit the model at all — which is why its every criterion has been zero in every table this field has drawn.

What the count decides

The number is not a curiosity, because it is the only thing standing between a design and its own model.

A design supported at kk distinct settings can compute, at each setting, the average of the responses observed there — an estimate of the true mean at that setting that assumes no model at all. The fitted surface also predicts each of those kk means. Comparing the two is a lack-of-fit test, and it has kpk - p degrees of freedom, however many runs the design has.

That is the quantity in the last column of the table. Nine settings and six parameters leave three degrees of freedom for lack of fit. Six settings leave none, and a design with none cannot be checked against anything: it fits its six parameters to six means exactly, every residual at a distinct setting is zero by construction, and the only residual variation left is replication within settings — which estimates σ\sigma and says nothing about whether the surface is the right shape.

Adding runs does not help. A saturated design replicated a hundred times has a hundred times the precision and still zero lack-of-fit degrees of freedom. The check needs distinct settings, and replication supplies none.

What the first spare setting costs

The figure above prices it, and the price is larger than it looks.

Going from six support points to seven costs nothing and gains 5.65 points of D-efficiency — 88.90% to 94.54%. That reads as though spare settings were free, and the reading is backwards: it is the saturated design that is expensive. Squeezing the support down to pp costs 11.1% of the determinant and 42.8% of the worst-case efficiency, and buys nothing at all.

The G-efficiency column is the one worth following. The saturated design’s worst prediction anywhere in the region is 10.49 against the bound of six, so its G-efficiency is 57.18%: at the region’s worst-served setting it predicts as precisely as the optimal design would on 57% of its runs. By nine support points that has risen to 82.76%.

Concentrating a design on fewer settings damages its worst case far more than its average, which is what a determinant and a minimax criterion measure respectively, and is the reason the two columns separate.

A design catalogue’s rows, re-read

The table’s rows are the designs this field has been comparing for four essays, and reading them by support count rather than by efficiency says something none of the efficiency tables did.

The 3² factorial at nine runs and the face-centred composite at thirteen visit the same nine settings. The composite design’s extra four runs are replicates of the centre, so it is the factorial with one setting repeated five times — which is exactly what this field found by another route, and which the support count says in one number rather than by comparing coordinates.

The rotatable composite at thirteen runs also visits nine settings, and they are different nine: its axial runs are at ±1.414 rather than ±1, outside the square. Same count, different places, and the count alone cannot tell them apart — which is the limit of this statistic and is worth stating beside its uses.

The factorial with centre runs visits five. Four corners and a centre, whatever the replication, and five settings cannot determine six parameters. Every criterion this field computes has been zero for that row in every table, and the reason has always been this count.

So the support count is a coarse instrument that answers one question exactly: can this design fit this model, and how much of it can be checked. It answers nothing about how well, which is what every other column is for.

Why nine, and why the same nine

The three criteria have different objectives, different derivatives and different weights, and all three return the same nine settings: the corners, the edge midpoints and the centre — which is the 323^{2} factorial.

The reason is a product structure. The model is a full quadratic in two factors and the region is a square, so both factorise: the model vector’s terms are products of one-dimensional terms, and the region is a product of two intervals. A one-dimensional quadratic on [1,1][-1, 1] has its D-optimal measure supported at three points — the two ends and the middle — and the product of two such measures is supported at the nine combinations.

That is the rediscovery this field already recorded: the search produces a design a catalogue contains, without being given it. What is new here is that the count rather than the positions is the part forced by the structure, and that the count is what the lack-of-fit test lives on.

The weights are not the same across criteria, and they are where the criteria disagree. D puts 0.0962 at the centre and A puts 0.2332 — two and a half times as much — from the same nine settings.

Where a A-optimal design puts its runs. The A-optimal measure over 121 candidate settings on a square region. It keeps 9 of them and discards the rest, and the 9 it keeps are the settings a catalogue would have offered without any of this arithmetic. What the search adds is the weights: 0.2332, 0.0978, 0.0940, which nine equal runs cannot express.
Fig. 3 The A-optimal measure’s weights on those nine settings. Same support, different distribution: a trace of an inverse is a sum over coefficients and is dominated by whichever is worst determined, while a determinant is a product and cannot let any direction go.

Six settings, and where they land

The saturated design is worth looking at rather than only pricing, because its settings are not the ones an experimenter would have guessed.

The six-point search returns three of the four corners and three settings that are not on the 3×33\times3 grid at all — (0.2,1)(0.2, 1), (1,0.6)(1, 0.6) and (0.2,0)(-0.2, 0). Only three of its six settings are settings the nine-point optimum uses.

That is not a rounding of the nine-point design. It is a different arrangement, and constraining the support changes where the support goes as well as how much of it there is. At seven settings the search comes back onto the grid entirely — seven of seven — and at eight it is seven of eight. The excursion is a property of the saturated case rather than of the search.

The reason is that with six settings and equal weights the design has to span the model’s six directions with nothing to spare, so every setting is carrying a direction on its own. The nine-point design can afford settings that are partly redundant, and redundancy is what a lack-of-fit test reads.

A design at its minimum support is a basis rather than a sample, and the two words describe different objects. Each of its settings is load-bearing, losing any one makes the matrix singular, and the previous field’s result about a lost run applies at its extreme here: the inflation is 1+1/(Np)1 + 1/(N-p) and a saturated design has Np=0N - p = 0.

The upper bound is never approached, and that is a fact about this region

Twenty-one is a long way from nine, and the gap deserves an explanation rather than a shrug.

Carathéodory’s bound is tight in the sense that designs attaining it exist, and it is loose for any particular problem because the information matrices of a structured candidate set do not fill the space they live in. Here the region is a square and the model is a product, so the optimum is a product measure and a product measure on a 3×33\times3 grid has nine support points by construction.

Change the region and the count changes. On a disc, the model no longer factorises with the region and the optimal measure is supported on a set that includes the boundary circle at several radii — which is the region-dependence this field’s last essay is about, seen in the support count rather than in the efficiency.

What does not change is the pair of bounds, because neither mentions the region. Six and twenty-one are properties of the model alone.

The variance touches p and never crosses it. d(x) = f(x)′M⁻¹f(x) along the diagonal of a square region, for the D-optimal measure. The line at 6 is the number of parameters in the model. Kiefer and Wolfowitz's theorem says a design is D-optimal exactly when the largest d anywhere in the region is p — not approximately, equals — so the optimal curve is tangent to that line at its support points and below it everywhere else. Here the largest value anywhere on a 41×41 grid is 6.000000000.
Fig. 4 The certificate for the nine-point measure, for reference: the prediction variance touches 6 at each of the support points and stays below it everywhere between. Nine tangencies is what nine support points look like on this axis, and the count is readable off the picture.
Four criteria on 5 designs, on a square region. Each design scored as an efficiency — its value over the best attainable — so four criteria in as many different units sit on one scale where 1 is the optimum. Rows are ordered by D. D picks 13-run exchange; A picks face-centred composite; G picks 13-run exchange; I picks face-centred composite. Every design has been scaled to just fit the region first, because a design run at settings the region does not contain is not a competitor on it. The disagreement is the point: the letter is a choice, and it is almost never reported as one.
Fig. 5 The same designs on the criteria this field usually ranks them by. The row that visits five settings is at zero on every one of them, and the three that visit nine are separated by fractions. The support count and the efficiencies are answering different questions about the same seven objects.

What an experimenter should take from the count

Three statements, and the third is the one that changes a decision.

Count the distinct settings, not the runs. A protocol described as “twenty runs” says nothing about whether the model can be checked. A protocol described as “six settings, replicated” says it cannot.

A saturated design is a decision to trust the model. It is a defensible decision — where the model is known from theory and the experiment is a measurement rather than an exploration, spending runs on lack of fit is spending them on a question already answered. What it is not is a default, and it is what a design chosen purely by a criterion tends towards when the criterion is evaluated on too few candidates.

And the first spare setting is the one to buy. From six to seven is 5.65 points of D-efficiency and the first lack-of-fit degree of freedom, at no cost in runs. From eight to nine is 1.41 points and the third. The returns fall off fast and the first one is nearly free, which is an unusually clean recommendation for this field.

The count is one of three numbers a design should be described by

Putting this beside the two results before it gives a short list, and the list is the practical content of the subject so far.

Runs. How much the experiment costs, and the only number most protocols give.

Distinct settings. Whether the model can be fitted at all — the count has to be at least pp — and how much of it can be checked, at kpk - p degrees of freedom.

And the certificate. The largest prediction variance anywhere in the region, divided into pp, which is a G-efficiency and needs no reference design because the equivalence theorem supplies the denominator.

Three numbers, all computable from the design’s own coordinates before anything is measured, and none of them requiring the optimal design to be found. A protocol that reports all three has said what its design is; one that reports the first has said what it costs.

What the search behind those numbers can and cannot claim

One caveat belongs in the essay rather than in a footnote, because it is the same caveat this field recorded about exact designs.

Choosing which kk settings to support is a combinatorial problem, not a convex one. The search here is point exchange from eight random starts, it stops at a local optimum, and which local optimum depends on where it started. The figures report the best of eight.

Two things make that acceptable here and both are worth stating. The nine-point answer agrees with the convex optimum’s support exactly, which is a check the search passes on the one case where the right answer is known independently. And the sequence is monotone — efficiency rises at every step — which a search that had stopped badly at one size would break.

What cannot be claimed is that the six-, seven- and eight-point designs are the best of their size. They are the best of eight starts, which is a lower bound on the best, and a better six-point design would make the saturated case look less bad rather than more. The essay’s argument survives either way, because it is about the direction of the trade rather than about its exact size.

What is claimed here, and what is not

The claim is about the support of an optimal design: that Carathéodory’s bound puts it between pp and p(p+1)/2p(p+1)/2, which is six and twenty-one for the quadratic in two factors; that all three criteria’s optima and three of the four catalogue designs visit exactly nine settings while the fourth visits five and cannot fit the model; that the number of distinct settings minus the number of parameters is the lack-of-fit degrees of freedom, however many runs the design has; and that constraining the support to six costs 11.1% of the determinant and 42.8% of the worst-case efficiency.

Every efficiency here is a ratio of determinants or of prediction variances computed from the designs themselves. The minimum-support designs are found by exchange from eight random starts, which is a local search and is reported as such.

What stays out: the lack-of-fit test’s power, which is the quantity that would say whether three degrees of freedom are enough and which needs an alternative model to be powered against — the design field measures the two-factor version at one degree of freedom and the general case is a larger question; designs on a disc, where the support count differs and the figures here would need a second region throughout; and Ds-optimality for a subset of parameters, where the relevant bound is different and the four lettered criteria supply the arithmetic.

Still open: the design that starts from runs already made

Every design in this field is chosen from nothing. The search is handed a region, a model and a run count, and it returns an arrangement as though no experiment had happened.

That is almost never the situation. A screening design has been run, a factorial has been done, four runs have been made down one edge because that is where the equipment was — and the question is where to put the next few. The equivalence theorem still applies and its statement changes in exactly one place: the number the largest prediction variance has to equal is no longer pp. It is a function of how much of the experiment is already spent, it reduces to pp when nothing is, and it equals pp again exactly when the runs already made can still be absorbed into the design that would have been chosen. That is augmenting a design that has already run.

The check, and the refusal

Three claims are gated. That a design supported at exactly pp settings has zero lack-of-fit degrees of freedom, which is a count and is the definition made checkable. That efficiency rises with every setting added, which is that monotonicity and would catch a search that had stopped at a worse local optimum on a larger support. And that the worst-case efficiency rises faster than the determinant’s across the whole sequence, which is the comparison the two columns exist to make.

The refusal is in the table: at least one design must visit too few settings to fit the model. The factorial with centre runs is required to fail, at five distinct settings against six parameters. A table in which every row could fit the model would be a table that never exercised the lower bound, and the bound would be a sentence rather than a measurement — which is the shape a check that has never rejected anything always has.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

D-optimalityDegrees of freedomDesign measureExperimental designG optimalityInformation matrixLack-of-fitOptimal designPrediction variance