Decided before the data

The run that did not happen

Lose one run from any orthogonal design and every coefficient's variance is multiplied by exactly 1 + 1/(N − p), and every pair of coefficients acquires a correlation of exactly 1/(N − p + 1) where there was none. The price is set by the design's spare capacity and by nothing else, and a saturated design cannot survive it at all.

Worth reading first: One factor at a time.

Everything in this field so far is about a design, and a design is a plan. What arrives is a dataset, and datasets lose runs: a sample is contaminated, an instrument drops out, a subject withdraws, a batch is scrapped, a technician writes down the wrong thing and the row is struck.

The question is what that costs, and it has an exact answer.

What one lost run costs a 16-run factorial fitting 11 coefficients. Every run is worth the same: dropping any one multiplies every coefficient's variance by 1.2000, which is 1 + 1/(N − p) with N = 16 and p = 11, and gives every pair of coefficients a correlation of 0.1667 where the complete design had exactly zero.
Fig. 1 A sixteen-run factorial in four factors fitting eleven coefficients — an intercept, four main effects and six two-factor interactions. Dropping any of its sixteen runs multiplies every coefficient’s variance by 1.2000, and every run gives the same number.

Why every run is worth the same

A complete factorial coded ±1 has a symmetry group that acts on its runs, and that is the whole explanation.

Flip the sign of a factor, or permute two factors, and the design maps to itself with its runs relabelled. Any run can be carried to any other run by such a map, so no run can be worth more than another: the arrangement contains no information that distinguishes them.

That is stronger than a simulation result and it is the kind of statement this field is for. It is not that the sixteen costs are close; they are the same number, to machine precision, and the claim made here is that equality rather than an agreement.

The closed form

The arithmetic behind the number is short enough to write out, and the result is more general than the design it is computed on.

For an orthogonal design with NN runs and pp coefficients, all columns ±1, the information matrix is NIN\mathbf{I} and every coefficient has variance σ2/N\sigma^{2}/N. Removing the run whose model vector is x\mathbf{x} leaves NIxxN\mathbf{I} - \mathbf{x}\mathbf{x}', and the Sherman–Morrison identity inverts that in one line. With xx=p\mathbf{x}'\mathbf{x} = p:

(NIxx)1=IN+xxN(Np)(N\mathbf{I} - \mathbf{x}\mathbf{x}')^{-1} = \frac{\mathbf{I}}{N} + \frac{\mathbf{x}\mathbf{x}'}{N(N-p)}

The diagonal gains 1/(N(Np))1/(N(N-p)) on top of 1/N1/N, so every variance is multiplied by

1+1Np1 + \frac{1}{N-p}

and the off-diagonal entries, which were exactly zero, become ±1/(N(Np))\pm 1/(N(N-p)). Dividing by the new diagonal gives a correlation of

1Np+1\frac{1}{N-p+1}

between every pair of coefficients — with signs given by the products of the lost run’s own model vector, so which pairs are positively correlated depends on which run was lost and the magnitudes do not.

At N=16N = 16 and p=11p = 11 that is 1+1/5=1.21 + 1/5 = 1.2 and 1/6=0.16671/6 = 0.1667, which is what the figure reports.

Two routes, which is the site’s habit

An algebraic result is a claim about a computation nobody ran, and a claim like that is worth counting out.

What one lost run costs, against how much room the design had. Six orthogonal designs, from a 8-run factorial fitting 7 coefficients to a 32-run one fitting 6. The counted inflation and the closed form 1 + 1/(N − p) agree to 8.9e-16 at every one of them, and the induced correlation is 1/(N − p + 1) by the same arithmetic. A saturated design, with no spare runs at all, cannot survive losing one.
Fig. 2 Six orthogonal designs, from an eight-run factorial fitting seven coefficients to a thirty-two-run one fitting six. For each, one run is removed, the information matrix is rebuilt from the remaining runs and inverted numerically, and the resulting inflation is compared with 1 + 1/(N − p). The largest disagreement over all six is 8.9 × 10⁻¹⁶.

The two routes share nothing but the design coordinates: one is a rank-one update identity applied symbolically, the other is a Gauss–Jordan inversion of an eleven-by-eleven matrix built by summing outer products. Agreeing to machine precision is evidence rather than tautology, and it is the same standard the equivalence theorem is gated at in the field next door.

The induced correlations agree the same way, which matters more than it looks: the correlation is the part of the result that a variance-only check would have missed entirely, and it is the part with the consequence.

Spare capacity is the whole of it

Reading the closed form as a statement about design rather than about algebra gives the rule.

NpN - p is the number of residual degrees of freedom — the runs the design has beyond what the model needs. The cost of losing a run is 1+1/(Np)1 + 1/(N-p), so:

design runs coefficients spare one run lost costs
2³ factorial, all interactions 8 7 1 ×2.0000
2³ factorial, main effects only 8 4 4 ×1.2500
2⁴ factorial, two-factor model 16 11 5 ×1.2000
2⁴ factorial, main effects only 16 5 11 ×1.0909
2⁵ factorial, two-factor model 32 16 16 ×1.0625
2⁵ factorial, main effects only 32 6 26 ×1.0385

Nothing in the column on the right depends on the number of factors, on which interactions are in the model, or on which run was lost. It depends on the spare capacity and on nothing else, and it falls away fast: at five spare runs a lost run costs 20%, at eleven it costs 9%, at twenty-six it costs 4%.

And at zero spare runs it costs everything. A saturated design — N=pN = p, which is the eight-run resolution-IV fraction a word of the defining relation buys — has no residual degrees of freedom, and removing a run leaves seven runs and eight coefficients. The information matrix is singular. The formula reads 1+1/01 + 1/0, and the design does not degrade: it stops being a design.

That is the practical form of the whole result. A design with no spare runs is one accident away from being unanalysable, and the number of spare runs is chosen when the design is, before anything can go wrong.

The correlation is the part that bites

The variance inflation is the number everybody looks at and the correlation is the one with the consequence, because it is not merely a loss of precision — it is a loss of the property the design was built to have.

A complete factorial is orthogonal: every off-diagonal of the inverse information matrix is exactly zero, so every coefficient can be estimated, interpreted and tested without reference to any other. That is what makes a factorial analysable by hand and reportable as a list.

Lose one run and every pair acquires a correlation of 1/(Np+1)1/(N-p+1). At the sixteen-run design that is 0.1667, which is small. What it is not is zero, and three things follow that were previously not true:

The estimates are no longer independent. A coefficient reported as 3.2 ± 0.4 is now 3.2 ± 0.44, correlated 0.17 with every other coefficient in the table.

The order the terms are dropped in matters. Sequential sums of squares depend on the order when the columns are not orthogonal, so the analysis of variance a factorial produces — which is order-independent by construction — becomes order-dependent, and two analysts dropping terms in different orders will report different tables.

And nothing says so. The design is still described as a 242^{4} factorial in the methods section, the software fits it without complaint, and the fifteen rows look like fifteen rows of a sixteen-run design.

That last is the recurring shape: the output is well-formed, every quantity in it is computed correctly, and a property the reader is relying on has quietly stopped holding. It is the same class as a fraction reporting an interaction as a main effect — there the columns coincide by construction and here they stop being independent by accident, and in both cases the analysis of variance is a valid table about something other than what its row labels say.

A composite design prices its runs differently

Everything above is about orthogonal designs, and the second-order designs this field spends most of its time on are not orthogonal.

What one lost run costs a 13-run composite design. dropping a corner run multiplies the worst variance by 1.667, on AB; dropping an axial run multiplies the worst variance by 1.667, on A; dropping a centre run multiplies the worst variance by 1.250, on I. The three kinds of run are not interchangeable.
Fig. 3 The thirteen-run central composite fitting six coefficients. Its runs come in three kinds and they cost three different amounts: losing a corner or an axial run multiplies the worst variance by 1.667, and losing a centre run multiplies it by 1.250.

The symmetry argument does not apply, because the design’s runs are not interchangeable: a corner, an axial run at ±1.414 and a centre run are three different objects and the design knows it.

The centre runs are the cheapest to lose, which is worth stating carefully because it is not what “cheapest” usually means. Losing a centre run costs 25% on the intercept and almost nothing on the squared terms; losing a corner costs 67% on the interaction. The five centre runs are the design’s slack, and a design with several of them is a design that can afford an accident — which is a use for centre runs that the uniform-precision argument does not mention and which is often the more valuable one.

There is also a run this design cannot afford to lose at all, and it is not on the figure because it does not exist here: at one centre run, losing it leaves four distinct levels in each factor and the design is still estimable, but at zero centre runs the composite design cannot separate the two squared terms from the intercept. The slack is load-bearing rather than decorative.

A central composite design, 13 runs. Adding 4 axial runs at ±√2 gives every factor three levels, which is the least that can estimate a squared term. The normal matrix now inverts, so each βᵢᵢ has an estimate of its own — and at exactly this axial distance the design is rotatable, which the next figure measures.
Fig. 4 The design being priced: four corners, four axial runs, five at the centre. The three kinds of run occupy three different distances from the middle, which is exactly why the symmetry argument that makes every factorial run equivalent has nothing to say here.

Spare capacity is not a bound once the design is not orthogonal

The closed form is a statement about orthogonal designs, and it is worth measuring how far it travels, because “residual degrees of freedom” is the quantity everybody carries and the formula makes it look like the whole story.

The Box–Behnken design in three factors: 15 runs, none at a corner. Twelve runs at the midpoints of the cube's edges and 3 at its centre. Every run holds one factor at zero, so no run puts all three factors at an extreme — which is what makes it runnable where a corner is not. The three panels are the design's coordinate projections, with repeated positions marked.
Fig. 5 The design that refuses the corners, which this section prices: fifteen runs fitting ten coefficients, so five spare — the same spare capacity as the sixteen-run factorial in the first figure, which paid ×1.2000 for a lost run.

The Box–Behnken has five spare runs and is not orthogonal: the largest correlation between its coefficients is −0.5547 before anything is lost. Lose one of its twelve edge runs and the worst variance is multiplied by 2.0000 — against the ×1.2000 the spare-capacity formula predicts for an orthogonal design with the same slack.

The central composite in three factors has seven spare runs, which is more, and pays more than the formula too: ×1.3786 for a corner, ×1.4971 for a centre run and ×1.5433 for an axial run, against a predicted ×1.1429.

So the closed form is a lower bound in practice rather than a rule, and it holds exactly for the one class of design where every column is independent of every other. The general statement is the one the algebra actually supports: the cost of a lost run is decided by how much of the design’s information that run was uniquely carrying, and orthogonality is the condition under which no run carries anything uniquely.

That also says which runs of a non-orthogonal design to protect. In the three-factor composite it is the axial runs — six of seventeen — because they are the only runs that let the three squared terms be told apart from each other, and losing one takes the design nearest to the boundary where it cannot do that at all.

What to do about it, which is a design decision

The result is about a loss that has already happened, and its use is before the data exists.

Do not read the spare-capacity formula as a guarantee. It is exact for an orthogonal design and optimistic for anything else, and the second-order designs this field spends most of its time on are all in the second class. A composite design with seven spare runs pays what an orthogonal design with two would, and the Box–Behnken with five pays what an orthogonal design with one would.

Build in spare capacity deliberately. The cost of a lost run is 1+1/(Np)1+1/(N-p), so a design with one spare run pays double and a design with five pays a fifth. If a run is likely to be lost — a long-running process, a fragile assay, human subjects — the run to add is not a replicate of the most interesting setting but any run, because the spare capacity is what is being bought.

Prefer slack that can be lost. In a composite design the centre runs are both the cheapest to lose and the ones that buy uniform precision, so adding one is doing two jobs at once. In a factorial there is no such distinction and the slack has to come from replication or from a smaller model.

And check the orthogonality rather than the design name. A dataset described as a 242^{4} factorial may not be one. The check is one line — the largest off-diagonal correlation of the inverse information matrix — and it reads exactly zero for a complete design and 1/(Np+1)1/(N-p+1) for one missing a run.

Prediction variance around a ring, three axial distances. Walked around a circle of radius 1 at 72 angles, with no simulation anywhere in it: the scaled prediction variance is a matrix computation on the design. At α = √2 it is flat to 7.6e-16 of its own value — the design says the same thing in every direction — and at α = 1 it varies by 47%, so a prediction towards a corner is worth measurably less than one along an axis.
Fig. 6 What the composite design’s precision looks like when nothing is missing: constant around a ring at α=2\alpha = \sqrt{2} to 1.1×10151.1 \times 10^{-15}, which is the rotatability the composite is built for. A lost run breaks that too, and it breaks it in a direction decided by which run was lost.

Reading the inflation back as a sample size

There is a second way to state the cost that an experimenter can act on, and it converts a multiplication into a count.

A variance multiplied by 1+1/(Np)1 + 1/(N-p) is the variance a design of N=N/(1+1/(Np))N' = N/(1+1/(N-p)) orthogonal runs would have had. So the sixteen-run factorial with eleven coefficients, having lost a run, is worth 16/1.2=13.3316/1.2 = 13.33 runs of the complete design — it has lost the equivalent of 2.67 runs, not one.

That is the number worth carrying. Losing one run costs more than one run, and how much more is the same spare-capacity ratio: the excess is 1/(Np)1/(N-p) of the whole design, so a design with five spare runs loses 2.67 and one with twenty-six loses 1.19.

The Box–Behnken is where it bites. Fifteen runs, five spare, and losing one edge run doubles the worst variance — so the fourteen remaining runs are worth 7.5 runs of the complete design. A single accident has cost half the experiment, at the design’s worst-affected coefficient, and the design still fits, still reports, and still describes itself as a Box–Behnken.

Why the result is about the design and not about the loss

One property of the closed form is easy to walk past and is the reason it is worth having: the cost does not depend on which run was lost, and for an orthogonal design it does not depend on the model’s contents either — only on how many coefficients there are.

That is unusual. Most statements about missing data are conditional on something about the missing observation: how extreme it was, where in the design it sat, whether its absence was related to the response. Here, for the class of design where the columns are independent, none of that enters the variance at all. Every run carries exactly 1/N1/N of the information about every coefficient, so losing any one removes exactly the same amount of every kind.

The consequence is that the cost can be quoted before the study runs, as a property of the plan. If a run is lost, every standard error rises by ten per cent and every pair of coefficients becomes correlated at 0.09 is a sentence an experimenter can write down at the design stage, alongside the power calculation, and act on by adding a run.

It is also why the non-orthogonal case is a different kind of statement. There the runs are not interchangeable, the cost depends on which one goes, and the only honest form of the sentence names the run: losing an axial run raises the worst standard error by 24%, losing a centre run by 22%, and losing a corner by 17%. Three numbers rather than one, and all three computable from the design alone.

What is claimed here, and what is not

The claim is what one missing run costs: that for any orthogonal design the variance inflation is exactly 1+1/(Np)1 + 1/(N-p) and the induced correlation exactly 1/(Np+1)1/(N-p+1), both verified against numerical inversion at six designs to 8.9×10168.9\times10^{-16}; that every run of a complete factorial costs the same by symmetry; that a saturated design is rendered unanalysable rather than degraded; and that a central composite prices its three kinds of run at 1.667, 1.667 and 1.250.

Nothing here is simulated. Every number is an inverse of an information matrix or a rank-one update of one.

What stays out: two or more lost runs, where the update is rank two and the cost depends on which pair in a way the rank-one case does not; the effect on the estimate rather than on its variance, which is zero in expectation for a lost run that is missing at random and is not zero when the reason for the loss is related to the response — a mechanism the missing-data field is entirely about, and which is the more serious problem in practice; and recovery designs, where the lost run is re-run and the question becomes whether the two batches are comparable.

The rank-one derivation assumes the lost run’s model vector has xx=p\mathbf{x}'\mathbf{x} = p, which is true for a ±1 design under any model whose terms are products of ±1 columns and is false for a design with a centre run in it. That is why the composite design is counted rather than derived, and why its three numbers are reported as three numbers rather than as a formula.

Still open: the loss that is not at random

Everything here treats the missing run as missing for reasons unconnected to what it would have produced — an instrument failure, a scheduling clash, a broken sample. Under that assumption the remaining estimates are unbiased and only their precision suffers, which is what the whole essay measures.

The other case is the interesting one and nothing above prices it. A run lost because the process failed at that setting, or a subject withdrawn because the treatment was intolerable, is a run whose absence is information about the response — and then the remaining fifteen rows are not a fifteen-run design with a known inflation, they are a biased sample from a sixteen-run one.

The check, and the refusal

Two claims are gated at machine precision. That every run of a complete factorial costs exactly the same, required as an equality across all sixteen rather than as an agreement within a tolerance. And that the counted inflation equals 1+1/(Np)1 + 1/(N-p) at six designs with spare capacities from one to twenty-six, with the induced correlations checked against 1/(Np+1)1/(N-p+1) by the same standard — because a check on the variances alone would have passed on an arithmetic that got every covariance wrong.

The refusal is the composite design, and it is required to fail the symmetry claim: its worst inflation must vary by more than a tenth across its runs. If every design priced its runs equally, the symmetry argument would be doing no work and the factorial result would be a fact about arithmetic rather than about orthogonality. A design that disagrees is what makes the agreement mean something, which is the standing form of every refusal in this collection.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Box behnken designCentral composite designCorrelationDegrees of freedomExperimental designFactorial designFractional factorialLeast squaresOrthogonalityStudy designVariance inflation