The run that did not happen
Worth reading first: One factor at a time.
Everything in this field so far is about a design, and a design is a plan. What arrives is a dataset, and datasets lose runs: a sample is contaminated, an instrument drops out, a subject withdraws, a batch is scrapped, a technician writes down the wrong thing and the row is struck.
The question is what that costs, and it has an exact answer.
Why every run is worth the same
A complete factorial coded ±1 has a symmetry group that acts on its runs, and that is the whole explanation.
Flip the sign of a factor, or permute two factors, and the design maps to itself with its runs relabelled. Any run can be carried to any other run by such a map, so no run can be worth more than another: the arrangement contains no information that distinguishes them.
That is stronger than a simulation result and it is the kind of statement this field is for. It is not that the sixteen costs are close; they are the same number, to machine precision, and the claim made here is that equality rather than an agreement.
The closed form
The arithmetic behind the number is short enough to write out, and the result is more general than the design it is computed on.
For an orthogonal design with runs and coefficients, all columns ±1, the information matrix is and every coefficient has variance . Removing the run whose model vector is leaves , and the Sherman–Morrison identity inverts that in one line. With :
The diagonal gains on top of , so every variance is multiplied by
and the off-diagonal entries, which were exactly zero, become . Dividing by the new diagonal gives a correlation of
between every pair of coefficients — with signs given by the products of the lost run’s own model vector, so which pairs are positively correlated depends on which run was lost and the magnitudes do not.
At and that is and , which is what the figure reports.
Two routes, which is the site’s habit
An algebraic result is a claim about a computation nobody ran, and a claim like that is worth counting out.
The two routes share nothing but the design coordinates: one is a rank-one update identity applied symbolically, the other is a Gauss–Jordan inversion of an eleven-by-eleven matrix built by summing outer products. Agreeing to machine precision is evidence rather than tautology, and it is the same standard the equivalence theorem is gated at in the field next door.
The induced correlations agree the same way, which matters more than it looks: the correlation is the part of the result that a variance-only check would have missed entirely, and it is the part with the consequence.
Spare capacity is the whole of it
Reading the closed form as a statement about design rather than about algebra gives the rule.
is the number of residual degrees of freedom — the runs the design has beyond what the model needs. The cost of losing a run is , so:
| design | runs | coefficients | spare | one run lost costs |
|---|---|---|---|---|
| 2³ factorial, all interactions | 8 | 7 | 1 | ×2.0000 |
| 2³ factorial, main effects only | 8 | 4 | 4 | ×1.2500 |
| 2⁴ factorial, two-factor model | 16 | 11 | 5 | ×1.2000 |
| 2⁴ factorial, main effects only | 16 | 5 | 11 | ×1.0909 |
| 2⁵ factorial, two-factor model | 32 | 16 | 16 | ×1.0625 |
| 2⁵ factorial, main effects only | 32 | 6 | 26 | ×1.0385 |
Nothing in the column on the right depends on the number of factors, on which interactions are in the model, or on which run was lost. It depends on the spare capacity and on nothing else, and it falls away fast: at five spare runs a lost run costs 20%, at eleven it costs 9%, at twenty-six it costs 4%.
And at zero spare runs it costs everything. A saturated design — , which is the eight-run resolution-IV fraction a word of the defining relation buys — has no residual degrees of freedom, and removing a run leaves seven runs and eight coefficients. The information matrix is singular. The formula reads , and the design does not degrade: it stops being a design.
That is the practical form of the whole result. A design with no spare runs is one accident away from being unanalysable, and the number of spare runs is chosen when the design is, before anything can go wrong.
The correlation is the part that bites
The variance inflation is the number everybody looks at and the correlation is the one with the consequence, because it is not merely a loss of precision — it is a loss of the property the design was built to have.
A complete factorial is orthogonal: every off-diagonal of the inverse information matrix is exactly zero, so every coefficient can be estimated, interpreted and tested without reference to any other. That is what makes a factorial analysable by hand and reportable as a list.
Lose one run and every pair acquires a correlation of . At the sixteen-run design that is 0.1667, which is small. What it is not is zero, and three things follow that were previously not true:
The estimates are no longer independent. A coefficient reported as 3.2 ± 0.4 is now 3.2 ± 0.44, correlated 0.17 with every other coefficient in the table.
The order the terms are dropped in matters. Sequential sums of squares depend on the order when the columns are not orthogonal, so the analysis of variance a factorial produces — which is order-independent by construction — becomes order-dependent, and two analysts dropping terms in different orders will report different tables.
And nothing says so. The design is still described as a factorial in the methods section, the software fits it without complaint, and the fifteen rows look like fifteen rows of a sixteen-run design.
That last is the recurring shape: the output is well-formed, every quantity in it is computed correctly, and a property the reader is relying on has quietly stopped holding. It is the same class as a fraction reporting an interaction as a main effect — there the columns coincide by construction and here they stop being independent by accident, and in both cases the analysis of variance is a valid table about something other than what its row labels say.
A composite design prices its runs differently
Everything above is about orthogonal designs, and the second-order designs this field spends most of its time on are not orthogonal.
The symmetry argument does not apply, because the design’s runs are not interchangeable: a corner, an axial run at ±1.414 and a centre run are three different objects and the design knows it.
The centre runs are the cheapest to lose, which is worth stating carefully because it is not what “cheapest” usually means. Losing a centre run costs 25% on the intercept and almost nothing on the squared terms; losing a corner costs 67% on the interaction. The five centre runs are the design’s slack, and a design with several of them is a design that can afford an accident — which is a use for centre runs that the uniform-precision argument does not mention and which is often the more valuable one.
There is also a run this design cannot afford to lose at all, and it is not on the figure because it does not exist here: at one centre run, losing it leaves four distinct levels in each factor and the design is still estimable, but at zero centre runs the composite design cannot separate the two squared terms from the intercept. The slack is load-bearing rather than decorative.
Spare capacity is not a bound once the design is not orthogonal
The closed form is a statement about orthogonal designs, and it is worth measuring how far it travels, because “residual degrees of freedom” is the quantity everybody carries and the formula makes it look like the whole story.
The Box–Behnken has five spare runs and is not orthogonal: the largest correlation between its coefficients is −0.5547 before anything is lost. Lose one of its twelve edge runs and the worst variance is multiplied by 2.0000 — against the ×1.2000 the spare-capacity formula predicts for an orthogonal design with the same slack.
The central composite in three factors has seven spare runs, which is more, and pays more than the formula too: ×1.3786 for a corner, ×1.4971 for a centre run and ×1.5433 for an axial run, against a predicted ×1.1429.
So the closed form is a lower bound in practice rather than a rule, and it holds exactly for the one class of design where every column is independent of every other. The general statement is the one the algebra actually supports: the cost of a lost run is decided by how much of the design’s information that run was uniquely carrying, and orthogonality is the condition under which no run carries anything uniquely.
That also says which runs of a non-orthogonal design to protect. In the three-factor composite it is the axial runs — six of seventeen — because they are the only runs that let the three squared terms be told apart from each other, and losing one takes the design nearest to the boundary where it cannot do that at all.
What to do about it, which is a design decision
The result is about a loss that has already happened, and its use is before the data exists.
Do not read the spare-capacity formula as a guarantee. It is exact for an orthogonal design and optimistic for anything else, and the second-order designs this field spends most of its time on are all in the second class. A composite design with seven spare runs pays what an orthogonal design with two would, and the Box–Behnken with five pays what an orthogonal design with one would.
Build in spare capacity deliberately. The cost of a lost run is , so a design with one spare run pays double and a design with five pays a fifth. If a run is likely to be lost — a long-running process, a fragile assay, human subjects — the run to add is not a replicate of the most interesting setting but any run, because the spare capacity is what is being bought.
Prefer slack that can be lost. In a composite design the centre runs are both the cheapest to lose and the ones that buy uniform precision, so adding one is doing two jobs at once. In a factorial there is no such distinction and the slack has to come from replication or from a smaller model.
And check the orthogonality rather than the design name. A dataset described as a factorial may not be one. The check is one line — the largest off-diagonal correlation of the inverse information matrix — and it reads exactly zero for a complete design and for one missing a run.
Reading the inflation back as a sample size
There is a second way to state the cost that an experimenter can act on, and it converts a multiplication into a count.
A variance multiplied by is the variance a design of orthogonal runs would have had. So the sixteen-run factorial with eleven coefficients, having lost a run, is worth runs of the complete design — it has lost the equivalent of 2.67 runs, not one.
That is the number worth carrying. Losing one run costs more than one run, and how much more is the same spare-capacity ratio: the excess is of the whole design, so a design with five spare runs loses 2.67 and one with twenty-six loses 1.19.
The Box–Behnken is where it bites. Fifteen runs, five spare, and losing one edge run doubles the worst variance — so the fourteen remaining runs are worth 7.5 runs of the complete design. A single accident has cost half the experiment, at the design’s worst-affected coefficient, and the design still fits, still reports, and still describes itself as a Box–Behnken.
Why the result is about the design and not about the loss
One property of the closed form is easy to walk past and is the reason it is worth having: the cost does not depend on which run was lost, and for an orthogonal design it does not depend on the model’s contents either — only on how many coefficients there are.
That is unusual. Most statements about missing data are conditional on something about the missing observation: how extreme it was, where in the design it sat, whether its absence was related to the response. Here, for the class of design where the columns are independent, none of that enters the variance at all. Every run carries exactly of the information about every coefficient, so losing any one removes exactly the same amount of every kind.
The consequence is that the cost can be quoted before the study runs, as a property of the plan. If a run is lost, every standard error rises by ten per cent and every pair of coefficients becomes correlated at 0.09 is a sentence an experimenter can write down at the design stage, alongside the power calculation, and act on by adding a run.
It is also why the non-orthogonal case is a different kind of statement. There the runs are not interchangeable, the cost depends on which one goes, and the only honest form of the sentence names the run: losing an axial run raises the worst standard error by 24%, losing a centre run by 22%, and losing a corner by 17%. Three numbers rather than one, and all three computable from the design alone.
What is claimed here, and what is not
The claim is what one missing run costs: that for any orthogonal design the variance inflation is exactly and the induced correlation exactly , both verified against numerical inversion at six designs to ; that every run of a complete factorial costs the same by symmetry; that a saturated design is rendered unanalysable rather than degraded; and that a central composite prices its three kinds of run at 1.667, 1.667 and 1.250.
Nothing here is simulated. Every number is an inverse of an information matrix or a rank-one update of one.
What stays out: two or more lost runs, where the update is rank two and the cost depends on which pair in a way the rank-one case does not; the effect on the estimate rather than on its variance, which is zero in expectation for a lost run that is missing at random and is not zero when the reason for the loss is related to the response — a mechanism the missing-data field is entirely about, and which is the more serious problem in practice; and recovery designs, where the lost run is re-run and the question becomes whether the two batches are comparable.
The rank-one derivation assumes the lost run’s model vector has , which is true for a ±1 design under any model whose terms are products of ±1 columns and is false for a design with a centre run in it. That is why the composite design is counted rather than derived, and why its three numbers are reported as three numbers rather than as a formula.
Still open: the loss that is not at random
Everything here treats the missing run as missing for reasons unconnected to what it would have produced — an instrument failure, a scheduling clash, a broken sample. Under that assumption the remaining estimates are unbiased and only their precision suffers, which is what the whole essay measures.
The other case is the interesting one and nothing above prices it. A run lost because the process failed at that setting, or a subject withdrawn because the treatment was intolerable, is a run whose absence is information about the response — and then the remaining fifteen rows are not a fifteen-run design with a known inflation, they are a biased sample from a sixteen-run one.
The check, and the refusal
Two claims are gated at machine precision. That every run of a complete factorial costs exactly the same, required as an equality across all sixteen rather than as an agreement within a tolerance. And that the counted inflation equals at six designs with spare capacities from one to twenty-six, with the induced correlations checked against by the same standard — because a check on the variances alone would have passed on an arithmetic that got every covariance wrong.
The refusal is the composite design, and it is required to fail the symmetry claim: its worst inflation must vary by more than a tenth across its runs. If every design priced its runs equally, the symmetry argument would be doing no work and the factorial result would be a fact about arithmetic rather than about orthogonality. A design that disagrees is what makes the agreement mean something, which is the standing form of every refusal in this collection.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A design is a number — both name central composite design, experimental design, factorial design, least squares
- A cut is not a polynomial, and it does not have to be — both name correlation, experimental design, orthogonality
- A dictionary that is neither — both name correlation, experimental design, orthogonality
- A search that is already the other — both name degrees of freedom, least squares, orthogonality
- The arcsine that closes it, and the error that was overstated — both name correlation, experimental design, orthogonality
- The degrees of freedom in the sums — both name correlation, degrees of freedom, orthogonality
Named objects
A flat tag is an object no other essay names yet.
Box behnken designCentral composite designCorrelationDegrees of freedomExperimental designFactorial designFractional factorialLeast squaresOrthogonalityStudy designVariance inflation