A run lost for a reason
Worth reading first: A design is a number · Three mechanisms and one dataset.
The run that did not happen priced a lost run exactly. Lose any one of a sixteen-run factorial’s runs, with eleven coefficients in the model, and every coefficient’s variance is multiplied by 1.2 and every pair becomes correlated at 1⁄6 — the same for every run, by the design’s symmetry. It assumed the run was lost for a reason unconnected to what it would have produced, and under that assumption the remaining estimates are unbiased and only their precision suffers. It ended on the case it could not price: a run lost because the process failed at that setting, whose absence is information about the response.
That case is priced here. The design is the same sixteen runs at ±1 in four factors, fitting an intercept, four main effects and six two-factor interactions. The true surface has a main effect of A equal to 1, of B equal to 0.5, an interaction AB of 0.5 and nothing else, with noise of standard deviation one. And a run fails when its response passes a threshold — a reactor that trips, an assay that saturates, a specimen that breaks before it can be measured.
Three reasons to lose the same runs
To separate the reason from the count, three mechanisms lose the same expected number of runs. With the threshold at 2.5, the response passes it on 1.26 runs per dataset on average, nearly all of them at the corner where A and B are both high, whose expected response is 2. The second mechanism loses runs at random from that corner only, with the probability that removes the same 1.26 on average; the third loses 1.26 runs from anywhere at random.
The first two mechanisms bias nothing. At random anywhere, or at random at the high corner, every coefficient fitted to the surviving runs is unbiased to within the count’s error — the second is the more instructive, since its losses are concentrated at exactly the corner the threshold empties. Losing runs because of where they were in the design is losing them because of something the design records, and least squares on the remaining rows is still an unbiased fit of the same surface. That is the earlier essay’s world, with a precision cost and nothing else.
The threshold is different. Fitted on the surviving runs, the intercept reads −0.130 instead of zero, A reads 0.869 — a bias of −0.131 on an effect of one — B reads 0.380, a bias of −0.120 on an effect of a half, and AB reads 0.383, a bias of −0.117. The interaction and the smaller main effect are each understated by about a quarter. The coefficients that had nothing to do with the response being high — C, D and every interaction not involving both A and B — are unbiased.
Where on the surface the bias lands
The four biased coefficients are biased together, and their sum is what an experimenter reading the fitted surface would see. At the corner where A and B are both high, the fitted height is the intercept plus A plus B plus AB, and the four biases add: the survivors’ fit puts that corner 0.498 below its true height of 2 — a quarter of it. At the opposite corner, where both are low, A and B enter with minus signs and AB with a plus, and the biases cancel to 0.005. At the two mixed corners they nearly cancel as well, to −0.023 and −0.001. The whole of the damage sits at the one corner where the runs were lost, which is also the corner the next experiment would be designed around, since it is where the process performs best. The corner’s own four runs, had they all survived, would have measured it with a standard error of a half; the fit’s error of half a unit there is as large as that standard error, systematic, and in the same direction on every dataset.
So the fitted surface is right where nothing was lost and wrong where the losses were, and wrong in the direction that hides the reason for them. An experiment run to find where the response is highest — the usual purpose of a factorial on a process — would find the high corner, report it a quarter lower than it is, and never learn that its own failures were the evidence. The runs lost to the threshold were the ones with the highest responses, and the fit that leaves them out describes a process that never reaches the threshold at all.
Why only the response biases
The bias comes from which values are lost, not which rows. At the high corner the expected response is 2, and a run there is lost exactly when its noise is above one half. The runs that survive at that corner are the ones whose noise happened to be low, and their average response is below the corner’s true mean. The fit on the survivors therefore sees a corner that is lower than it is, and every coefficient that contributes to that corner’s height — the intercept, A, B and AB, all positive there — is pulled down together.
That is the distinction the assumption that identifies the mechanism made for missing outcomes: a value missing for a reason recorded in the data is missing at random given what was recorded, and a value missing for a reason connected to itself is not. In a designed experiment the design matrix is everything recorded about a run besides its response, so a loss decided by the settings is ignorable and a loss decided by the response is not — whatever the number of runs involved. Dropping the incomplete rows is safe exactly when that holds, and here the threshold breaks it at the one corner the experiment was most interested in.
The interval does not show it. The fit on the survivors reports its own standard errors from the residuals, and the bias of AB is about four tenths of a standard error. The nominal 95% interval for AB still covers 94.0% of datasets, and for A 94.1%. An analyst reading the intervals would see slightly wider ones than the full design gives, for the one or two runs lost, and nothing else.
How the bias grows with the losses
Lowering the threshold loses more runs and biases more. At a threshold of 3.5 the response passes it on 0.26 runs per dataset and AB is biased by −0.031; at 3, 0.64 runs and −0.067; at 2.5, 1.26 runs and −0.117; at 2, 2.09 runs and −0.181; at 1.5, 3.08 runs and −0.248 — half the interaction’s true size. The bias per run lost falls a little as the threshold drops, because a lower threshold starts removing runs whose responses were not extreme, but the total keeps rising.
The precision cost the earlier essay measured, by contrast, does not depend on why the runs were lost, and it is the smaller of the two. Losing one run multiplied each variance by 1.2, raising AB’s standard error from 0.250 to 0.274. A threshold that removes a quarter of a run on average already adds a bias of 0.031 — more than that whole increase in the standard error.
When the corner is gone
A response-driven loss has a failure the random one does not. Its losses pile up at one corner, and the four runs where A and B are both high are the only ones that tell the AB interaction apart from the other three combinations. When all four are lost the interaction is not merely biased but unestimable: the surviving information matrix is singular. At a threshold of 2.5 that happens on 0.88% of datasets, with the eleven-coefficient model; at 2 on 6.30%, and at 1.5 on 24.6%, by then also because too few runs survive to leave a residual degree of freedom. The bias figures above are averages over the datasets that could still be fitted, so they understate the problem by exactly the datasets where it was worst.
A random loss of 1.26 runs never empties a corner in this design’s twenty thousand datasets. The response-driven loss does it about one time in a hundred and fourteen, because the response is the thing it is following.
Repairing it by treating the lost runs as censored
The response of a lost run is not unknown: it is known to be above the threshold. That is a censored observation, and the textbook repair is to fit the surface by maximum likelihood with the lost runs entering as censored — contributing the probability of being above the threshold rather than a density at a value. The fit here is by expectation–maximisation: replace each lost run by its expected response above the threshold under the current fit, add its conditional variance to the residuals, refit, and repeat until the coefficients stop moving.
With eleven coefficients, it overshoots. At a threshold of 2.5 the censored fit’s bias on AB is +0.066, against the survivors’ −0.117: it removes the bias and puts back half of it with the opposite sign. Knowing the noise’s standard deviation exactly, rather than estimating it from five residual degrees of freedom, helps a little and does not fix it: A’s bias is then +0.046. At a threshold of 2, the overshoot is +0.105 on AB.
With six coefficients — the intercept, the four main effects and AB, dropping the five interactions the true surface does not have — the same repair is nearly unbiased: +0.020 on AB at a threshold of 2.5, +0.017 at 2 and −0.015 at 1.5, where three runs are lost on average. The survivors’ bias is the same for either model, about −0.12 on AB at 2.5, so the difference is entirely in the repair.
The repair’s error is also not a matter of the noise being estimated from too few residuals. With eleven coefficients, fixing the noise at its true value moves A’s bias from +0.064 to +0.046, and with six coefficients from +0.016 to +0.012: a quarter of the overshoot at most. Most of it is the model’s size.
Why the repair needs spare capacity
The reason is leverage, and it connects this essay’s repair to the earlier essay’s price. In an orthogonal design every run has leverage — eleven sixteenths with eleven coefficients — which is the share of a run’s fitted value determined by the run itself. With leverage that high, a censored run’s fitted value is mostly set by its own likelihood contribution, and the likelihood of a single observation known only to lie above a threshold keeps rising as its mean is pushed upward: one censored observation, on its own, would set its mean at infinity. The other runs restrain it by the remaining five sixteenths. That restraint is not enough, and the coefficients that make the lost run’s fitted value high are pushed too far.
With six coefficients each run’s leverage is six sixteenths, and ten sixteenths of each fitted value come from the other runs. The censored run’s upward pull is outweighed, and the likelihood settles where the survivors and the censoring together put it. The repair works when the design has spare capacity — the same that set the earlier essay’s price of for a run lost at random. A design with five spare runs pays 1.2 for a random loss and repairs a censored one only by overshooting; a design with ten pays 1.1 and repairs it. Spare capacity bought against accidents is also what makes an informative loss repairable, which is a second reason to buy it.
What the smaller model assumes
Dropping five interactions to make the repair work is a trade, and the trade has a price the measurements here do not show, because the true surface really has no interactions besides AB. A six-coefficient fit assumes the five dropped interactions are zero; if one of them were not, the smaller model would be biased whether or not any run was lost, and the censored repair would be fitting the wrong surface well.
So the choice is between two assumptions about things the experiment cannot fully check: that the dropped terms are absent, or that the lost runs are missing for reasons unconnected to their responses. With eleven coefficients the analyst assumes nothing about the surface and must either accept the survivors’ bias or a repair that overshoots; with six, the analyst assumes a simpler surface and gets a repair that works. The design that avoids the trade is the one with spare capacity built in from the start — more runs, not fewer terms — and that is a decision made before any run fails. The run that did not happen gave the arithmetic for its cost; this essay gives a second reason to pay it.
What a report of an experiment with failed runs should state
Why each run was lost. A run lost to a scheduling clash and a run lost because the process tripped are different events, and only the second biases the fit. The reason is usually known to whoever ran the experiment and rarely appears in the dataset.
The settings of the lost runs. Losses concentrated at one corner are the signature of a response-driven failure, and four lost runs at the high corner of two factors mean their interaction was not estimated at all, whatever the software reports.
Whether the lost runs were treated as censored, and how many coefficients the model had. With five spare degrees of freedom a censored fit overcorrects by about half the bias it removes; with ten it is close to unbiased. The curve that survives censoring is the same likelihood in its natural setting, where every subject contributes and the model is small.
Counted, on what
Twenty thousand datasets at each threshold and mechanism, generated from the stated surface with unit normal noise, one stated seed per mechanism. The surviving runs are fitted by least squares; datasets whose survivors leave the model singular or without a residual degree of freedom are counted as unfittable and left out of the bias averages. The censored fit is the expectation–maximisation of the normal censored likelihood, run to a change below 10⁻¹⁰ in every coefficient, with the noise either estimated or fixed at its true value. With the bias of a coefficient’s estimate measured over twenty thousand datasets, its standard error is about 0.002, so every bias quoted above to three decimals is several times its own error except those reported as zero. One surface is measured; a surface whose high corner is less extreme, or a threshold below the response rather than above it, changes the sizes and not the mechanism.
Still open: a threshold the analyst does not know
Everything here assumes the threshold is known — the reactor’s trip point, the assay’s ceiling. Often it is not: a process “failed” without anyone recording the level at which it fails, and the only information is that the run did not produce a usable value. Then the censoring point is a parameter too, estimated from the pattern of which settings failed, and the likelihood has to carry both the surface and the threshold.
Whether the surface can still be recovered when the threshold is unknown, how much the extra parameter costs in a design with ten spare runs, and whether a design can be augmented — as one that has already run was — with runs placed to locate the threshold rather than to estimate the surface, is the measurement this leaves. Three levels and a ring showed that where a design puts its runs decides what it can see; a design that expects failures at one corner might do better to put its spare runs just below it.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The part the rule already took — both name experimental design, least squares, leverage, orthogonality
- A point ordinary on every axis — both name experimental design, least squares, leverage
- A probe chosen from the design — both name experimental design, leverage, orthogonality
- A quantity that loses to a heuristic — both name experimental design, leverage, orthogonality
- A search that is already the other — both name degrees of freedom, least squares, orthogonality
- A term built from the others — both name experimental design, least squares, leverage
Named objects
A flat tag is an object no other essay names yet.
BiasCensoringDegrees of freedomEm algorithmExperimental designFactorial designLeast squaresLeverageMaximum likelihoodMissing at randomMissing not at randomOrthogonality