One factor at a time
Worth reading first: The variance removed before the data · Randomisation is not balance.
An experiment has a budget of runs and several factors to investigate: temperature and pressure, dosage and timing, three settings on a machine. The instinct — and the advice in most laboratory training — is to vary one at a time, holding everything else fixed, so that any change observed can be attributed to the thing that was changed.
It is a worse use of the same runs, by a factor that has a closed form, and in the presence of an interaction it produces a recommendation that is wrong rather than merely imprecise.
The two ways to spend eight runs
One factor at a time fixes a baseline with every factor low, then raises each factor in turn. With three factors that is four conditions — the baseline and three single changes — and each effect is estimated by comparing one condition with the baseline.
A factorial design runs every combination. With three factors at two levels each that is eight conditions, and a factor’s main effect is the difference between the average of the four runs where it was high and the average of the four where it was low.
The second uses every run to estimate every effect. That is the whole of the efficiency argument and it is called hidden replication: the four runs contributing to the first factor’s effect are the same four contributing to the second factor’s and the third’s, so nothing is spent twice and nothing is spent once.
The ratio, exactly
Both designs are unbiased for the main effects, so the comparison is of variance, and both have closed forms. With N runs, σ² the noise variance and k factors:
factorial: 4σ²/N — the difference of two means of N/2 runs each
one at a time: 2σ²(k+1)/N — each effect from two conditions of N/(k+1) runs
The ratio is (k + 1)/2, and it does not depend on the noise, the effects, or the number of runs. At three factors the one-at-a-time design’s estimates are twice as variable; at six factors, three and a half times.
Measured over two thousand experiments at equal runs, the ratio comes out at 2.00 at three factors and 3.46 at six, against predictions of 2.0 and 3.5.
The factor of two, and where it comes from
The efficiency has a closed form and it is worth deriving before the interaction argument, because the two are separate complaints and only one of them is about precision.
In a factorial, a main effect is the average of the half of the runs where the factor is high minus the average of the half where it is low. Every one of the N runs is in one of the two averages, so the effect’s variance is .
In one-factor-at-a-time with k factors, the N runs are divided among k + 1 conditions, and an effect compares two of those conditions. Each average is built from runs, so the variance is .
The ratio is
At three factors the factorial is twice as precise; at seven factors, four times. Equivalently, the one-factor-at-a-time design needs times as many runs to match it, and the penalty grows without limit as more factors are added.
The reason is one sentence. In a factorial every run contributes to every effect; in one-factor-at-a-time each run contributes to one effect and the baseline contributes to all of them, which is why the baseline is the most replicated condition in a design that never uses it for anything but subtraction.
The comparison has to be at equal runs, and it is easy to get wrong
The measurement above hides a trap that this site walked into while building it, and it is worth recording because the wrong version looks like a result.
A factorial design over k factors needs a multiple of 2ᵏ runs; a one-at-a-time design needs a multiple of k + 1. Those rarely coincide. Give both designs “about twenty runs” and the factorial gets 16 while the one-at-a-time gets 18, or the factorial gets 8 and the other 6 — and the measured ratio then includes the run-count difference along with the efficiency difference.
At two factors that produced a measured ratio of 1.98 against a predicted 1.5: a 32% discrepancy that reads as the closed form being wrong, and is entirely the comparison being unfair. Fixing it required choosing run counts divisible by both 2ᵏ and k + 1 — 12 runs at two factors, 8 at three, 80 at four — after which the agreement is exact.
The general lesson is one this site keeps meeting: a comparison between two procedures has to hold constant the thing being spent. Otherwise the answer measures the budget rather than the method.
Then the argument that makes the efficiency beside the point
An efficiency ratio of two is a strong argument and it is not the strong argument. The strong one is that a one-at-a-time design cannot see an interaction at all, and in the presence of one its conclusion can be wrong in a way no amount of replication fixes.
Take a response with two factors:
y = 3x₁ + 2x₂ − 5x₁x₂
on factors coded −1 and +1. From the baseline at (low, low), raising the first factor alone gains 16, and raising the second alone gains 14. Both single changes are large improvements, both would be found convincingly by any experiment with enough replication, and both are correctly measured.
The one-at-a-time analysis adds them: if raising the first helps and raising the second helps, raise both. That combination gains 10 — worse than either single change — and it is a corner the design never ran.
Across three thousand simulated experiments the one-at-a-time procedure recommends the best combination 0% of the time. The factorial design, on the same number of runs and the same noise, finds it 99.2% of the time.
Why the failure is total rather than probabilistic
The 0% is not a rounding of a small number and it is worth being precise about why.
One factor at a time measures each factor’s effect at the baseline setting of the others. That is a real quantity, correctly estimated, and it is not the main effect — it is the simple effect at one corner. Adding two simple effects measured at the same corner predicts the diagonal corner only if the factors do not interact.
When they do, the prediction is wrong by exactly the interaction term, every time, in the same direction. Noise moves the estimates around; it does not move the conclusion, because the conclusion follows from the two single changes being positive, which they reliably are.
So the design fails deterministically rather than unluckily, and — the part that matters — it contains no information that it has failed. Every measurement it made is correct. There is no diagnostic, no residual, no warning. The experiment that would reveal the problem is the one the design does not run.
Interactions are not exotic
The natural response is that interactions are a special case, and the two-factor square above was built to have one.
The response has an answer in three parts.
They are common in every applied field that has looked. A drug that helps at one dose and harms at another in combination with a second drug; a catalyst effective only above a temperature; a teaching method that works for prepared students and not for others. Wherever the mechanism has a bottleneck, raising one input helps only when the other is not limiting, which is an interaction.
A factorial design measures them for free. The eight runs that estimate three main effects also estimate all three two-factor interactions and the three-factor one. Not as an extra: the same runs, different contrasts, no additional cost.
A one-at-a-time design cannot measure them at any price, because the runs that would reveal an interaction are the combinations it does not visit. The design does not produce a weak estimate of an interaction. It produces none.
That asymmetry is why the efficiency argument is the smaller one. The design that is twice as precise is also the one that answers a question the other cannot ask.
What eight runs actually buy
It is worth listing what a three-factor design of eight runs estimates, because the list is longer than most people expect and it is the concrete form of the hidden-replication argument.
Three main effects, each from all eight runs — four high against four low.
Three two-factor interactions, each also from all eight runs, as the difference between the effect of one factor when a second is high and when it is low.
One three-factor interaction, again from all eight.
Seven estimated quantities from eight runs, each using every observation, all mutually orthogonal so that no estimate disturbs another. The one-at-a-time design with the same eight runs — a baseline and three single changes, doubled up — estimates three quantities, each from four runs, and nothing else.
Orthogonality is worth a sentence of its own because it is what makes the analysis simple rather than merely efficient. Each effect is estimated independently of the others, so dropping a term from the model does not move the remaining estimates. That is exactly the property a regression with correlated predictors lacks, and it is a property the design creates rather than one the data happens to have.
Why the advice persists
The one-at-a-time rule is not a mistake anyone makes carelessly. It is taught, it is intuitive, and its appeal survives being told the arithmetic, so it is worth naming what is right about it.
It is the correct rule for debugging, and the two get confused. When a system has broken and the question is which change caused it, changing one thing at a time is exactly right, because the target is a single cause and the interactions are not of interest. Experimentation to understand a response surface is a different task with a different answer.
Attribution feels safer. With one factor changed, whatever moved is attributable to that factor — a real property, and it holds for the factorial design too, where each effect is a contrast that is orthogonal to the others by construction.
The design fails invisibly. As established above, every measurement a one-at-a-time experiment makes is correct. Nothing in its output signals the missing corner, and so nobody who has used it has seen it fail. This is the shape of failure this site keeps finding — an estimator that is confidently wrong with no distress signal — and it is far more durable than a method that produces obvious nonsense.
And sequential experimentation genuinely works, in a form that resembles it. Changing a factor, moving in the direction that improved things, and repeating is the method of steepest ascent, and it is a serious technique. The difference is that it uses a small factorial design at each step to decide the direction, rather than a single change to decide it.
The connection to the rest of this site
Two other fields turn out to be about the same arithmetic, and stating that makes the design decisions transferable rather than a separate body of lore.
A factorial experiment estimates several effects at once, so it produces several p-values, and everything the corrections field says applies: seven effects tested at 0.05 will produce a spurious one about a third of the time. The design does not create that problem — a one-at-a-time programme testing the same factors has it too, spread across more experiments where it is harder to notice — but it makes it visible in one table, which is where it can be handled.
And the runs are allocated to conditions, which makes every design in this essay a candidate for the blocking and randomisation of the two before it. A factorial design run in blocks, with the allocation randomised within them, gets all three gains at once, and they multiply rather than compete.
Fractional designs, when even the factorial is too many runs
The objection to factorial designs is arithmetic: 2ᵏ runs is 32 at five factors and 1,024 at ten.
The standard resolution is to run a fraction of the design — half, a quarter, a sixteenth — chosen so that the main effects remain estimable while the highest-order interactions are given up. A half-fraction at five factors is 16 runs and still estimates all five main effects with every run contributing to each.
The price is aliasing: some effects become indistinguishable from others. In a half-fraction, each main effect is confounded with a four-factor interaction, which almost nobody minds; in a sixteenth-fraction, main effects can be confounded with two-factor interactions, which matters a great deal given the essay above.
The quantity that describes this is the design’s resolution, and it is worth knowing as a piece of vocabulary because it is what a design catalogue is indexed by. Resolution III confounds main effects with two-factor interactions; resolution IV keeps main effects clear but confounds two-factor interactions with each other; resolution V keeps both clear.
The point for a reader is that fractional designs are how the run count is controlled, and that the choice is which interactions to sacrifice — a decision taken in advance, on subject-matter grounds, and recorded. A one-at-a-time design makes the same sacrifice implicitly and completely, sacrificing every interaction without saying so.
What the two arguments come to
Two claims, both measured, and they are of different kinds.
Efficiency. At equal runs, one factor at a time produces estimates (k+1)/2 times as variable — 2.00 at three factors, 3.46 at six, against closed forms of 2.0 and 3.5. That is a cost, and it is recoverable by running more experiments.
Validity. With interacting factors, the one-at-a-time design recommends a combination that is worse than either single change, 0% right against the factorial’s 99.2%, and holds no evidence that anything went wrong. That is not recoverable by running more of the same experiment, because the experiment does not visit the conditions that would show it.
The practical instruction is short. Vary everything at once, on a designed grid, and use a fraction of it when the full one is too large. The runs are not more numerous, the analysis is not harder, and the design answers the question about combinations that the alternative silently answers wrong.
Two levels, and when two is not enough
Every design here uses two levels per factor, which is a deliberate restriction with a clear boundary.
Two levels estimate a straight line through the factor’s range, which is all that is needed while the question is which factors matter and in which direction. That is the screening stage, and it is where most experimental budgets should go first, because a factor that does nothing can be dropped before anything expensive is spent on it.
Two levels cannot detect curvature. A factor whose response peaks in the middle of the range tested shows the same average at both ends, so its main effect is zero and the design reports no effect at all — the one failure mode a two-level factorial shares with the design it beats.
The standard repair is cheap and worth knowing: add a few runs at the centre of every factor’s range. If the centre runs sit near the average of the corners, the response is close to planar; if they sit well above or below, there is curvature, and the design says so with three or four extra runs. Establishing that a response surface is curved and then fitting it needs more levels, and that is the response-surface stage rather than the screening one.
The order — screen at two levels, check the centre, then model whatever survives — is a decision about where to spend runs, made before the data exists, which is what this whole field is.
What is left
The last essay of this field is about the number every design decision eventually converts into: how many units are needed to detect an effect of a given size.
It is also where this site closes a gap it has recorded since its first essays. Every power number here has been simulated, with nothing independent to disagree with. The non-central t supplies the closed form, and the two routes now have to agree — which is the standard every other quantity on this site has been held to.
What links here
Computed from the collection, not written here: the essays that point at this one.
Reads more easily once this is understood
Essays that name this one as worth reading first.
- A design is a number
- The design that cannot see a curve
- The optimum is a ratio, and its interval is sometimes the whole line
- Three levels, and the ring where the design says the same thing
- Walking up the gradient
- The word a fraction costs
- The design that refuses the corners
- The run that did not happen
Named objects
A flat tag is an object no other essay names yet.
Factorial designHidden replicationInteractionMain effectReplication