The word a fraction costs
Worth reading first: One factor at a time.
Changing everything at once estimates each effect from every run, which is why a factorial beats the one-at-a-time design by a factor of . The cost is that the design has runs, and doubles.
At five factors that is thirty-two runs to estimate six things anybody usually cares about — an intercept and five main effects. The obvious move is to run half of them, and the question is what half a design cannot do.
Why it is an identity
The word confounded suggests something statistical, and nothing statistical is involved.
Building the fraction by setting means that in every run of the design, the column of values and the column of values are the same column. Not similar, not correlated — equal, entry by entry, in all eight rows.
Least squares fits coefficients to columns. Two identical columns cannot be given different coefficients, so the arithmetic does the only thing available: it reports one number for the pair. The expected value of that number is the sum of the two true effects.
exactly, for any sample size, any noise level and any number of replicates. Running the design a thousand times narrows the interval around the sum and does not separate the pair, because there is nothing to separate — the design contains no information that distinguishes them.
Which columns coincide is decided by the defining relation. Setting makes a column of all ones, which is the intercept’s column, and the shorthand for that is . Multiplying any effect by the word gives its partner: , so estimates .
Resolution is the length of the shortest word
The defining relation here has one word and it has four letters, so the design is called resolution IV. The number is the length of the shortest word, and it is the only summary of a fraction worth carrying because it says which orders of effect get mixed.
A word of length confounds any effect of order with an effect of order . So:
- Resolution III: a word of three letters confounds a main effect () with a two-factor interaction (). Main effects are contaminated by the interactions that are most likely to be real.
- Resolution IV: main effects are confounded with three-factor interactions, which are usually small. Two-factor interactions are confounded with each other.
- Resolution V: main effects with four-factor interactions, two-factor interactions with three-factor ones. Everything an experimenter usually wants is clear.
What the contamination is worth
The alias chain says what is added; how much it matters is a separate number and it is computable exactly, because the bias is an expectation rather than a sample.
The third row is the one to look at. The true effect of C is −1 and the design reports −0.20 — the same sign, a fifth of the size, and an experimenter reading it would conclude that C barely matters. One three-factor interaction of 0.8, which is a modest thing to suppose exists, has removed four fifths of a real effect.
Move the slider to zero and every bar is exact. That frame is the control rather than a decoration: it says the contamination comes from the interactions being real, not from the design being a fraction. A fraction is unbiased when its assumption holds, and its assumption is about quantities it cannot measure.
At interactions of 2 the reported effect of C is +1.00 — the wrong sign. There is no amount of data that reveals this, and no residual pattern that flags it, because the model fitted has no term for the thing that is contaminating it.
Every run, every effect, still
One property survives fractionation intact and it is the one that made the factorial worth having in the first place, so it deserves saying before the costs are counted.
In the eight-run fraction, every one of the eight runs contributes to the estimate of every effect. The estimate of is the average of the four runs where is high minus the average of the four where it is low — all eight runs, each used once. That is the hidden replication the essay that priced one factor at a time found: the one-at-a-time design would need two runs per effect and would estimate each from two observations, and this estimates each from eight.
So the fraction is not a retreat towards one-at-a-time. It keeps the whole of the efficiency argument and gives up something the one-at-a-time design never had either — it too confounds, worse and without a defining relation to say how. The comparison that matters is against the full factorial, not against the arrangement the factorial replaced.
What the design cannot test
The usual defence of resolution IV is that three-factor interactions are rare and small, and it is a good defence. It is also an assumption, and the point worth making is about its status rather than its plausibility.
A resolution-IV fraction has eight runs and estimates eight quantities: the intercept, four main effects, and three pairs of two-factor interactions. There are no degrees of freedom left. Even if there were, the three-factor interactions have no column of their own — their columns are the main effects’ columns — so there is nothing to test them against.
So the design is in the position a design that cannot see a curve is in for curvature, one order up: the thing that would invalidate the analysis is not merely unestimated, it is unrepresented. The repair there was centre runs, which buy one number back. The repair here is more runs of the same kind, which is to say a bigger fraction.
Two fractions of the same size
The number of runs does not say what a design can estimate, and the cleanest demonstration is two designs with the same run count.
gives the word , length five, resolution V: main effects are clear of everything up to four-factor interactions, and two-factor interactions are clear of each other. gives the word , length three, resolution III: the main effect of is confounded with the interaction , which is a quantity an experiment studying and has every reason to expect is real.
Sixteen runs each. One design answers the question and one does not, and nothing about the run count distinguishes them. The relation is the design. A fraction described only by its size — a half fraction was run — has not been described.
Sixteen runs, and an effect that disappears
The resolution-III half fraction in the table above is worth following through, because it is the arrangement an experimenter would build by accident.
Five factors, sixteen runs, and the generator chosen as — which is what happens when the fifth factor is added late and assigned to a column that looked free. The defining word is , so
Three of the five main effects are confounded with two-factor interactions among themselves. If and interact — which is the thing a factorial is run to find out — then the reported effect of the newly added factor is that interaction, and the newly added factor will look important whether or not it does anything.
The failure has a particular shape that makes it hard to catch. The design is orthogonal, every diagnostic is clean, every standard error is what it should be, and the analysis of variance is perfectly well formed. The residuals carry no pattern, because the model has absorbed the interaction into a main effect rather than leaving it out. It is the same class as an estimator returning the edge of its own range: the output is a valid object and it is an object about something else.
Compare , the same sixteen runs, resolution V. There estimates , and a four-factor interaction is a quantity almost nobody has ever needed. The difference between the two designs is a choice made in an afternoon, and it decides whether the experiment can answer its question.
Where the fractions sit against one another
Three arrangements have now been priced in this field and it is worth putting them on one line, because they answer three different questions and get compared as though they answered one.
One factor at a time estimates each effect from two runs, misses interactions entirely and recommends settings it never tried. It is the arrangement with no design in it.
The complete factorial estimates every effect from every run and confounds nothing. It is what a fraction is a fraction of, and it costs .
A fraction estimates every effect from every run and confounds each with a set determined by its generators. Its efficiency per run is the factorial’s exactly, and what it additionally gives up is identifiability.
That second distinction is the one that gets blurred, so it is worth the arithmetic. In any design whose columns are orthogonal and coded ±1, the variance of every coefficient is — the information matrix is times the identity and there is nothing else in it. So halving the runs doubles every variance, which is the ordinary and expected cost of a smaller experiment, and it is the cost an experimenter has already decided to pay.
The alias structure is a second cost on top of that, and it is of a different kind. A fraction is not a noisier factorial: its coefficients have exactly the standard errors that any orthogonal design of its size would give, and they are estimates of different quantities. More runs of the same fraction shrink the interval around the sum and never split it.
The practical form of that is a sentence worth carrying. A fraction’s precision improves with replication and its confounding does not, so an experimenter who suspects the alias structure is biting cannot answer the suspicion by running the same design again.
What a resolution buys, as a count
The three resolutions are usually described in words, and the words hide that they are counts of the same thing.
A word of length in the defining relation means that every effect of order is confounded with one of order . So the question is my main effect safe is the question is large enough that effects of that order are negligible, and the orders are ordered by how often anybody has seen one:
| resolution | main effects are mixed with | two-factor interactions are mixed with |
|---|---|---|
| III | two-factor interactions | main effects |
| IV | three-factor interactions | each other |
| V | four-factor interactions | three-factor interactions |
Reading down the first column is the standard advice and reading down the second is where the useful distinction is. Resolution IV protects the main effects and does nothing for the interactions: at four factors the pairs AB–CD, AC–BD and AD–BC are inseparable, so a design that finds a large interaction cannot say which of two it is. That is often exactly the finding an experiment is run to produce.
Resolution V is where both columns are safe, and at five factors it costs sixteen runs against thirty-two — a half rather than a quarter. The step from IV to V is the expensive one and it is the one that buys the answer to the second question.
What an experimenter should decide, and in which order
The arithmetic supports an order of decisions, and it is not the order the run count suggests.
First, which interactions can be assumed away. That is a subject-matter question, it is the assumption the whole fraction rests on, and it is not answerable from the data the design will produce. An experimenter who cannot name the interactions they are willing to confound has not yet chosen a design.
Then, the resolution that assumption implies. If two-factor interactions among the factors are plausible — which is the usual case — the answer is IV at minimum and V if the interactions are themselves of interest.
Then, the runs that resolution costs, which is the only step where the budget enters. At five factors, resolution V is sixteen runs and resolution IV is also sixteen; at six factors, V is thirty-two and IV is sixteen, and that is where the budget starts making the decision.
Reversing the order — starting from the affordable run count and taking whatever resolution comes with it — is what produces the design in the section above, and it is what the table of fractions is for. The two sixteen-run designs there cost the same and answer different questions, so a budget alone cannot pick between them.
This is the same shape as the decision the optimal-design field makes explicitly: a design is chosen against a criterion, the criterion encodes what the experiment is for, and starting from the size is starting from the one input that carries no information about the question.
What is claimed here, and what is not
The claim is what a fractional factorial confounds and what that costs: that the alias chain is an identity about the columns of the design rather than an approximation about the data; that the bias on each reported main effect is the exact sum of its chain, so a true effect of −1 is reported as −0.20 when every interaction is 0.8 and as +1.00 when every interaction is 2; that resolution is the length of the shortest word in the defining relation and decides which orders get mixed; and that two fractions of the same size can differ by two levels of it.
Nothing in this essay is simulated. Every number is a column identity or an expectation computed from the design.
One qualification on the bias numbers. They are computed with every interaction set to the same value, which is a stated construction rather than a realistic world: it makes the bias a function of how many contaminants a chain has and of nothing about which they are, which is what a figure about the design should show. Real interactions differ in size and sign, and a chain with two contaminants of opposite sign can come back unbiased by cancellation. The construction is an even-handed case rather than a worst one, and the worst case is what a design guarantees against the least favourable setting rather than what it does on average.
What stays out: how to choose generators to maximise resolution, which is a combinatorial problem with a well-developed literature and published tables; foldover designs, which add a second fraction to break specific alias chains and are the standard repair when a resolution-III screen finds something; Plackett–Burman designs, whose aliasing is partial rather than complete and whose chains are therefore fractions rather than identities; and the whole question of which interactions to expect, which is subject-matter knowledge and is what the choice of generators is supposed to encode.
Still open: a design that refuses a setting
Every design so far is built on a corner — the combination where every factor is at its extreme at once. The factorial is corners, the fraction is half the corners, and the central composite is corners plus axial runs.
A corner is sometimes not a setting. A temperature at the equipment’s limit combined with a pressure at the vessel’s limit; two ingredients each at their maximum in a formulation that must sum to one; a dose at the maximum tolerated combined with a duration at the maximum permitted. What the arrangement looks like when the corner is removed, and what predicting at a place the design never visits costs, is what a design that refuses the corners costs — where a fifteen-run design predicts the corner 1.84 times worse than a seventeen-run one that goes there.
The check, and the refusal
Two claims are gated. That every effect’s alias chain has exactly as many members as the defining relation has words, which is the group structure stated as a count and would fail on any error in the word list. And that the defining relation contains no word of length two, which would confound two main effects and produce a design nobody would call a design.
The refusal is the one that keeps the contamination honest: at a contamination of zero, every reported main effect must equal its own value to machine precision. A check that found bias at every setting of the slider would be reporting a defect in the arithmetic rather than a property of the fraction, and the zero frame is the only one that can tell the two apart. It is also the frame a check written to demonstrate the problem would have been most likely to leave out — which is the standing shape of a check that has never rejected anything.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- Balancing what is known in advance — both name confounding, interaction, study design
- Three arms and three scores — both name experimental design, interaction, study design
- A copula that halves a marginal — both name experimental design, interaction
- A cut is not a polynomial, and it does not have to be — both name experimental design, interaction
- A dictionary that is neither — both name experimental design, interaction
- A symmetry that was not enough — both name experimental design, interaction
Named objects
A flat tag is an object no other essay names yet.
Alias structureConfoundingDesign resolutionExperimental designFactorial designFractional factorialInteractionLeast squaresMain effectStudy design