Decided before the data

The word a fraction costs

A half fraction estimates each main effect as an exact sum of that effect and everything it is confounded with — no error term, no sample-size argument. With every interaction at 0.8 the design reports a true effect of −1 as −0.20, and the design cannot test the assumption that makes the number mean anything.

Worth reading first: One factor at a time.

Changing everything at once estimates each effect from every run, which is why a factorial beats the one-at-a-time design by a factor of (k+1)/2(k+1)/2. The cost is that the design has 2k2^{k} runs, and 2k2^{k} doubles.

At five factors that is thirty-two runs to estimate six things anybody usually cares about — an intercept and five main effects. The obvious move is to run half of them, and the question is what half a design cannot do.

What a 8-run fraction of 4 factors confounds. The defining relation is I = ABCD, so the resolution is 4. A is estimated as A + BCD; B is estimated as B + ACD; C is estimated as C + ABD; D is estimated as D + ABC. Each of those is an identity about the design rather than an approximation about the data.
Fig. 1 The half fraction of four factors built by setting D = ABC. Eight runs, and the estimate of each main effect is the sum of that effect and a three-factor interaction. Every line is an identity about the arrangement of the runs.

Why it is an identity

The word confounded suggests something statistical, and nothing statistical is involved.

Building the fraction by setting D=ABCD = ABC means that in every run of the design, the column of DD values and the column of ABCABC values are the same column. Not similar, not correlated — equal, entry by entry, in all eight rows.

Least squares fits coefficients to columns. Two identical columns cannot be given different coefficients, so the arithmetic does the only thing available: it reports one number for the pair. The expected value of that number is the sum of the two true effects.

E[D^]=D+ABC\mathbb{E}[\hat{D}] = D + ABC

exactly, for any sample size, any noise level and any number of replicates. Running the design a thousand times narrows the interval around the sum and does not separate the pair, because there is nothing to separate — the design contains no information that distinguishes them.

Which columns coincide is decided by the defining relation. Setting D=ABCD = ABC makes ABCDABCD a column of all ones, which is the intercept’s column, and the shorthand for that is I=ABCDI = ABCD. Multiplying any effect by the word gives its partner: A×ABCD=BCDA \times ABCD = BCD, so A^\hat A estimates A+BCDA + BCD.

Resolution is the length of the shortest word

The defining relation here has one word and it has four letters, so the design is called resolution IV. The number is the length of the shortest word, and it is the only summary of a fraction worth carrying because it says which orders of effect get mixed.

A word of length rr confounds any effect of order mm with an effect of order rmr - m. So:

  • Resolution III: a word of three letters confounds a main effect (m=1m=1) with a two-factor interaction (rm=2r-m=2). Main effects are contaminated by the interactions that are most likely to be real.
  • Resolution IV: main effects are confounded with three-factor interactions, which are usually small. Two-factor interactions are confounded with each other.
  • Resolution V: main effects with four-factor interactions, two-factor interactions with three-factor ones. Everything an experimenter usually wants is clear.
What a 8-run fraction of 4 factors confounds. The defining relation is I = ABCD, so the resolution is 4. AB is estimated as AB + CD; AC is estimated as AC + BD; AD is estimated as AD + BC; BC is estimated as BC + AD; BD is estimated as BD + AC; CD is estimated as CD + AB. Each of those is an identity about the design rather than an approximation about the data.
Fig. 2 The other half of the same design’s relation. At resolution IV the main effects are safe from two-factor interactions and the two-factor interactions are confounded in pairs: AB with CD, AC with BD, AD with BC. A design that finds a large interaction here cannot say which of the two it is.

What the contamination is worth

The alias chain says what is added; how much it matters is a separate number and it is computable exactly, because the bias is an expectation rather than a sample.

What a resolution-4 fraction reports, with every interaction at 0.8Each main effect's least-squares estimate is the sum of its whole alias chain, so with every two- and three-factor interaction equal to 0.8 the reported values are A = 5.80 against a true 5, B = 2.80 against a true 2, C = -0.20 against a true -1, D = 3.80 against a true 3. Nothing here is simulated: the bias is an expectation computed from the design.A5.80true 5, aliased with BCDB2.80true 2, aliased with ACDC-0.20true -1, aliased with ABDD3.80true 3, aliased with ABCresolution 4, every interaction 0.8the estimate is the chain, not the effect
Fig. 3 The eight-run resolution-IV fraction with every two- and three-factor interaction set to 0.8. Each reported main effect is its own value plus its chain, so A comes back as 5.80 against a true 5 and C comes back as −0.20 against a true −1. Nothing here is simulated.

The third row is the one to look at. The true effect of C is −1 and the design reports −0.20 — the same sign, a fifth of the size, and an experimenter reading it would conclude that C barely matters. One three-factor interaction of 0.8, which is a modest thing to suppose exists, has removed four fifths of a real effect.

Move the slider to zero and every bar is exact. That frame is the control rather than a decoration: it says the contamination comes from the interactions being real, not from the design being a fraction. A fraction is unbiased when its assumption holds, and its assumption is about quantities it cannot measure.

At interactions of 2 the reported effect of C is +1.00 — the wrong sign. There is no amount of data that reveals this, and no residual pattern that flags it, because the model fitted has no term for the thing that is contaminating it.

Every run, every effect, still

One property survives fractionation intact and it is the one that made the factorial worth having in the first place, so it deserves saying before the costs are counted.

In the eight-run fraction, every one of the eight runs contributes to the estimate of every effect. The estimate of AA is the average of the four runs where AA is high minus the average of the four where it is low — all eight runs, each used once. That is the hidden replication the essay that priced one factor at a time found: the one-at-a-time design would need two runs per effect and would estimate each from two observations, and this estimates each from eight.

So the fraction is not a retreat towards one-at-a-time. It keeps the whole of the efficiency argument and gives up something the one-at-a-time design never had either — it too confounds, worse and without a defining relation to say how. The comparison that matters is against the full factorial, not against the arrangement the factorial replaced.

What the design cannot test

The usual defence of resolution IV is that three-factor interactions are rare and small, and it is a good defence. It is also an assumption, and the point worth making is about its status rather than its plausibility.

A resolution-IV fraction has eight runs and estimates eight quantities: the intercept, four main effects, and three pairs of two-factor interactions. There are no degrees of freedom left. Even if there were, the three-factor interactions have no column of their own — their columns are the main effects’ columns — so there is nothing to test them against.

So the design is in the position a design that cannot see a curve is in for curvature, one order up: the thing that would invalidate the analysis is not merely unestimated, it is unrepresented. The repair there was centre runs, which buy one number back. The repair here is more runs of the same kind, which is to say a bigger fraction.

What a resolution-4 fraction reports, with every interaction at 0. Each main effect's least-squares estimate is the sum of its whole alias chain, so with every two- and three-factor interaction equal to 0 the reported values are A = 5.00 against a true 5, B = 2.00 against a true 2, C = -1.00 against a true -1, D = 3.00 against a true 3. Nothing here is simulated: the bias is an expectation computed from the design.
Fig. 4 The same design where the assumption holds exactly. Every reported main effect is its own value, to machine precision, on eight runs rather than sixteen. This is what a fraction buys when it is right, and the previous figure is what it costs when it is not.

Two fractions of the same size

The number of runs does not say what a design can estimate, and the cleanest demonstration is two designs with the same run count.

Every fraction worth running at 5 factors. the full factorial: 32 runs, nothing confounded; half — E = ABCD: 16 runs, resolution 5; half — E = AB: 16 runs, resolution 3; quarter — D = AB, E = ACD: 8 runs, resolution 3. Two fractions of the same size can differ by a whole level of resolution, so the number of runs does not say what a design can estimate.
Fig. 5 Five factors, four arrangements. The two sixteen-run halves differ by two whole levels of resolution — one is V and the other III — purely because of which word was chosen as the generator. Both cost the same sixteen runs.

E=ABCDE = ABCD gives the word ABCDEABCDE, length five, resolution V: main effects are clear of everything up to four-factor interactions, and two-factor interactions are clear of each other. E=ABE = AB gives the word ABEABE, length three, resolution III: the main effect of EE is confounded with the interaction ABAB, which is a quantity an experiment studying AA and BB has every reason to expect is real.

Sixteen runs each. One design answers the question and one does not, and nothing about the run count distinguishes them. The relation is the design. A fraction described only by its size — a half fraction was run — has not been described.

What a 8-run fraction of 5 factors confounds. The defining relation is I = ABD = ACDE = BCE, so the resolution is 3. A is estimated as A + BD + CDE + ABCE; B is estimated as B + AD + ABCDE + CE; C is estimated as C + ABCD + ADE + BE; D is estimated as D + AB + ACE + BCDE; E is estimated as E + ABDE + ACD + BC. Each of those is an identity about the design rather than an approximation about the data.
Fig. 6 The quarter fraction, where eight runs have to hold five factors. Every main effect is confounded with a two-factor interaction — A with BD, B with CE, and so on — which is resolution III, and it is the best that eight runs can do at five factors. The size forces it; no choice of generators escapes it.

Sixteen runs, and an effect that disappears

The resolution-III half fraction in the table above is worth following through, because it is the arrangement an experimenter would build by accident.

Five factors, sixteen runs, and the generator chosen as E=ABE = AB — which is what happens when the fifth factor is added late and assigned to a column that looked free. The defining word is ABEABE, so

E^ estimates E+AB,A^ estimates A+BE,B^ estimates B+AE\hat{E} \text{ estimates } E + AB, \qquad \hat{A} \text{ estimates } A + BE, \qquad \hat{B} \text{ estimates } B + AE

Three of the five main effects are confounded with two-factor interactions among themselves. If AA and BB interact — which is the thing a factorial is run to find out — then the reported effect of the newly added factor EE is that interaction, and the newly added factor will look important whether or not it does anything.

The failure has a particular shape that makes it hard to catch. The design is orthogonal, every diagnostic is clean, every standard error is what it should be, and the analysis of variance is perfectly well formed. The residuals carry no pattern, because the model has absorbed the interaction into a main effect rather than leaving it out. It is the same class as an estimator returning the edge of its own range: the output is a valid object and it is an object about something else.

Compare E=ABCDE = ABCD, the same sixteen runs, resolution V. There E^\hat E estimates E+ABCDE + ABCD, and a four-factor interaction is a quantity almost nobody has ever needed. The difference between the two designs is a choice made in an afternoon, and it decides whether the experiment can answer its question.

Where the fractions sit against one another

Three arrangements have now been priced in this field and it is worth putting them on one line, because they answer three different questions and get compared as though they answered one.

One factor at a time estimates each effect from two runs, misses interactions entirely and recommends settings it never tried. It is the arrangement with no design in it.

The complete factorial estimates every effect from every run and confounds nothing. It is what a fraction is a fraction of, and it costs 2k2^{k}.

A fraction estimates every effect from every run and confounds each with a set determined by its generators. Its efficiency per run is the factorial’s exactly, and what it additionally gives up is identifiability.

That second distinction is the one that gets blurred, so it is worth the arithmetic. In any design whose columns are orthogonal and coded ±1, the variance of every coefficient is σ2/N\sigma^{2}/N — the information matrix is NN times the identity and there is nothing else in it. So halving the runs doubles every variance, which is the ordinary and expected cost of a smaller experiment, and it is the cost an experimenter has already decided to pay.

The alias structure is a second cost on top of that, and it is of a different kind. A fraction is not a noisier factorial: its coefficients have exactly the standard errors that any orthogonal design of its size would give, and they are estimates of different quantities. More runs of the same fraction shrink the interval around the sum and never split it.

The practical form of that is a sentence worth carrying. A fraction’s precision improves with replication and its confounding does not, so an experimenter who suspects the alias structure is biting cannot answer the suspicion by running the same design again.

Four corners and 5 runs at the centre. The centre runs add two things at once. They estimate σ from replicates at one setting, which assumes nothing about the surface, and they supply the one contrast that sees curvature — the corner mean minus the centre mean, which estimates Σβᵢᵢ. It is one number: the design still cannot say which factor the curvature is in. 5 centre runs give 4 degrees of freedom for the pure-error estimate, and that is what sets the test's power.
Fig. 7 The two-factor design this field started from, for scale. At two factors there is nothing to fractionate: four corners, three effects, one degree of freedom. Everything in this essay begins at four factors, which is where 2k2^{k} first produces a number an experimenter wants to halve.

What a resolution buys, as a count

The three resolutions are usually described in words, and the words hide that they are counts of the same thing.

A word of length rr in the defining relation means that every effect of order mm is confounded with one of order rmr - m. So the question is my main effect safe is the question is r1r - 1 large enough that effects of that order are negligible, and the orders are ordered by how often anybody has seen one:

resolution main effects are mixed with two-factor interactions are mixed with
III two-factor interactions main effects
IV three-factor interactions each other
V four-factor interactions three-factor interactions

Reading down the first column is the standard advice and reading down the second is where the useful distinction is. Resolution IV protects the main effects and does nothing for the interactions: at four factors the pairs AB–CD, AC–BD and AD–BC are inseparable, so a design that finds a large interaction cannot say which of two it is. That is often exactly the finding an experiment is run to produce.

Resolution V is where both columns are safe, and at five factors it costs sixteen runs against thirty-two — a half rather than a quarter. The step from IV to V is the expensive one and it is the one that buys the answer to the second question.

What an experimenter should decide, and in which order

The arithmetic supports an order of decisions, and it is not the order the run count suggests.

First, which interactions can be assumed away. That is a subject-matter question, it is the assumption the whole fraction rests on, and it is not answerable from the data the design will produce. An experimenter who cannot name the interactions they are willing to confound has not yet chosen a design.

Then, the resolution that assumption implies. If two-factor interactions among the factors are plausible — which is the usual case — the answer is IV at minimum and V if the interactions are themselves of interest.

Then, the runs that resolution costs, which is the only step where the budget enters. At five factors, resolution V is sixteen runs and resolution IV is also sixteen; at six factors, V is thirty-two and IV is sixteen, and that is where the budget starts making the decision.

Reversing the order — starting from the affordable run count and taking whatever resolution comes with it — is what produces the E=ABE = AB design in the section above, and it is what the table of fractions is for. The two sixteen-run designs there cost the same and answer different questions, so a budget alone cannot pick between them.

This is the same shape as the decision the optimal-design field makes explicitly: a design is chosen against a criterion, the criterion encodes what the experiment is for, and starting from the size is starting from the one input that carries no information about the question.

What is claimed here, and what is not

The claim is what a fractional factorial confounds and what that costs: that the alias chain is an identity about the columns of the design rather than an approximation about the data; that the bias on each reported main effect is the exact sum of its chain, so a true effect of −1 is reported as −0.20 when every interaction is 0.8 and as +1.00 when every interaction is 2; that resolution is the length of the shortest word in the defining relation and decides which orders get mixed; and that two fractions of the same size can differ by two levels of it.

Nothing in this essay is simulated. Every number is a column identity or an expectation computed from the design.

One qualification on the bias numbers. They are computed with every interaction set to the same value, which is a stated construction rather than a realistic world: it makes the bias a function of how many contaminants a chain has and of nothing about which they are, which is what a figure about the design should show. Real interactions differ in size and sign, and a chain with two contaminants of opposite sign can come back unbiased by cancellation. The construction is an even-handed case rather than a worst one, and the worst case is what a design guarantees against the least favourable setting rather than what it does on average.

What stays out: how to choose generators to maximise resolution, which is a combinatorial problem with a well-developed literature and published tables; foldover designs, which add a second fraction to break specific alias chains and are the standard repair when a resolution-III screen finds something; Plackett–Burman designs, whose aliasing is partial rather than complete and whose chains are therefore fractions rather than identities; and the whole question of which interactions to expect, which is subject-matter knowledge and is what the choice of generators is supposed to encode.

Still open: a design that refuses a setting

Every design so far is built on a corner — the combination where every factor is at its extreme at once. The factorial is corners, the fraction is half the corners, and the central composite is corners plus axial runs.

A corner is sometimes not a setting. A temperature at the equipment’s limit combined with a pressure at the vessel’s limit; two ingredients each at their maximum in a formulation that must sum to one; a dose at the maximum tolerated combined with a duration at the maximum permitted. What the arrangement looks like when the corner is removed, and what predicting at a place the design never visits costs, is what a design that refuses the corners costs — where a fifteen-run design predicts the corner 1.84 times worse than a seventeen-run one that goes there.

The check, and the refusal

Two claims are gated. That every effect’s alias chain has exactly as many members as the defining relation has words, which is the group structure stated as a count and would fail on any error in the word list. And that the defining relation contains no word of length two, which would confound two main effects and produce a design nobody would call a design.

The refusal is the one that keeps the contamination honest: at a contamination of zero, every reported main effect must equal its own value to machine precision. A check that found bias at every setting of the slider would be reporting a defect in the arithmetic rather than a property of the fraction, and the zero frame is the only one that can tell the two apart. It is also the frame a check written to demonstrate the problem would have been most likely to leave out — which is the standing shape of a check that has never rejected anything.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Alias structureConfoundingDesign resolutionExperimental designFactorial designFractional factorialInteractionLeast squaresMain effectStudy design