What a diagnostic plot is showing

Three runs at the end of the line

With one run at the top of a regression's range, a bad observation there and a line that bends there produce data with exactly the same distribution, so no residual, influence measure or test can say which happened. Move six of twelve runs to the ends, three at each, and a discrepancy of four σ is named correctly as a bad run 84.1% of the time and as a bend 91.6% — while the slope's standard error falls from 0.0836 to 0.0709. The design, not the diagnostic, decides whether the question has an answer.

Worth reading first: What normal actually looks like · The line that one point drew.

Residuals are not the errors found that a residual carries only part of its error’s variance, and the part it loses is largest at the points with the most leverage: at a leverage of 0.663 the residual keeps 34% of the error’s variance and the fit absorbs the rest. It ended on the one thing standardising cannot repair — the information the point never gave — and on a claim about design: a design with replication at its extreme points can tell a bad observation from a wrong model, and a design without it cannot, whatever residual is plotted.

This essay makes that claim exact. It is stronger than it sounds. Without replication the two explanations are not merely hard to separate; they produce data with the same distribution, so there is nothing in the data to separate them by.

A discrepancy of 4 σ at the top of the range: one bad run, or a line that bends there, in two designs of twelve runs. The same twelve errors in every panel. With one run at the top (upper panels) a bad observation and a bent line produce identical data, point for point. With three runs at the top (lower panels) the bad observation leaves one run far from its two neighbours and the bent line moves all three together. The dashed line is the true line; the grey line is the least-squares fit.
Fig. 1 Two designs of twelve runs over x from 0 to 11, and two explanations of a four-σ discrepancy at the top of the range: one run there is bad, or the true line bends there. Every panel uses the same twelve errors. The dashed line is the truth before the discrepancy; the grey line is the least-squares fit; red points are the runs at the top.

Two explanations, one set of data

Take twelve runs, one at each of x=0,1,…,11x = 0, 1, \dots, 11, and a straight line with normal errors of known spread σ\sigma. Something goes wrong at the top of the range, where the response comes out four σ\sigma above the line. There are two ways that can happen.

One observation is bad. The instrument slipped, a sample was mislabelled, the run was recorded wrong. The line is fine; one error is not.

The model is wrong at that end. The response saturates, or a second mechanism switches on at high doses, and the true mean at x=11x = 11 is off the line. Every run made there would be off it.

With one run at x=11x = 11, the two explanations change the data in exactly the same way: they add four σ\sigma to the one number recorded there. The upper panels of the hero figure are not similar; they are identical, point for point, because the same errors were drawn and the same four σ\sigma was added to the same observation. A bad observation of size δ\delta at a single point and a bend of size δ\delta confined to that point are two stories about one vector of numbers. Their likelihoods are equal for every dataset, so every statistic has the same distribution under both, and no rule — residual plot, studentised residual, Cook’s distance, lack-of-fit test, a statistician’s eye — can call one more often under the first than under the second.

That is a statement about identifiability, not about power. More data at the other eleven points does not help, and a more careful diagnostic does not help. The design asked the top of the range one question, and one answer cannot say whether the question was misheard or the answer is surprising.

Three runs at the top

Move six of the twelve runs: three at x=0x = 0, three at x=11x = 11, and six spread over x=3x = 3 to 88. The same two explanations now leave different traces in the lower panels.

A bad observation moves one of the three runs at the top. The other two sit where the line predicts, and the bad one stands apart from its own replicates. A bend moves all three together, so they agree with one another and disagree with the line.

The two traces have two statistics, and under normal errors the statistics are independent of each other:

  • the spread of the replicates about their own mean, W=∑(yi−yˉtop)2/σ2W = \sum (y_i - \bar y_{\text{top}})^2 / \sigma^2, which is χ22\chi^2_2 if the three share a mean — whatever that mean is — and noncentral with λ=δ2⋅2/3\lambda = \delta^2 \cdot 2/3 if one of them is off by δ\delta;
  • the top mean against the line fitted to the other nine runs, which is shifted by δ\delta under a bend and by only δ/3\delta/3 under a bad observation.

A bend cannot move WW at all, because it moves the replicates together. That is the whole of the separation. WW is the classical pure-error sum of squares, the part of the residual variation that no choice of model can explain because it is variation among runs at the same setting; the rest of the residual sum of squares is lack of fit. With one run at each setting, pure error has zero degrees of freedom and the split does not exist. In the replicated design it has four — two from the three runs at each end — and those four are what remain available when the spread is not known in advance: an estimate of σ\sigma from the replicates alone, which a bad run at the top inflates and a bend does not touch. The known-spread arithmetic below is the clean version of the argument, and the unknown- spread version keeps its structure while paying for the four degrees of freedom in power.

How well the two are separated

A simple rule reads WW first — significant at 5%, call it a bad run — and otherwise calls a significant top mean a bend.

Telling a bad run from a bent line with three runs at the top of the range, and flagging either with one. At a discrepancy of four σ the replicated design names a bad run as one 84.1% of the time and a bend as one 91.6%; it calls a bend a bad run 5% of the time by construction and a bad run a bend 3.8%. The spread design flags a four-σ discrepancy 91.9% of the time whatever its cause, and can name the cause no better than a coin.
Fig. 2 With three runs at the top, the chance of each verdict under each explanation against the size of the discrepancy; with one run at each x, the chance that anything is flagged, which is the same whichever explanation is true.
discrepancy bad run named as bad bend named as bend bend called bad bad run called a bend one run at each x: flagged, either cause
2σ 28.95% 44.41% 5.00% 6.82% 38.97%
3σ 58.40% 76.42% 5.00% 6.47% 71.20%
4σ 84.10% 91.56% 5.00% 3.82% 91.91%
5σ 96.36% 94.70% 5.00% 1.26% 98.74%

At four σ\sigma the replicated design names a bad run correctly 84.10% of the time and a bend 91.56%, and it mistakes one for the other 5.00% and 3.82% of the time. The spread design flags a four-σ\sigma discrepancy 91.91% of the time — slightly more often than the replicated design flags a bad run, 87.92%, because three runs at the top dilute one bad one — and it says nothing about which explanation holds. Its chance of naming the cause is a coin’s, not approximately but exactly.

The errors in the replicated column have a structure worth noticing. Calling a bend a bad run happens at exactly the test’s 5% at every size, because a bend leaves WW central. Calling a bad run a bend falls with the size of the discrepancy, because a larger bad observation is ever more obviously apart from its siblings. The design turns an unanswerable question into a test with a known error rate.

What the six moved runs cost

Replication looks like a luxury — six of twelve runs spent on only two settings — and the natural worry is that it buys the diagnosis with the estimate. It does the opposite.

Twelve runs spread one to an x, against twelve with three at each end: what each design can do. Moving six runs to the ends lowers the slope's standard error from 0.0836 to 0.0709 and the largest leverage from 0.295 to 0.235, raises the power against a bend of 0.05(x − 5.5)² from 44.7% to 66.5%, and makes a bad run nameable. It observes 8 distinct x values instead of 12, so a departure confined to x = 1, 2, 9 or 10 is invisible to it.
Fig. 3 Six measures of what each design of twelve runs can do: the slope’s standard error, the largest leverage, the power against a gentle bend, the chance of flagging and of naming a four-σ bad run, and the number of distinct x values observed.

The slope is estimated more precisely. Its standard error falls from 0.0836 to 0.0709, because the slope’s precision is set by the spread of the x values and runs at the ends contribute most to it. No run carries as much leverage: the largest falls from 0.295 to 0.235, since three runs share the top where one held it alone — which is the lesson of the essay on residuals turned into a design rule, and it also keeps any single bad run from drawing the line itself. A smooth bend is easier to detect: against a quadratic departure of 0.05(x−5.5)20.05(x - 5.5)^2 the power rises from 44.7% to 66.5%, because a curvature term is estimated from the ends against the middle, and the ends are better known.

The cost is coverage. The replicated design observes eight distinct values of xx where the spread design observes twelve, and a departure confined to the settings it skips — x=1,2,9x = 1, 2, 9 or 1010 — is invisible to it, not weakly but entirely. That is the trade, stated exactly: replication makes the ends answerable and leaves gaps between them unasked. A study that has reason to expect a local feature in the interior — a threshold, a step, a narrow resonance — needs runs where the feature might be, and the two designs here are two answers to two different expectations.

What a residual plot recovers, and what it cannot

The diagnostics this line of essays has measured all work on the residuals, and it is worth being exact about which of them the design changes.

Twenty residual plots calibrated the eye against a correct model; the band the eye was standing in for priced its multiplicity; the essay on residuals showed that a raw residual understates exactly the influential points. Each of those is about seeing a discrepancy. None is about explaining one, and the design fact here sits beneath all of them.

When the high-leverage point really had the largest error, which plot shows it. Across designs from evenly spread to nine tenths clumped, on the samples where the high-leverage point's error really was the largest the model made. An evenly spread design shows it on 37% of those samples with raw residuals and 47% with standardised ones. At a clump of 90% the raw plot shows it on 0.0%.
Fig. 4 When the high-leverage point in a clumped design really did have the largest error, how often a raw and a standardised residual plot show it as the largest. Standardising recovers part of what the raw plot loses; neither can say why the point is large.

On the spread design, a four-σ\sigma discrepancy at the top produces a large studentised residual at the top run and a large Cook’s distance there; a careful reader sees it 92% of the time. What the reader then does — delete the run as an outlier, or add a curvature term to the model — is a choice with no support in the data at all. Deleting it is right under one explanation and hides a real feature under the other; adding the term is right under the second and fits noise under the first. The choice is usually made by what the analyst expected, and the residual plot, which looks like evidence, has contributed nothing to it.

On the replicated design the same residual plot has something to show. The bad run is a single point far from two neighbours at the same xx; the bend is three points together off the line. A reader looking at twelve points can see the difference, and the pure-error test makes the seeing into a number.

When the extreme is where the question is

The practical case for replicating the ends is strongest exactly where the ends matter most. A dose–response study’s highest dose is where toxicity or saturation would appear; a calibration’s extreme concentrations are where the instrument is least linear; an extrapolation leans entirely on the last point observed, as the prediction interval makes plain. In each case a surprising value at the extreme is either the most important finding in the study or the least interesting kind of error, and a single run there cannot say which.

The replicated design also answers a question the spread design never asks: is the spread of the errors the same at the ends as in the middle? Pure error at x=0x = 0 and at x=11x = 11 is an estimate of σ\sigma at each end with no model involved, and a heteroscedastic response shows up there directly. On the spread design the same question can only be asked of residuals, whose spreads already differ by the design with the model exactly right.

None of this requires a large study. Twelve runs were enough, and the rearrangement cost nothing in slope precision. What it requires is deciding before the runs are made that the ends are where a surprise would need an explanation — which is the same decision, in design form, as naming an analysis in advance: the evidence that can settle a question has to be collected before the question is asked.

The same fact in a factorial

Response-surface designs met this problem long ago and solved it in the same way, which is some evidence that the solution is not a quirk of straight lines. A two-level factorial puts every run at a corner, where a squared term equals one, so it cannot see curvature at all; the standard repair is a few runs at the centre. Those centre runs do two jobs at once, and they are the two jobs the replicates at the top did here. Their mean against the corners’ average is a test of curvature — the analogue of the top mean against the line — and their spread among themselves is pure error, an estimate of σ\sigma that no model assumption touches — the analogue of WW.

The factorial’s version makes one point clearer than the straight line does. The centre runs are replicates because the design has nowhere else to put an estimate of pure error that is free of the model, and the corners are unreplicated because each corner is already doing work in the effect estimates. A design that replicated only its corners would separate a bad corner from a curved surface in the way three runs at x=11x = 11 separate a bad run from a bend, at the cost of runs a factorial would rather spend elsewhere. Which of those is worth more is the same question as the straight-line one, asked in more dimensions, and the answer in both is that the design has to be chosen with the diagnostic in mind rather than the diagnostic chosen afterwards.

What to do with a spread design already run

Most data were not collected on a replicated design, and a reader with a surprising point at the end of a spread design still has to write something. The exact result above says what that something cannot be: a conclusion about the cause, drawn from the data.

What it can be is both answers. Fit the line with the point and without it; fit the line with a term that allows a bend at the top and without one. Report the slope and the prediction at the extreme under each, and say plainly that the data cannot choose between them. That is not a failure of analysis, and it is more useful to a reader than a confident deletion: it shows how much of the conclusion rests on a question the study did not ask. If the four readings agree, the point did not matter; if they disagree, the next study knows exactly where it needs three runs instead of one.

What replication at the end separates, and what the spread design keeps

With one run at the top of the range, a bad observation of size δ\delta there and a bend of size δ\delta confined there produce identical data from identical errors, so every diagnostic has the same distribution under both and none can name the cause better than chance.

With three runs at the top, the replicates’ spread about their own mean is central chi-square under a bend and noncentral under a bad run, and a rule reading it first names a four-σ\sigma bad run correctly 84.10% of the time and a four-σ\sigma bend 91.56%.

The replicated design estimates the slope more precisely — a standard error of 0.0709 against 0.0836 — carries a smaller largest leverage, 0.235 against 0.295, and detects a quadratic bend of 0.05(x−5.5)20.05(x - 5.5)^2 with power 66.5% against 44.7%. It observes eight settings instead of twelve and is blind to departures confined to the four it skips.

Every rate is exact, with the spread known: normal tail areas for the top mean, and a noncentral chi-square computed as a Poisson mixture of central ones for the replicates’ spread. The hero figure uses one seeded set of twelve errors in all four panels, and the spread design’s two panels hold the same numbers, point for point.

Not claimed: that the rule used is optimal, or that the spread is known in practice. With it estimated, from the interior residuals and the replicates at both ends, the pure-error test becomes an F test on few degrees of freedom and loses power, though the identity that separates the explanations — a bend cannot move the replicates’ spread — holds exactly with the spread unknown. And the bend here is confined to the top setting; a curvature that spreads across several settings leaves traces in the interior that a spread design can read, which is a different and weaker version of the question.

Still open: how many replicates, and where

Three runs at each end is one choice. Two would give the pure-error test one degree of freedom at each end and very little power; four would give more power at the top and fewer interior settings. The trade between the number of replicates at the extremes and the number of distinct interior settings has a best point for any stated pair of worries — a bad run at an end and a local feature somewhere inside — and it has not been found here.

There is a second version of the question that matters more for large studies. Replicating every setting is what makes pure error available everywhere, and it halves the number of settings for a given budget. Replicating only the settings with the highest leverage — the ends in a straight-line study, the corners in a factorial — may buy almost all the diagnostic value for a fraction of the cost, and whether it does is a calculation over designs that this essay has only started.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Experimental designInfluenceLack of fitLeverageModel diagnosticsOutlierPure-errorResidual plot