What a diagnostic plot is showing

Residuals that look too normal

The quantile plot most often read is of a regression's residuals, and residuals look more normal than the errors behind them. Each residual is a mix of all the errors, and the mix moves towards a bell as the fit takes more columns. On forty rows with skewed errors, the straightness test flags 82.9% of samples of the errors themselves, 79.1% of the residuals of a line, 71.4% with five columns, 51.2% with ten and 23.0% with twenty — far below the 49.5% that twenty errors would give, so the fit costs more than the observations it uses. Light tails vanish faster: 73.0% on the errors, 4.7% on twenty-column residuals. The band's level holds at every design, and one point with leverage 0.73 costs almost nothing; what costs is the number of columns.

Worth reading first: What normal actually looks like.

A number for the shape priced five one-number summaries of a normal quantile plot’s departure from its line, each a correct 5% test on forty normal observations, and found that none dominates: straightness catches a skewed source 83% of the time and light tails 31%, kurtosis the reverse, and reading all five and reporting whichever looks bad rejects 13.9% of normal samples. Every sample there was a batch of observations, drawn independently.

The quantile plot most often drawn is not of a batch. It is of a regression’s residuals, which are not the errors: their spreads differ by the design and they are correlated with each other through the fit. The essay ended by asking whether the five summaries keep their level on residuals, and how much of their power a high-leverage design costs. The level turns out to be nearly untouched and the power not — and the design feature that costs power is not the one the question named.

How often a quantile plot's shape test sees a skewed error distribution, on the errors and on residuals of fits of more columnsOn the forty errors themselves the test flags 82.9%. On the residuals of fits with 2, 5, 10, 20 columns: 79.1%, 71.4%, 51.2%, 23.0%. On 40 − p errors, which is what losing the fitted columns' worth of observations alone would give: 80.6%, 76.1%, 69.1%, 49.5%.00.2500.5000.75010251020columns fitted to forty rowsshare of samples the straightness test flagsthe residuals of the fit40 − p errors, unmixedall forty errors3,000 samples a point, each read at its own 5% pointfitting mixes the errors towards normal
Fig. 1 How often the straightness test sees a skewed error distribution in forty rows, read on the errors themselves, on the residuals of least-squares fits with more and more columns, and on 40 − p errors — what simply losing the fitted columns’ worth of observations would cost. The slider changes the error distribution.

What a residual is made of

A least-squares fit’s residuals are the errors multiplied by the matrix I−HI - H, where HH is the hat matrix. Each residual is therefore a weighted sum of every error in the data — mostly its own, with a little of every other row’s subtracted — and the weights on the other rows grow with the number of columns the fit uses. With two columns, an intercept and a slope, each residual is its own error less a small share of the rest. With twenty columns on forty rows, the hat matrix has trace twenty, half of every residual’s variance has been absorbed by the fit, and what is left is a thorough mixture.

Mixtures of independent draws move towards the normal, which is the central limit theorem in a place nobody put it. A residual plot is a plot of partial sums of errors, and the more of each residual is made of other rows’ errors, the more normal the plot looks, whatever the errors are.

The band’s level holds

The first question the essay on the shape left was whether the five summaries, read on residuals against the band for forty independent normal observations, still hold their 5% level. They do, very nearly.

The level of the five shape tests on residuals, read against the band for forty independent normals, by design. a line, even x: 5.8%, 5.4%, 4.5%, 5.4%, 5.3%; a line, lognormal x: 5.7%, 5.4%, 4.8%, 5.3%, 5.7%; a line, one far point: 6.0%, 6.1%, 4.8%, 5.8%, 5.2%; five columns: 5.9%, 5.7%, 4.5%, 5.1%, 5.7%; ten columns: 4.6%, 5.1%, 5.4%, 4.4%, 5.0%; twenty columns: 4.8%, 5.1%, 5.3%, 4.9%, 4.8%.
Fig. 2 The share of samples with normal errors in which each of the five shape tests rejects, read on the residuals of six designs against the 5% points for forty independent normals. Each bar spans the five tests’ rates for one design.

Across six designs — a line with an evenly spread covariate, a line with a lognormal covariate, a line with one point at leverage 0.73, and fits with five, ten and twenty columns — the five tests reject between 4.4% and 6.1% of samples with normal errors. The largest excesses are on the line with one far point, where the residual at that point has only a quarter of the others’ variance and the standardised residuals are a little uneven; the largest shortfalls are on the ten-column fit. None is far enough from 5% to matter.

The reason is the same mixing. Normal errors mixed by I−HI - H are still normal, jointly: the residuals are exactly a normal vector with a singular covariance, and after standardising by their own sample spread they look, to a shape functional, very much like a smaller normal sample. So the band drawn for independent normals is close to right for them, and a line-up whose null panels are simulated from the fitted model — the correct construction — makes almost no difference to the level. At twenty columns, the straightness test calibrated on the design’s own residuals flags a skewed error 22.1% of the time against 23.0% calibrated on independent normals.

The power does not

On the errors themselves, the straightness test flags a skewed error distribution — lognormal with σ = 0.5 — in 82.9% of samples of forty. On the residuals of a line it flags 79.1%; of five columns, 71.4%; of ten, 51.2%; of twenty, 23.0%. The test has not changed and the errors have not changed. The fit has made them look normal.

Forty skewed errors and the residuals of a 20-column fit to them, on one normal quantile plot. The errors are lognormal with σ = 0.5. Their straightness statistic is 0.1259 and their skewness 1.161; the residuals of a 20-column least-squares fit to the same errors give 0.0259 and 0.276.
Fig. 3 One sample of forty skewed errors on a normal quantile plot, beside the residuals of a twenty-column fit to the same errors. The errors bend away from the line in both tails; the residuals lie along it.

In the single sample drawn above, the errors’ straightness statistic is 0.1259 and their skewness 1.161; the residuals of a twenty-column fit to the same errors give 0.0259 and 0.276. The skewness that was plainly there is a quarter of its size after the fit.

How often a quantile plot's shape test sees a light-tailed error distribution, on the errors and on residuals of fits of more columns. On the forty errors themselves the test flags 73.0%. On the residuals of fits with 2, 5, 10, 20 columns: 64.2%, 42.7%, 17.7%, 4.7%. On 40 − p errors, which is what losing the fitted columns' worth of observations alone would give: 69.1%, 62.3%, 56.9%, 36.8%.
Fig. 4 The same comparison for uniform errors read by the kurtosis test. The errors’ light tails are the first thing a fit mixes away: at twenty columns the test flags residuals no more often than it flags normal ones.

Light tails suffer most, because a light-tailed distribution is the one a mixture most quickly turns into a bell. The kurtosis test flags uniform errors in 73.0% of samples of the errors, and in 64.2%, 42.7%, 17.7% and 4.7% of residuals from fits of two, five, ten and twenty columns — at twenty columns, no better than chance. Heavy tails are the most robust to mixing, since one very large error dominates any sum it is in: the worst-point test flags errors from a t distribution on five degrees of freedom in 33.6% of samples of errors and 13.2% of twenty-column residuals.

It costs more than the columns’ worth of data

A natural guess is that a fit with pp columns leaves n−pn - p observations’ worth of information about the errors’ shape, so the residuals of twenty columns on forty rows should behave like twenty errors. They behave much worse. Twenty independent skewed errors are flagged by the straightness test 49.5% of the time; the residuals of a twenty-column fit, 23.0%. Thirty errors are flagged 69.1% of the time; ten-column residuals, 51.2%.

For light tails the gap is wider still. Twenty independent uniform errors are flagged by the kurtosis test 36.8% of the time and thirty of them 56.9%; the residuals of twenty and ten columns, 4.7% and 17.7%. A fit of twenty columns has not left the equivalent of twenty errors to look at. For this purpose it has left almost nothing.

The difference is that n−pn - p unmixed errors keep their shape, and residuals do not have n−pn - p unmixed errors in them: they have nn mixtures, each of all nn errors, constrained to a subspace of dimension n−pn - p. The information about the errors’ mean and slope is removed by the fit, as the counting guess supposes; the information about their shape is spread thinly across every residual, and the mixing dilutes it further than the dimension count suggests.

Which summary survives the mix

The five summaries lose power at different rates, and the order is instructive. Against skewed errors at ten columns, the skewness test keeps 55.2% of samples against its 82.9% on the errors, straightness 51.2% against 82.9%, the worst-point test 42.9% against 73.4%, kurtosis 31.2% against 48.5%, and the outer-tails sum 5.5% against 12.1%. The two that read the whole distribution’s lean — skewness and straightness — lose least, because a mixture of skewed errors is still somewhat skewed; the tests that read the extremes lose most, because the extremes are exactly what mixing averages into the middle.

So the ordering what each number sees established on raw samples is not preserved on residuals. A viewer who learned to look at the tails of a quantile plot for trouble is looking at the part of a residual plot the fit has most thoroughly smoothed.

What each summary sees of one departure: light tails (uniform), forty observations. Power at 5%: straightness 30.8%, worst point 18.0%, outer eighths 49.0%, skewness 0.2%, kurtosis 72.5%. Any of the five at 5% each: 76.1%, at a false-alarm rate of 13.9%. The search corrected to 5%: 37.4%. A line-up of twenty read by a viewer running the same search: 38.3%.
Fig. 5 The five shape summaries’ power against each departure on forty independent observations, from the essay on the shape — the baseline every residual rate above is measured against.

One far point is not the problem

The essay on the shape expected a high-leverage design to cost the most, and it costs very little. A line whose covariate has one point at leverage 0.73 — a point whose own value accounts for nearly three quarters of its fitted value — flags skewed errors 80.6% of the time on its residuals, against 79.1% for an evenly spread covariate and 82.9% on the errors. A lognormal covariate, with its largest leverage 0.35, gives 80.0%.

The high-leverage point’s residual is small, because the fit follows it, and that one residual carries almost no information about its error. But it is one residual of forty. The other thirty-nine are mixed about as little as an even design mixes them, and they carry the shape. What spreads the damage across every residual is the number of columns, since each column adds a direction the fit removes from every row at once.

That reverses the usual ordering of worries. The points a slope rests on found a single high-leverage point deciding a slope; for the residual plot’s view of the errors’ shape it is nearly irrelevant, and a model with many covariates — a fit that looks unremarkable, with no point of high leverage — is the one whose residual plot says least.

Why the line-up does not rescue it

The band the eye was standing in for and the essay on the shape both found the line-up a clean repair for the reading of a quantile plot: twenty panels, nineteen from the null, and the viewer picks the odd one out, which holds exactly one in twenty for any reading the viewer uses. For residuals the correct line-up simulates its null panels from the fitted model — normal errors, the same design, refitted — so that each null panel carries the design’s distortions.

That repairs the level, which was not broken, and cannot repair the power, which was. The real panel is a plot of mixed errors, and the null panels are plots of mixed normal errors; mixing has moved the real panel towards the nulls, and no choice of null panels undoes the mixing. A line-up of residuals from a twenty-column fit asks the viewer to find a panel that the fit has already made look like the others.

Whether it matters that the errors look normal

The residual plot is usually drawn to check an assumption, and it is worth asking which assumption. The plot is about the wrong quantity found that a t interval needs the sampling distribution of its estimate to be normal, not the data: a two-lump source fails every quantile band and its interval still covers 94.80%. A regression coefficient is a weighted sum of all the errors, which is the same mixing that makes the residuals look normal, so when the residual plot has lost its power to see the errors’ shape, the coefficient’s sampling distribution is usually close to normal for the same reason.

That softens the loss for the most common use of a residual plot and sharpens it for the others. A model used for inference about its coefficients can tolerate errors the residual plot cannot see, because the coefficients mix the errors as the residuals do. A model used for prediction intervals, or for any statement about individual future observations — a band allowed a few misses was one — cannot, because a single future observation carries its own error unmixed, and that error’s shape is exactly what the residuals no longer show.

Why the columns count and the far point does not

The two design features do different things to the hat matrix. A far point makes one diagonal element of HH large and leaves the others small, so one residual is heavily mixed — it is mostly the fit’s echo of its own error — and thirty-nine are not. More columns make every diagonal element larger at once, since the trace of HH is the number of columns, and every residual is mixed a little more. A penalty is a trace found the same trace governing how much a fit flatters itself; here it governs how much of each error’s shape the fit has absorbed, and the absorption is spread over every row.

What a design can do for its own residual plot

A study that wants its errors’ shape to be checkable can arrange for it before the data exist. The residuals of a fit mix every error; the deviations of replicated runs from their own group means mix only the errors within each group, and they are untouched by how many columns the model has. Three runs at the end of the line found replicates at the extremes separating a bad run from a bend in the relation; the same replicates supply a pure-error sample whose shape can be read without the fit’s mixing, because a group of three deviations from their mean is mixed with two other errors rather than with forty.

The same arithmetic says what the alternatives cost. A study that fits twenty columns to forty rows has spent half its rows’ information on the model and, it turns out, more than half of its ability to see the errors’ shape; a study that fits five columns and replicates keeps most of both. Where the errors’ shape matters — prediction, tolerance limits, any statement about an individual rather than an average — it is worth designing for, in the same way the design a report could state is worth stating: before any outcome, from the design alone.

What a residual plot can and cannot be asked

It can be asked about a single wild observation. Heavy tails and outliers survive mixing best, because one large error is large in every residual it enters. A residual plot remains a reasonable place to look for a recording error or a contaminated row, and two numbers for the fit’s geometry gave the deleted residual that reads one row at a time for exactly that job.

It cannot be asked whether the errors are normal, once the fit has many columns. At twenty columns on forty rows, residuals from uniform errors pass the kurtosis test as often as residuals from normal ones do. A clean residual plot from a large model is a statement about the model’s size, not about the errors.

It should be read with the ratio of columns to rows in mind. Up to a few columns the residual plot sees most of what an error plot would; at a quarter of the rows it has lost two fifths of its power against skewness and three quarters against light tails; at half the rows it has lost most of both.

The level is not where to spend effort. A carefully simulated null band for residuals is right and hardly differs from the ordinary one. The effort belongs to the question the plot is being asked and whether a residual plot can answer it at the model’s size.

What is measured here and what is not

Read against the band for forty independent normals, the five shape tests reject between 4.4% and 6.1% of normal-error samples on residuals of six designs.

On forty rows, the straightness test flags skewed errors in 82.9% of samples of the errors and 79.1%, 71.4%, 51.2% and 23.0% of residuals from fits of two, five, ten and twenty columns; twenty unmixed errors give 49.5%.

A line with one point at leverage 0.73 flags skewed errors in 80.6% of residual samples.

Every rate is counted over three thousand samples a point with each test read at a 5% point simulated from eight thousand samples, on designs fixed once from stated seeds; the residuals are computed exactly from each design’s own I−HI - H.

Not measured: studentised residuals, which rescale each residual by its own leverage but cannot unmix them; designs with many more rows than columns, where the mixing is slight; and generalised linear models, whose residuals have no single natural definition and whose plots mix differently.

Still open: a plot that unmixes

The loss comes from the mixing, so the natural question is whether residuals can be transformed back towards something unmixed. There are constructions that produce n−pn - p residuals that are exactly independent and identically distributed under normal errors — recursive residuals, computed one row at a time as each new observation is predicted from the fit to the rows before it — and they are unmixed in the sense that each depends on one new error and a finite past.

Whether a quantile plot of recursive residuals recovers the power the ordinary residuals lose, how much the ordering of the rows matters to it, and whether it keeps the level for a design with a far point, are measurements of the same kind as the ones above and have not been made here.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Hat matrixLeverageModel diagnosticsNormalityQ–Q plotResidualStatistical powerVisual inference