Concept

Residual plot — where it appears

The residuals of a fitted model drawn against the fitted values or a predictor, looked at for structure the model has missed. A single plot showing nothing is weak evidence, which is why the honest comparison is against a wall of plots from a model that is known to be correct.

Named by 7 essays across 5 fields — each of them below, with the objects they name alongside it.

One relationship at five designs, residual spread 1.00. Every panel has the same slope of 1, the same intercept of 0 and the same residual standard deviation of 1.00. Only the range of x differs. R-squared runs from 0.021 to 0.849, and the estimated residual spread is 0.9932 in all five.

R² is a property of the design

One line, one residual spread, five studies that differ only in how far apart they placed their x values. R² runs from 0.021 to 0.849 and the estimated residual spread is 0.993 in every one of them. Nothing about the relationship changed.

spread · Summary
Four datasets, slope 0.50, R² 0.67. Every one of these fits reports the same slope to two decimals and the same R². Only the first is a linear relationship with noise: the second is a curve, the third is a line with one outlier, and the fourth has its slope set by a single point.

Four datasets, one summary

Four datasets agree on slope, intercept and R² to two decimals. One is a linear relationship, one is a curve, one is a line with an outlier, and one has its slope set by a single point. The summary cannot tell them apart and neither can any other summary.

regression · Summary
The correction, at a generating α of -0.2. Each point is one step: the gap at the end of yesterday against the change in y today. The fitted slope is -0.202 against the -0.2 the data was generated from, which means 20% of any disagreement between y and its long-run relation with x is undone in a single step. A shock therefore has a half-life of 3.1 steps. Neither series is stationary; the relation between them is.

The model that corrects its error

A cointegrated pair can always be written as a mechanism — today's change in y depends on yesterday's disagreement between y and its long-run relation with x. The coefficient of that disagreement is recovered from data that never saw it — and on unrelated series the same fit produces one a t table would call real 41% of the time.

cointegration · Dependence
Twenty samples of 40, every one of them genuinely normal. Each panel is a quantile-quantile plot of 40 draws from a normal distribution. The worst point in the worst panel sits 0.87 standard deviations off the line. Anything a reader would reject here would be a false alarm.

What normal actually looks like

A single quantile plot of forty normal points wanders enough to look suspicious. Twenty of them, all genuinely normal, show what the noise looks like — and any single panel a reader would have rejected is in there.

normal · Qq
How far apart four summaries put the quartet. Each bar is the largest value minus the smallest across the four datasets. Pearson spans 0.0018, which is the construction working. Distance correlation — the measure that is zero if and only if the variables are independent — spans 0.101, and Spearman spans 0.491.

The summary that was meant to work

Distance correlation is zero if and only if two variables are independent, which is exactly the guarantee a correlation coefficient lacks. Run on the four datasets that share a correlation, it spreads them by 0.10 — and Spearman, which guarantees nothing, spreads them by 0.49.

spread · Summary
What a design does to the residuals of a correct model. Every residual has standard deviation sigma times the square root of one minus its leverage. On this design the leverages run from 0.045 to 0.663, so the residual spreads differ by a factor of 1.68 — and the model is exactly right. The high-leverage point's residual averages 0.46 of the fitted spread where a typical point's averages 0.79.

Residuals are not the errors

A residual's standard deviation is σ√(1 − hᵢᵢ), so a design whose leverages run from 0.045 to 0.663 produces residuals whose spreads differ by a factor of 1.68 with the model exactly right. On the samples where the high-leverage point really did have the largest error, a raw residual plot shows it as the largest on 0.0% of them.

lineup · Qq
Twenty residual plots from data where the model is exactly right, n = 24. Every panel is a correctly specified linear model with normal errors. The apparent curvature, funnelling and outliers are all produced by noise, and the largest single residual across the twenty is 2.13 standard deviations of the error. This is the reference nobody has when judging a real residual plot.

Twenty residual plots

Judging whether a residual plot looks wrong requires knowing what a correct one looks like, and almost nobody has seen twenty of those. Here they are, from a model that is exactly right, at the sample size that matters.

regression · Qq

Named alongside it

The objects these essays reach for when they reach for this one.

Model diagnosticsLeverageSummary statisticsAnscombe quartetCorrelationExperimental designNormalityQ–Q plotSample sizeVisual inferenceAutocorrelationCointegration

All concepts