Residual plot — where it appears
Named by 7 essays across 5 fields — each of them below, with the objects they name alongside it.
R² is a property of the design
One line, one residual spread, five studies that differ only in how far apart they placed their x values. R² runs from 0.021 to 0.849 and the estimated residual spread is 0.993 in every one of them. Nothing about the relationship changed.
Four datasets, one summary
Four datasets agree on slope, intercept and R² to two decimals. One is a linear relationship, one is a curve, one is a line with an outlier, and one has its slope set by a single point. The summary cannot tell them apart and neither can any other summary.
The model that corrects its error
A cointegrated pair can always be written as a mechanism — today's change in y depends on yesterday's disagreement between y and its long-run relation with x. The coefficient of that disagreement is recovered from data that never saw it — and on unrelated series the same fit produces one a t table would call real 41% of the time.
What normal actually looks like
A single quantile plot of forty normal points wanders enough to look suspicious. Twenty of them, all genuinely normal, show what the noise looks like — and any single panel a reader would have rejected is in there.
The summary that was meant to work
Distance correlation is zero if and only if two variables are independent, which is exactly the guarantee a correlation coefficient lacks. Run on the four datasets that share a correlation, it spreads them by 0.10 — and Spearman, which guarantees nothing, spreads them by 0.49.
Residuals are not the errors
A residual's standard deviation is σ√(1 − hᵢᵢ), so a design whose leverages run from 0.045 to 0.663 produces residuals whose spreads differ by a factor of 1.68 with the model exactly right. On the samples where the high-leverage point really did have the largest error, a raw residual plot shows it as the largest on 0.0% of them.
Twenty residual plots
Judging whether a residual plot looks wrong requires knowing what a correct one looks like, and almost nobody has seen twenty of those. Here they are, from a model that is exactly right, at the sample size that matters.
Named alongside it
The objects these essays reach for when they reach for this one.
Model diagnosticsLeverageSummary statisticsAnscombe quartetCorrelationExperimental designNormalityQ–Q plotSample sizeVisual inferenceAutocorrelationCointegration