Concept

Model diagnostics — where it appears

The checks made on a fitted model to see whether its assumptions hold — residual plots, influence measures, tests for structure left over. Their weakness is that a plot showing nothing is evidence only that this plot shows nothing, which is why twenty residual plots from a correct model are worth looking at.

Named by 13 essays across 8 fields — each of them below, with the objects they name alongside it.

Four datasets, slope 0.50, R² 0.67. Every one of these fits reports the same slope to two decimals and the same R². Only the first is a linear relationship with noise: the second is a curve, the third is a line with one outlier, and the fourth has its slope set by a single point.

Four datasets, one summary

Four datasets agree on slope, intercept and R² to two decimals. One is a linear relationship, one is a curve, one is a line with an outlier, and one has its slope set by a single point. The summary cannot tell them apart and neither can any other summary.

regression · Summary
What the plot says, and what the interval does, at n = 40. For each source: how often a quantile plot of the data leaves its pointwise band, and how often the 95% t interval for the mean misses. The two-lump source leaves the band on 100% of samples and its interval covers 94.80%; the t on three degrees of freedom leaves it on 57% and covers 95.73%, the best of the five.

The plot is about the wrong quantity

A t interval needs the sampling distribution of the mean to be normal, not the data. A two-lump source leaves its quantile band on 100% of samples of forty and its interval covers 94.80%; a t on three degrees of freedom leaves it on 57% and covers 95.73%, the best of five sources.

lineup · Qq
What each probe can see. How far apart the two components of the admissible set are on each probe, over the spread inside a component, on 100 designs whose set is enumerated and split. It is the population quantity a chain is trying to report. The separating direction itself reads 10.5646; the projected fourth power 5.0800, the design's own leverage 3.8362, the modelled active set 1.9529, the counted active set 2.0170 and a random direction in the same subspace 0.9422. The two active-set probes beat the random direction and lose to both of the earlier field's, which is the field's answer to the question that opened it.

A quantity that loses to a heuristic

Leverage is a heuristic about which units a balancing rule has most to say about. The constraint's active set is the thing the rule actually does. As a probe, the heuristic wins by 4.4 paired standard errors.

blocked · Randomisation
The window a whitening wants is not the memory of the errors. Regret under a five-period moving average as the tapered estimate is given more lags, over 120 draws at n = 120. The best window is L = 30; the automatic bandwidth is 4 and the error model's own likelihood chooses 9.7 on average. Both land in the same place and both are short, and the reason is the taper: a Bartlett weight at lag k is 1 − k/(L + 1), so a window of 8 keeps 0.556 of whatever the fourth lag carries and a window of 30 keeps 0.871. A window has to be several times the memory before it stops removing the memory. The dashed line is the rule told the errors are a first-order autoregression, which needs no window at all.

The window a whitening wants

Every law here is best whitened by a window several times longer than its own memory, including the one whose memory ends at the fourth lag. The three ways of choosing it from the sample all land in the same place, and it is the wrong one.

general · Order-selection
How often each probe finds a split that is there. The share of 34 designs — every one of them enumerated to be in two components — on which a two-chain test of 800 draws declares the split, by probe. The fourth power as the earlier fields use it finds it on 55.9%, so it misses 44.1% of the sets that have one. The same column projected off the rule's span finds it on 88.2%, and the separating direction itself on 91.2%. The design's own leverage, chosen without any dictionary, gets 79.4%. A random direction in the same subspace gets 44.1%, and the direction chosen for being concentrated gets 38.2% — worse than random, which is what a heuristic that finds the wrong structure looks like from the outside.

What a chosen probe finds

On a chain of eight hundred draws the probe the earlier fields use misses 44% of the sets that are split. Its own residual off the rule's span misses 12%, for one least-squares fit.

aimed · Randomisation
The threshold buys accuracy and spends exceedances. The mean squared error of the estimated shape against the threshold, split into the square of its bias and its spread, over 600 records of 2000 readings from a a normal parent. At the 0.9 quantile 199 exceedances are left, the bias is -0.1708, the spread is 0.0701 and the total error is 0.0341. The bias falls as the threshold rises because the exceedances get closer to being generalised Pareto; the spread rises because there are fewer of them. The sum is smallest at the 0.925 quantile, at 0.0340, of which 80.6% is still bias — so even the best threshold on this grid is one where accuracy, not spread, is the binding constraint.

The threshold is a dial

A peaks-over-threshold analysis has one knob, and raising it buys accuracy with exceedances. For a normal parent the error is smallest at the 0.925 quantile and 80.6% of it is still bias there — and both diagnostics practitioners use to set the knob lose to a fixed 0.90 rule, one by a factor of 1.590 and one by 11.881.

extreme · Extremes
How far apart four summaries put the quartet. Each bar is the largest value minus the smallest across the four datasets. Pearson spans 0.0018, which is the construction working. Distance correlation — the measure that is zero if and only if the variables are independent — spans 0.101, and Spearman spans 0.491.

The summary that was meant to work

Distance correlation is zero if and only if two variables are independent, which is exactly the guarantee a correlation coefficient lacks. Run on the four datasets that share a correlation, it spreads them by 0.10 — and Spearman, which guarantees nothing, spreads them by 0.49.

spread · Summary
What a design does to the residuals of a correct model. Every residual has standard deviation sigma times the square root of one minus its leverage. On this design the leverages run from 0.045 to 0.663, so the residual spreads differ by a factor of 1.68 — and the model is exactly right. The high-leverage point's residual averages 0.46 of the fitted spread where a typical point's averages 0.79.

Residuals are not the errors

A residual's standard deviation is σ√(1 − hᵢᵢ), so a design whose leverages run from 0.045 to 0.663 produces residuals whose spreads differ by a factor of 1.68 with the model exactly right. On the samples where the high-leverage point really did have the largest error, a raw residual plot shows it as the largest on 0.0% of them.

lineup · Qq
One of them is mostly leverage. How much of the design's own leverage direction each active-set probe carries, once both are standardised and projected off the rule's span — which is what a probe is, so it is the comparison that matters. Over 189 designs the modelled active set agrees with leverage at |r| = 0.8359 ± 0.0114 and the counted one at 0.4239 ± 0.0216. So the modelled probe is largely leverage under another name and the counted one is genuinely a different direction — and the counted one is the worse probe, at 0.3854 of alignment against 0.4272. What the active set contains beyond leverage points away from where the set splits.

Counting it exactly does not help

If a modelled active set lost because the model was crude, the exact one would win. It is computed at a cost no trial can pay, and it is worse — so the approximation was never what was costing the probe.

blocked · Randomisation
An AR(1) at φ = 0.5, 200 observations. The bars are the measured correlations; the curve is φᵏ, which is what an AR(1) must have. The band is ±1.96/√n, where an independent series would stay. The first bar is 0.53 against a band of ±0.14.

The check before the standard error

One number decides whether every interval in an analysis is trustworthy, and the check for it flags a lag-one correlation of 0.5 nine times in ten — and one of 0.2 only one time in five, where the interval already covers 88.6% instead of 95%.

timeseries · Dependence
Twenty residual plots from data where the model is exactly right, n = 24. Every panel is a correctly specified linear model with normal errors. The apparent curvature, funnelling and outliers are all produced by noise, and the largest single residual across the twenty is 2.13 standard deviations of the error. This is the reference nobody has when judging a real residual plot.

Twenty residual plots

Judging whether a residual plot looks wrong requires knowing what a correct one looks like, and almost nobody has seen twenty of those. Here they are, from a model that is exactly right, at the sample size that matters.

regression · Qq
The pairing recovers most of it and passes nothing. How much of the separating direction each probe carries, over 100 designs of 14 units whose admissible set is enumerated and split in two. The four rows the pairing adds are the dominant direction of what each blocking matrix keeps past its degrees, and the cut that direction's signs induce. Counted, they read 0.5266 and 0.5258 against the counted per-unit share's 0.3854 — most of the gap between that share and the design's own leverage at 0.5395, closed. Modelled, they read 0.4274 and 0.4954 against 0.4272. Nothing built from the active set passes leverage, and the projected fourth power is still ahead of all of them at 0.6583.

A set of pairs, not a vector

The active set is a graph on the units, and every probe built from it so far has been its degree. Read as a graph it recovers 0.1326 of the alignment the summary lost — and draws level with leverage rather than passing it.

blocked · Randomisation
Two far rows, and the line with one of them deleted. Twenty clean points and two rows near x = 9. The slope is −0.511 with every row, −0.376 with one far row deleted, and 0.495 with both deleted. Deleting one of them barely moves the line, because the other is still there.

Two points that hide each other

One far observation among twenty-one has a Cook's distance of 24.1. Put a second beside it and the two read 0.966 and 0.772, neither crossing 1, while together they reverse the slope and deleting both moves the fit by 53.3.

regression · Leverage

Named alongside it

The objects these essays reach for when they reach for this one.

LeverageExperimental designAssignment mechanismConnected componentCovariate balanceExact enumerationProjectionRandomisation testResidual plotHat matrixImbalanceMarkov chain Monte Carlo

All concepts