Outlier — where it appears
Named by 2 essays across 2 fields — each of them below, with the objects they name alongside it.
Three runs at the end of the line
With one run at the top of a regression's range, a bad observation there and a line that bends there produce data with exactly the same distribution, so no residual, influence measure or test can say which happened. Move six of twelve runs to the ends, three at each, and a discrepancy of four σ is named correctly as a bad run 84.1% of the time and as a bend 91.6% — while the slope's standard error falls from 0.0836 to 0.0709. The design, not the diagnostic, decides whether the question has an answer.
Two numbers for the fit's geometry
No summary of the dependence between x and y can flag Anscombe's third and fourth datasets. A summary of the fit's own geometry can, and it takes two numbers of different kinds: the largest leverage, which reads the design before any y is seen and flags the fourth with certainty, and the largest deleted residual against a Bonferroni threshold, which flags the third at 1,192 and raises a false alarm on 4.78% of honest datasets. The rule of thumb most often quoted, Cook's distance over 4/n, flags the honest line too — and 66.45% of honest datasets like it.
Named alongside it
The objects these essays reach for when they reach for this one.
LeverageModel diagnosticsAnscombe quartetBonferroniCook's distanceExperimental designFalse positiveInfluenceLack of fitPure-errorResidual plotStudentised residual