What a diagnostic plot is showing

A number for the shape

Five one-number summaries of a quantile plot's departure from its line, each a correct 5% test on forty normal observations, and no two agree about what matters. Straightness catches a skewed source 83% of the time and light tails 31%; kurtosis catches light tails 72% and skewness 0.2%. Reading all five and reporting whichever looks bad rejects 13.9% of genuinely normal samples. Correcting that search to 5% costs almost half the power against the departure it would have caught best — and a line-up of twenty panels gets the same correction for nothing.

Worth reading first: What normal actually looks like.

The band the eye was standing in for priced the confidence band software draws on a quantile plot: it holds each point at 95%, so a genuinely normal sample leaves it somewhere 45% of the time. It ended on the limitation no widening repairs. A band judges points one at a time, and the reading a person makes of a quantile plot is about shape — the ends curling away, an S through the middle, one point off on its own. It proposed a remedy and a doubt. The remedy: summarise the plot’s departure by a single number with an honest reference distribution. The doubt: which number, since each is powerful against some departures and blind to others, and a reader looking at a plot is running an unspecified search over all of them.

This essay runs the search on the page. It takes five such numbers, prices each against five departures, and then prices the search itself.

How often each one-number summary of a quantile plot detects five departures from normality, forty observations. Power at 5% against five departures. The straightness statistic detects the skewed source 83.3% of the time and light tails 30.8%; kurtosis detects light tails 72.5%; skewness detects light tails 0.2%. Rejecting when any of the five exceeds its own 5% point rejects 13.9% of normal samples; correcting that search to 5% gives it 37.4% power against light tails.
Fig. 1 The power of five one-number summaries of a quantile plot, each calibrated to reject 5% of normal samples of forty, against five departures from normality; and in the last two columns, rejecting when any of the five exceeds its own 5% point, and the same search corrected to 5% overall. Darker cells are higher power.

Five numbers, each a correct test

Each summary is computed on the sample standardised by its own mean and standard deviation, compared with the expected positions of forty normal order statistics — the same frame a quantile plot draws. Each rejects when it is large, and each has its 5% threshold calibrated on twenty thousand genuinely normal samples, so each is an exact 5% test.

  • Straightness: one minus the squared correlation between the sorted sample and the expected positions, which is Shapiro and Francia’s statistic. It asks whether the points lie on some straight line.
  • The worst point: the largest distance of any point from its expected position, in units of that point’s own pointwise band — the band a plotting program draws, read as a test.
  • The outer eighths: the five points at each end, their summed distance outward from the line. It asks whether the tails are too long or too short.
  • Skewness: the absolute sample skewness. It asks whether the plot bends one way at both ends.
  • Kurtosis: the absolute excess kurtosis. It asks whether the plot bends opposite ways at the two ends.

A sixth candidate, the summed squared distance between the plot and the line, turns out not to be one. On a standardised sample it equals a constant minus a multiple of the correlation in the first summary, so it is the straightness test exactly, under another name — a reminder that two summaries that look different on a plot can be one test.

None of them dominates

The hero table is the whole finding at a glance, and its pale cells are the point.

departure straightness worst point outer eighths skewness kurtosis
skewed (lognormal) 83.3% 73.5% 12.4% 83.2% 48.2%
heavy tails (t, 5 df) 35.8% 32.6% 10.0% 28.8% 34.0%
light tails (uniform) 30.8% 18.0% 49.0% 0.2% 72.5%
two lumps 26.4% 21.3% 51.1% 1.0% 60.0%
one wild point 58.5% 58.9% 6.8% 56.6% 58.4%

Every column has a row where another column beats it by more than twenty points. Straightness is the best or nearly the best summary against a skewed source and a wild point, and it catches light tails less than half as often as kurtosis does. Kurtosis is the best against light tails and lumps and catches the skewed source 35 points less often than straightness. The outer eighths catch two lumps better than anything but kurtosis and almost never notice a skewed source or a single wild point, because a standardised sample’s extreme points are pulled in by the very spread they inflate.

Skewness against a uniform source is the sharpest case. Its power is 0.2%, far below its 5% size: a uniform sample of forty has a sample skewness that is smaller and less variable than a normal sample’s, so a test for asymmetry rejects light-tailed data less often than it rejects normal data. A reader who checked skewness and found none has, for this departure, looked in exactly the one direction that is guaranteed to show nothing.

The heavy-tailed row is pale everywhere. A t distribution on five degrees of freedom has tails heavy enough to break a variance estimate now and then and, at forty observations, no summary catches it more than 36% of the time; the best, straightness, is barely ahead of the worst point and kurtosis. Heavy tails at this strength are a departure that forty observations mostly cannot see, whatever number is computed, and a plot that looks acceptable says less about them than about any other row. It is the departure for which the reassurance a clean quantile plot gives is least deserved, and — since a heavy tail is what inflates the occasional variance estimate — one of the more consequential to miss.

The worst-point statistic — the band itself, read as a test — is never the best and never the worst. That is what a pointwise band is: a summary built to catch any single point that strays, which makes it moderately sensitive to everything that moves a point and specifically sensitive to nothing.

The search a reader runs

Nobody looking at a quantile plot consults one of these. The eye checks whether the plot is straight, whether any point is far off, whether the ends curl, whether it bends one way or both — which is all five, and it reports whichever of them looks worst.

The false-alarm rate of reading one summary after another off a quantile plot of forty normal observations. Straightness alone rejects 5% of normal samples. Adding the worst point, the outer eighths, the skewness and the kurtosis, each at its own 5%, raises the rate to 5.0%, 6.8%, 11.1%, 12.3%, 13.9%. Five independent tests would reach 22.6%; these are correlated, so the search costs less than that and more than one test.
Fig. 2 The rate at which genuinely normal samples of forty are called non-normal when the summaries are consulted in turn, each at its own 5%, and the rate if the five were independent.

Consulting the summaries one after another, each at its own 5%, the false-alarm rate climbs from 5.0% for straightness alone to 6.8%, 11.1%, 12.3% and 13.9% with all five. Five independent tests would reach 22.6%; these are correlated — the worst point adds 1.8 points to straightness because the two largely look at the same thing, and the outer eighths add 4.3 because they look where the first two do not — so the search costs less than independence would and nearly three times one test.

That is twenty analyses of nothing in visual form. Each summary is a correct test; the family is not; and the reader who says “the plot looks fine” or “the plot looks off” has run the family and reported the most extreme member without saying so. Against the departures it produces a large number — the uncorrected search detects light tails 76.1% of the time — but at a false-alarm rate of 13.9%, not 5%.

What correcting the search costs

The search can be corrected. Take each summary’s null p-value, keep the smallest, and reject when it falls below the level at which the smallest of five null p-values falls only 5% of the time — here 1.78%. That restores a 5% test that looks in all five directions.

What each summary sees of one departure: light tails (uniform), forty observationsPower at 5%: straightness 30.8%, worst point 18.0%, outer eighths 49.0%, skewness 0.2%, kurtosis 72.5%. Any of the five at 5% each: 76.1%, at a false-alarm rate of 13.9%. The search corrected to 5%: 37.4%. A line-up of twenty read by a viewer running the same search: 38.3%.straightness30.8%worst point18.0%outer eighths49.0%skewness0.2%kurtosis72.5%any of the five, uncorrected76.1%the search, corrected to 5%37.4%a line-up of twenty, same search38.3%20,000 normal samples calibrate every thresholdthe uncorrected bar is not a 5% test
Fig. 3 Against light tails at forty observations: each summary’s power, the uncorrected search, the search corrected to 5%, and a line-up of twenty panels read by a viewer who runs the same search. The dashed line is 5%. The slider changes the departure.

The corrected search has power 37.4% against light tails, where kurtosis alone had 72.5% — it gives up almost half of what the right single summary would have found, because it spends its 5% across five directions and four of them are looking elsewhere. Against the skewed source it has 75.6% against straightness’s 83.3%; against two lumps 38.6% against kurtosis’s 60.0%. The price of not knowing which departure to look for runs from almost nothing to half the power against the departure that is there, and it is largest where the departure is one only a single summary can see — a skewed source, which three summaries catch, costs 9% of the best power; light tails, which one catches well, cost 48%. That is not a defect of the correction; it is the arithmetic of a search, the same that naming analyses in advance measured for twenty analyses.

It also says what a normality test should be chosen for. A study that knows its data are likely to be skewed — waiting times, concentrations, incomes — should test straightness or skewness and nothing else, at full power. A study worried about bounded or bimodal data should test kurtosis. A study that wants to catch anything should accept the corrected search’s lower power as the honest cost of wanting that, rather than run the uncorrected one and call it a 5% test.

What the line-up does instead

What normal actually looks like set the data’s quantile plot among nineteen made from genuinely normal samples and asked a reader to pick the odd one out. The arithmetic of that device is now easy to state exactly.

A reader who ranks twenty panels by any statistic and points at the most extreme picks the real data by chance with probability exactly 1/20, because under normality the twenty panels are exchangeable. That holds whatever the statistic is — straightness, kurtosis, the smallest of five p-values, or whatever the eye actually computes — and so the line-up’s false-alarm rate is 5% for every viewer, including one running an unspecified search. The correction the search needed is applied by the device, without anyone having to know what was searched.

Its power is close to the corrected search’s. A viewer running the five-summary search on a line-up picks light-tailed data 38.3% of the time, against the corrected test’s 37.4%; skewed data 69.7% against 75.6%; one wild point 50.8% against 56.6%. The line-up loses a few points because it compares the data with nineteen draws rather than with the whole null distribution. In exchange it needs no calibration, no list of summaries and no agreement about which one the reader used — which is why a procedure that looks less rigorous than a computed band has a rigour a band cannot get.

What each summary sees of one departure: skewed (lognormal, σ = ½), forty observations. Power at 5%: straightness 83.3%, worst point 73.5%, outer eighths 12.4%, skewness 83.2%, kurtosis 48.2%. Any of the five at 5% each: 87.5%, at a false-alarm rate of 13.9%. The search corrected to 5%: 75.6%. A line-up of twenty read by a viewer running the same search: 69.7%.
Fig. 4 The same comparison against the skewed source, where two summaries catch the departure more than four times in five and one catches it one time in eight.

What a normality check is for

The power table invites a question that it cannot answer by itself: detecting a departure from normality is only useful if something would be done differently once it is detected. For the commonest use of a quantile plot — deciding whether a t interval or test can be trusted — the plot is about the wrong quantity. A t interval needs the mean to be close to normal, not the data, and the departures a summary catches best are not the ones that damage the interval most: two lumps, which kurtosis catches 60% of the time at forty, leave a t interval’s coverage almost untouched, while a skewed source, which three summaries catch well, is the one that unbalances its two tails.

So the corrected search’s lost power is not always a loss worth regretting. A normality check used to decide whether a t test may be run is a pre-test, and a pre-test that switches methods on a noisy verdict is typically worse than committing to either method — the same finding a robust standard error produced for a different choice. The honest use of these summaries is descriptive: to say how the data depart, with the uncertainty that forty observations leave, and to choose a method for the question actually asked rather than for the shape.

That is also where the table is most useful. A reader who has decided that skewness is the departure that matters, because the analysis downstream is a one-sided bound or a mean of waiting times, can use the one summary that looks in that direction at full power — 83% at forty against a lognormal of this strength — and ignore the rest without paying for a search. A reader who has not decided should look at a line-up, which pays for the search automatically, and should expect to miss a departure a targeted summary would have caught.

Twenty panels and the reader’s own eye

The line-up’s arithmetic above assumed a reader who ranks panels by some statistic, and a human reader is not that; attention drifts, one panel’s oddity primes the next, and the statistic being used may change from panel to panel. The guarantee survives all of it. Under normality the twenty panels are exchangeable whatever the reader does, so the chance of picking the real data is exactly 1/20 for a reader who computes something, a reader who guesses, and a reader who changes their mind halfway. The power is where readers differ, and the table says how: a reader who knows which departure to look for is running a targeted summary rather than a search, and can expect the targeted summary’s power rather than the search’s. That is the case for looking at twenty residual plots from a correct model before judging one — not that it makes the eye a better search, but that it teaches the eye which direction the one plot in front of it should be read in.

A band on a standardised sample

One measurement here bears on the band that started this line of argument, and it is worth recording because it moves the number the band essay reported.

The pointwise band there was built for the order statistics of a normal sample with known mean and spread, and a sample drawn that way leaves it somewhere 45% of the time. A quantile plot of real data is almost always drawn after standardising by the sample’s own mean and spread, and a standardised sample hugs its line more closely: estimating the location and scale removes exactly the two directions in which a whole sample’s order statistics wander together. Measured on twenty thousand standardised normal samples of forty, the worst point’s 5% threshold sits at 1.002 half-bands — the pointwise band is, to three digits, already a 5% simultaneous band for a standardised sample of this size.

That does not rescue the band as a reader of shape; the table above shows the worst-point test weak against every departure except a wild point. It does say that the widening factor of 1.502 reported for known parameters is the right correction for a sample compared with a specified normal distribution and too large for one compared with the best-fitting normal. Which of the two a plotting program’s band assumes is worth checking before reading it, and it is rarely stated.

What five correct tests add up to, and what a line-up does not need

At forty observations each of five one-number summaries of a quantile plot is an exact 5% test, and none dominates: every one is beaten by another by more than twenty points on some departure, and skewness detects light tails 0.2% of the time.

Reading all five, each at its own 5%, rejects 13.9% of normal samples. The search corrected to 5% has 37.4% power against light tails where kurtosis alone had 72.5%, and 75.6% against a skewed source where straightness had 83.3%.

A line-up of twenty is a 5% test for every viewer and, read by a viewer running the same search, has power within six points of the corrected search on every departure measured.

Every threshold is calibrated on twenty thousand seeded normal samples of forty, and every power is counted on four thousand samples from each departure. The line-up’s power is computed exactly from the data sample’s null tail probability — the chance that nineteen null panels all fall below it — rather than by drawing line-ups.

Not claimed: that these five are the summaries the eye uses, or that the eye uses a fixed one. The line-up’s guarantee is precisely that it does not matter. Not claimed either that forty is typical; at larger samples every summary gains power and the pale cells darken, but the ordering of columns within each row — which summary sees which departure — is a property of the departures and changes little.

Still open: the plot of residuals

Every sample here is a single batch of observations. The quantile plot most often read is of a regression’s residuals, which are not identically distributed — their spreads differ by the design — and are correlated with each other through the fit. The null distribution of every summary above changes when it is computed on residuals, and it changes by an amount that depends on the design’s leverages.

The line-up handles that as well, if its null panels are made correctly: simulate new responses from the fitted model, refit, and plot those residuals, so that each null panel carries the design’s own distortions. Whether the five summaries keep their ordering on residuals, how much of the corrected search’s power a high-leverage design costs, and whether a line-up built this way loses more than a few points against the corrected test, are measurements of the same kind as the ones above, on a null that has to be simulated rather than assumed.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

KurtosisThe line-up testMultiple comparisonsNormalityQ–Q plotSkewnessStatistical powerVisual inference