the-normal-law-and-where-it-stops

The normal density at sigma = 1.00

The bands hold 68.27%, 95.45%, 99.73% of the mass. Those figures are integrals of the curve drawn, not the memorised 68-95-99.7.

The distribution itselfslider: the standard deviation, 0 positionswide40 views

What else it draws

The same object, drawn to answer the other questions the essays put to it.

The bands hold 68.27%, 95.45%, 99.73% of the mass. Those figures are integrals of the curve drawn, not the memorised 68-95-99.7.

The source is one-sided and skewed. At n = 8 the standardised sum has skew 0.695, and the theory says 2/sqrt(n) = 0.707 — so the convergence is visible AND its rate is predicted.

Relative error against the exact binomial. At n = 1280 the error at the median is 0.96% and three sigma out it is 25.7% — a factor of 27. The tail is where the approximation is used.

Approximate tail probability divided by the exact one. At one sigma the ratio is 0.98; at 4 sigma it is 0.104, so a rare event is understated by a factor of 10.

Each panel is a quantile-quantile plot of 40 draws from a normal distribution. The worst point in the worst panel sits 0.87 standard deviations off the line. Anything a reader would reject here would be a false alarm.

The two-sided 95% critical value is 2.571 for t(5) and 1.960 for the normal — 31% wider. Using the normal at this sample size makes every interval too short by that much.

At δ = √n·μ/σ = 1 the exact law of the squared estimate has mean 2.00, variance 6.00 and skewness 2.177; the delta method's normal has mean 1.00, variance 4.00, no skewness, and 30.85% of its mass below zero, where a square cannot go. The Kolmogorov distance between them is 0.3085.

Each dot is one pair (zx, zy) drawn around (3, 1). Outside the horizontal band |zy| > 1.96 Fieller's set is a bounded interval, with probability 17.01%; inside the band and outside the disc of radius 1.96 it is everything outside an interval, 75.03%; inside the disc it is the whole line, 7.96%. On 40,000 counted draws Fieller covers ρ = 3.00 95.21% of the time and the delta interval 82.48%.

The content of the band is a random variable. Across 20,000 normal samples of 10 it averages 91.1%, its fifth percentile is 74.7%, and it falls short of 95% on 59.9% of samples. The band that is drawn to show where 95% of the data lies.

At ten observations it is 3.382 sample standard deviations, against the 1.96 an interval for the mean uses. The two only converge in the hundreds: at n = 300 the factor is still 2.106.

Half-widths in sample standard deviations: 0.468 for the mean, 2.14 for one future observation, 2.75 to hold 95% of the population. The first two differ by exactly the square root of n + 1, which is 4.58 here.

The interval between the extremes of n draws holds at least 95% of the population with probability 1 - n p^(n-1) + (n-1) p^n, whatever the population is. Reaching 95% confidence takes 93 observations.

The chance that the smallest and largest of 20 observations hold at least 90% of the population, counted under a normal, an exponential, a uniform and a t on three degrees of freedom. The closed form gives 60.8% and knows nothing about which.

How often the band holds the 95% of the population it promises. On the normal it is 95.1%, as designed. On a Laplace it is 78.4%, and on a uniform it is 99.5% — conservative rather than short, which is the direction nobody predicts.

Both tails should be 2.5%. At 8 observations the interval falls short of the mean on 9.75% of samples and overshoots on 0.31%. At 500 they are 3.31% and 2.05%, and the total is 5.36% — which a coverage table reports as very nearly right.

Both bars should read 2.5%. The three symmetric sources — normal, Laplace and a t on five degrees of freedom — are balanced whatever their tails do. The two skewed ones are not: the exponential's misses split 7.9 to 1 and the lognormal's 61 to 1.

Student's t is derived on the assumption that the sample mean and the sample variance are independent. Here their correlation is 0.7023 at 8 observations and 0.7093 at 800 — a constant rather than a small-sample effect.

Forty units split between two groups, with the second group's variance a stated multiple of the first's, and the two population means equal. A 5% test should reject 5% of the time. The pooled test runs from 0.55% to 18.91% across this region; Welch's runs from 4.63% to 5.51%.

The pooled test has 38 degrees of freedom at every split, because it counts them. Welch computes them from the two estimated variances, so they are not integers and they differ from sample to sample — the standard deviation of 8-and-32's is 7.36.

The outer band is left by 5% of genuinely normal samples — which is what a reader is using a band for. The inner one holds each point separately at 95%, which is what software draws, and 45.0% of genuinely normal samples step outside it. The outer is the inner widened by a factor of 1.502.

For each source: how often a quantile plot of the data leaves its pointwise band, and how often the 95% t interval for the mean misses. The two-lump source leaves the band on 100% of samples and its interval covers 94.80%; the t on three degrees of freedom leaves it on 57% and covers 95.73%, the best of the five.

The distance between the exact distribution function and the normal one peaks at 0.0133, at z = -0.01. The Berry–Esseen bound is 0.1146, 8.62 times the real worst error, and larger than the whole 2.5% tail a two-sided test reads.

The bound divided by the largest distance between the exact distribution function and the normal one. A fair coin reads 1.19, so the bound is nearly attained; an exponential source reads 8.62 and a gamma of shape 16 reads 23.61.

The bound on the distance from the normal falls as one over the square root of the sample size. It drops below a 2.5% tail at 2,103 draws, below 0.5% at 52,573 and below 0.05% at 5,257,207. Before each of those points it cannot vouch for that tail at all.

Each curve is an approximation divided by the exact gamma tail, so 1 is exact. Six standard deviations out at n = 10: the normal gives ×0.0000670, one Edgeworth term ×0.00159, two ×0.0167 and the saddlepoint ×1.0007.

The sum cannot fall below -2.24 standard deviations, and the exact tail runs down to zero there. The one-term correction crosses zero at about -1.97 and reports -6.20e-3 at -2.19 — a negative probability.

One to six standard deviations out on an exponential sum. The saddlepoint's relative error at a single draw runs from 1.5e-3 to 1.6e-2, and at thirty from 2.2e-5 to 1.3e-4. The normal's at 30 draws runs from 7.6e-3 to 1.5e+3.

Both with a continuity correction. The saddlepoint's ratio stays within 0.07% of exact across 5 thresholds; the normal's falls to ×0.051 at 5 standard deviations.

The lower tail turns negative 1.97 standard deviations below the mean at 5 draws, 3.13 at a hundred and 9.83 at 100,000. The threshold recedes like the sixth root of the sample size, following the dashed prediction, so there is no sample size at which the region is gone.

The largest threshold before each approximation's ratio to the exact tail first leaves 1 ± 0.1. At a hundred draws: the normal 1.66, one Edgeworth term 3.09, two 4.11, and the saddlepoint past 12, the edge of the range searched. At a hundred thousand draws the normal has reached only 4.66.

On exponential samples the sample skewness has a median of 0.93 at ten observations and 1.85 at three hundred, against a true value of 2. The correction built on it gives a median of ×0.834 of the exact tail at ten, where the true-skewness correction gives ×1.081 and the normal ×0.617.

From 10 observations, the band for one future value has factor 2.371, for ten 3.716, for a hundred 4.942 and for a thousand 6.008. The 95%-content tolerance factor is 3.382 and is passed by m = 10; the Bonferroni stretch of the prediction factor reaches 7.567 at a thousand.

The interval holds one future value 95% of the time. It holds all of the next ten 67.9% of the time, against the 59.9% that 0.95 to the tenth power gives, because the future values share the sample's mean and spread and succeed or fail together.

Each cell should read 2.5 / 2.5. Two identical exponentials at 20 and 20 read 2.04 / 2.22; the worst cell, a wide exponential against a normal at 8 and 32, reads 9.79 / 0.47.

Five pairings of populations at five splits each. Whatever makes the difference skewed — unequal sizes, unequal spreads, opposite skews — the cells fall on one rising curve, with a rank correlation of 0.997.

The gap between the low and high rejection rates, as forty units are split between the groups. The skewness of the difference is zero at 29.6 units in the wider group, where the split is 2.83 to one; the split that minimises the standard error puts 26.7 there.

Both bars should read 2.5%. The t interval misses below the mean on 6.38% of samples. Widened until its total is exactly 5%, it misses below on 4.69% and above on 0.30%. Hall's transformation misses below on 3.31% and above on 1.96%.

The nominal rate is 2.5%. At fifteen observations the t interval's upper limit is exceeded on 8.31% of samples, the widened interval's on 4.93%, the shifted interval's on 7.99% and Hall's on 4.49%.

The t interval's upper limit sits 0.526 standard deviations above the mean on average and is exceeded on 8.31% of samples. A fixed one-sided multiplier tuned to 2.5% needs 0.841. Hall's transformation spends 1.044 and is exceeded on 4.49%.

Where it is used

26 essays draw this figure, each at the numbers its own argument is about, so the same picture answers 26 different questions.

All 80 figures

FieldsThreadsSeriesConceptsAll essaysSearch