Field

The distribution itself

Sums converge on it, which is the most-quoted theorem in the subject. What the quoting leaves out is the rate: the middle converges quickly and the tail does not, and the tail is where the approximation is actually read. The same theorem carried through a function stops being a normal limit where the function is flat, and carried through a ratio it forbids every interval that is always finite.
8 exponential draws, standardised, against the normal. The source is one-sided and skewed. At n = 8 the standardised sum has skew 0.695, and the theory says 2/sqrt(n) = 0.707 — so the convergence is visible AND its rate is predicted.

Sums of almost anything

The theorem says sums converge on one shape whatever they are sums of, which is remarkable and true. Watching it happen from a one-sided skewed source, with the rate of convergence predicted in advance, is more convincing than watching the shape appear.

Where the normal approximation converges, and where it does not. Relative error against the exact binomial. At n = 1280 the error at the median is 0.96% and three sigma out it is 25.7% — a factor of 27. The tail is where the approximation is used.

The tail converges last

The central limit theorem is usually shown as a shape arriving. What the demonstration leaves out is the rate — and the rate is wildly different in the middle and in the tail, which is where every approximation in the subject is actually read.

Twenty samples of 40, every one of them genuinely normal. Each panel is a quantile-quantile plot of 40 draws from a normal distribution. The worst point in the worst panel sits 0.87 standard deviations off the line. Anything a reader would reject here would be a false alarm.

What normal actually looks like

A single quantile plot of forty normal points wanders enough to look suspicious. Twenty of them, all genuinely normal, show what the noise looks like — and any single panel a reader would have rejected is in there.

The normal density at sigma = 1.00. The bands hold 68.27%, 95.45%, 99.73% of the mass. Those figures are integrals of the curve drawn, not the memorised 68-95-99.7.

The shape, and where its mass is

68, 95, 99.7 is recited more often than any other set of numbers in the subject. They are integrals of a specific curve, they are worth computing rather than remembering, and the third one is the one people misuse.

The squared estimate 1 standard errors from the flat point, exact and linearised. At δ = √n·μ/σ = 1 the exact law of the squared estimate has mean 2.00, variance 6.00 and skewness 2.177; the delta method's normal has mean 1.00, variance 4.00, no skewness, and 30.85% of its mass below zero, where a square cannot go. The Kolmogorov distance between them is 0.3085.

Where the derivative is zero

The delta method reads a standard error off a tangent line, and at a flat point the tangent says the spread is zero. The interval built on it for a squared mean covers 99.991% there and 85.978% one and a half standard errors away, with nearly every miss on the same side — and the law it should have used is a χ², not a normal.

The two standardised means, and the shape of Fieller's set for their ratio at a = 3, d = 1. Each dot is one pair (zx, zy) drawn around (3, 1). Outside the horizontal band |zy| > 1.96 Fieller's set is a bounded interval, with probability 17.01%; inside the band and outside the disc of radius 1.96 it is everything outside an interval, 75.03%; inside the disc it is the whole line, 7.96%. On 40,000 counted draws Fieller covers ρ = 3.00 95.21% of the time and the delta interval 82.48%.

A ratio whose interval has to be the whole line

The delta interval for a ratio of two means covers 95.61% when the denominator is eight standard errors from zero and 1.10% at a ten-thousandth of one, and ten times as wide it still covers only 3.48%. Linearising is not the fault. Gleser and Hwang proved that every interval that is always finite fails the same way, so an interval that keeps its promise has to be the whole line some of the time.

The law is the eigenvalues, and nothing else. The mean and the skewness of n(ĝ − g) at the stationary point, measured over 40,000 draws, against the closed forms ½ Σλ and 2√2 Σλ³ ⁄ (Σλ²)^(3⁄2). The worst disagreement anywhere is 0.028. In one variable the second-order law is a single χ² and its sign is the sign of g″; here it is a weighted sum with the Hessian's eigenvalues as weights, so a bowl and a valley differ in both moments and a saddle has both equal to zero.

A flat point with more than one direction

At a stationary point of a function of several means the second-order law is ½ Z′HZ, so the bias is half the Hessian's trace — 2.008 for a bowl, 5.028 for a valley, and −0.006 for a saddle, where the eigenvalues cancel. The saddle's coverage is the worst of the three.

All essays