Series

Rate — the series

4 essays on one idea, from the one that introduces it to the one that assumes the rest.
  1. Where the normal approximation converges, and where it does not. Relative error against the exact binomial. At n = 1280 the error at the median is 0.96% and three sigma out it is 25.7% — a factor of 27. The tail is where the approximation is used.

    The tail converges last

    The central limit theorem is usually shown as a shape arriving. What the demonstration leaves out is the rate — and the rate is wildly different in the middle and in the tail, which is where every approximation in the subject is actually read.

    part 2 · normal
  2. The normal approximation's error on a sum of 100 exponential draws, under its Berry–Esseen bound. The distance between the exact distribution function and the normal one peaks at 0.0133, at z = -0.01. The Berry–Esseen bound is 0.1146, 8.62 times the real worst error, and larger than the whole 2.5% tail a two-sided test reads.

    A bound written for a coin

    The Berry–Esseen theorem guarantees how far a standardised sum can be from the normal, and the guarantee is true. On an exponential source it is 8.62 times the real worst error at every sample size, the worst error sits at the centre rather than in a tail, and at a hundred draws the bound is larger than the 2.5% tail it would be asked to vouch for.

    part 3 · expansion
  3. The upper tail of 10 exponential draws: normal, one Edgeworth term, two Edgeworth terms, each against the exact tail. Each curve is an approximation divided by the exact gamma tail, so 1 is exact. Six standard deviations out at n = 10: the normal gives ×0.0000670, one Edgeworth term ×0.00159, two ×0.0167 and the saddlepoint ×1.0007.

    A correction that goes below zero

    One Edgeworth term takes the normal approximation's error at two standard deviations from 38% to 8% on ten exponential draws, and stretches the range within 10% of the truth from 1.66 to 3.09 standard deviations at a hundred. It also turns negative in the short tail at every sample size — past 3.13 standard deviations at a hundred draws and 9.83 at a hundred thousand — because the region recedes only as the sixth root of n.

    part 4 · expansion
  4. The upper tail of 10 exponential draws: normal, two Edgeworth terms, saddlepoint, each against the exact tail. Each curve is an approximation divided by the exact gamma tail, so 1 is exact. Six standard deviations out at n = 10: the normal gives ×0.0000670, one Edgeworth term ×0.00159, two ×0.0167 and the saddlepoint ×1.0007.

    An approximation built at the threshold

    The saddlepoint approximation reads the tail of a sum of five exponential draws to within 0.19% six standard deviations out, where the normal is short by a factor of more than sixty thousand. It is within 2.2% out to ten standard deviations on a single draw, where there is nothing to average, and within 1.1% on a binomial whose expected count is one. It works because it is built where the tail is read rather than at the mean.

    part 5 · expansion

All series