Past the first term of the normal approximation

A bound written for a coin

The Berry–Esseen theorem guarantees how far a standardised sum can be from the normal, and the guarantee is true. On an exponential source it is 8.62 times the real worst error at every sample size, the worst error sits at the centre rather than in a tail, and at a hundred draws the bound is larger than the 2.5% tail it would be asked to vouch for.

Worth reading first: The tail converges last.

The tail converges last measures the normal approximation against an exact answer and quotes the theorem usually offered in its defence: the Berry–Esseen bound, which says the absolute error falls like one over the square root of the sample size. The essay names the bound and does not measure it.

It is worth measuring, because it is the only guarantee in the subject. Every other statement about how good a normal approximation is at a given sample size is either a rule of thumb or a simulation of one case. Berry–Esseen is a theorem, it holds for every source with a finite third moment, and it gives a number. The question is what the number is good for.

The normal approximation's error on a sum of 100 exponential draws, under its Berry–Esseen boundThe distance between the exact distribution function and the normal one peaks at 0.0133, at z = -0.01. The Berry–Esseen bound is 0.1146, 8.62 times the real worst error, and larger than the whole 2.5% tail a two-sided test reads.0.0000.0320.0640.096-4-2024standard deviations from the mean|exact − normal|the bound, 0.1146a 2.5% tailworst 0.0133exact incomplete gamma against the normalbound ÷ worst error = 8.62
Fig. 1 The distance between the exact distribution of a standardised sum of a hundred exponential draws and the normal, at every threshold. The solid rule is the Berry–Esseen bound on that distance; the dashed one is a 2.5% tail. The slider is the number of draws summed.

What the theorem says

For independent draws with mean μ\mu, standard deviation σ\sigma and third absolute moment ρ=EXμ3\rho = E|X - \mu|^3, the distribution function FnF_n of the standardised sum satisfies

supzFn(z)Φ(z)    Cρσ3n\sup_z \bigl|F_n(z) - \Phi(z)\bigr| \;\le\; \frac{C\,\rho}{\sigma^3 \sqrt{n}}

at every sample size, not merely in the limit. The constant has been whittled down for eighty years; the best value known for identically distributed draws is 0.4748.

Three things about the statement matter for what follows. It is about the distribution function, so it bounds the absolute difference between two probabilities. It is a supremum, so it bounds the worst threshold and says nothing about where that threshold is. And the source enters through one number, ρ/σ3\rho/\sigma^3, so two sources with the same standardised third moment get the same bound however differently their sums actually behave.

The source where the exact answer is free

A sum of nn exponential draws is a gamma variable with shape nn, so its distribution function is an incomplete gamma function and can be evaluated to full precision at any threshold. The exponential is also as skewed as an ordinary measured quantity gets — skewness 2 — which is the case the bound is invoked for. Its standardised third absolute moment is 12/e212/e - 2, 2.4146.

At a hundred draws the bound is therefore 0.4748×2.4146/100.4748 \times 2.4146 / 10, which is 0.1146.

The exact distance, found by evaluating both functions across the whole range and refining around the largest gap, is 0.0133. The bound is 8.62 times the error it bounds.

That ratio does not move with the sample size. At thirty draws the bound is 0.2093 and the worst error 0.0243; at a thousand, 0.0363 and 0.0042. Both fall as one over the square root of nn, which is the theorem’s claim and is correct, and they fall in step, so the slack is a constant factor that no amount of data removes.

Where the error actually is

The hero figure’s curve has one feature and it is in the middle. The largest distance between the exact distribution and the normal on an exponential sum sits at z=0.008z = -0.008 at a hundred draws — the centre of the distribution, where nobody reads a probability.

The reason is the first correction term of the Edgeworth expansion, measured properly in a correction that goes below zero. A skewed sum’s distribution function differs from the normal by γ(z21)φ(z)/(6n)\gamma\,(z^2 - 1)\,\varphi(z) / (6\sqrt{n}), and the factor (z21)φ(z)(z^2 - 1)\,\varphi(z) is largest in size at z=0z = 0, where it is φ(0)\varphi(0). For the exponential that predicts a worst error of γ/(62πn)\gamma / (6\sqrt{2\pi n})0.0133 at a hundred draws, to the four digits measured.

Out at a two-sided 5% threshold the absolute error is smaller. At 1.96 standard deviations above the mean the exact tail is 3.02% against the normal’s 2.50%, a difference of 0.52 points; below the mean it is 1.92% against 2.50%. Those are not small in relative terms — the upper tail is a fifth heavier than the normal says — but they are well inside what the bound allows, and the bound would allow them to be twenty times larger.

Six sources, one constant

If the bound is loose by a factor of eight on a source as skewed as the exponential, the constant must be set by something else. It is.

How far the Berry–Esseen bound sits above the real error, six sources, n = 100. The bound divided by the largest distance between the exact distribution function and the normal one. A fair coin reads 1.19, so the bound is nearly attained; an exponential source reads 8.62 and a gamma of shape 16 reads 23.61.
Fig. 2 The bound divided by the real worst error, for six sources at a hundred draws. The two coins sit near the line at one; every smooth source sits far to the right of it.
source ρ/σ3\rho/\sigma^3 bound worst error bound ÷ error
a fair coin 1.0000 0.0475 0.0398 1.19
a 10% coin 2.7333 0.1298 0.0832 1.56
gamma, shape 16 1.6534 0.0785 0.0033 23.61
gamma, shape 4 1.8206 0.0864 0.0066 13.00
exponential 2.4146 0.1146 0.0133 8.62
chi-square on 1 df 3.0729 0.1459 0.0188 7.76

A fair coin nearly attains the bound; no smooth source comes within a factor of seven of it. The coin’s error is not a matter of shape at all. A sum of coin flips takes whole-number values, its distribution function is a staircase, and a smooth curve cannot pass closer than half a step to a staircase at the step. At a hundred flips the middle step is 0.0796 high, half of it is 0.0398, and that half-step is the worst error.

So the constant in the theorem is a statement about discreteness. It has to cover the lattice, because the theorem is written for every source with a third moment and a lattice is one of them. A continuous source never has to pay for a jump, and its real error is governed by its skewness through a much smaller coefficient.

The ordering inside the smooth rows makes the same point from the other side. A gamma of shape 16 has skewness 0.5 and is nearly symmetric; its real worst error is tiny, but its ρ/σ3\rho/\sigma^3 is 1.6534, well over half the exponential’s, because the third absolute moment is large for any spread-out distribution whether or not it is skewed. The bound reads the wrong property. It charges a symmetric source for spread and a skewed one only a little more, while the error itself is driven by the skew.

This is also why a proportion’s interval oscillates as the sample grows while a mean’s does not: the half-step is the same object. A binomial is a lattice, the lattice is what the constant was written for, and the error that the Berry–Esseen constant has to allow for is the one that makes a sample of twenty cover twelve points worse than a sample of nineteen.

What a bound on the distribution function can vouch for

The practical use anybody has for a normal approximation is a tail probability — a p-value, a control limit, a coverage. The bound is on the distribution function, so it bounds a tail probability directly: a normal tail of 2.5% is within 0.1146 of the true tail at a hundred exponential draws.

That is the whole of what the theorem certifies, and read as a certificate it says that the true tail lies somewhere between 0 and 13.96%. The interval contains every p-value anyone has ever reported as significant and most of the ones reported as not.

Where the Berry–Esseen bound falls below a tail probability, exponential source. The bound on the distance from the normal falls as one over the square root of the sample size. It drops below a 2.5% tail at 2,103 draws, below 0.5% at 52,573 and below 0.05% at 5,257,207. Before each of those points it cannot vouch for that tail at all.
Fig. 3 The bound as a function of the number of draws, on logarithmic axes, against three tail probabilities. A tail is vouched for only to the right of the point where the bound crosses it.

A bound can only certify a tail smaller than itself, so the useful question is the sample size at which it drops below the tail being read. Because the bound falls as one over the root, a tail five times smaller needs twenty-five times the data:

tail the bound is below it from
2.5% 2,103 draws
0.5% 52,573 draws
0.05% 5,257,207 draws

At 2,103 draws the bound equals the tail, so the certificate reads “between 0 and 5%” — still not a statement anyone could act on. To pin a 2.5% tail to within a fifth of itself the bound would have to be 0.005, which is the second row: fifty-two thousand exponential draws. By then the real error at that threshold is 0.00024, about a twentieth of what the bound still allows.

The shape of the failure is not that the bound is wrong. It is that a uniform guarantee on an absolute error is the wrong kind of object for a tail. A tail probability is small, so the error that matters is relative, and a bound that is the same size at every threshold is enormous relative to any threshold far enough out to be interesting. The essay before this one made the distinction between absolute and relative error in words; this is what it costs in draws.

A smaller source of skew, and a larger one

The ratio of 8.62 is a property of the exponential. A different smooth source has a different ratio, and it is worth seeing which way the change goes.

The normal approximation's error on a sum of 30 chi-square draws, under its Berry–Esseen bound. The distance between the exact distribution function and the normal one peaks at 0.0344, at z = -0.02. The Berry–Esseen bound is 0.2664, 7.75 times the real worst error, and larger than the whole 2.5% tail a two-sided test reads.
Fig. 4 A sum of thirty chi-square draws on one degree of freedom — the squared normal, skewness 2.83. The worst error is larger and still at the centre, and the bound is larger again.

A chi-square on one degree of freedom is more skewed than the exponential and its sums converge more slowly: at thirty draws the worst error is 0.0344, against 0.0243 for the exponential at the same size. The bound grows too, to 0.2664, and the ratio falls only to 7.75. At the other end a gamma of shape 16 is barely skewed, its sums are nearly normal already, and the ratio is 23.61.

The more normal a source already is, the more the bound overstates the error. That is backwards from what a user of the bound would want. The cases where a normal approximation is safe are the cases the bound is least able to say so, and the cases where it is least safe are the cases where the bound is merely very loose rather than absurdly loose.

The bound’s real job

None of this is a complaint about the theorem, which does exactly what it says and is one of the more remarkable results in the subject. Its job is to establish a rate: that the error falls as one over the root of the sample size and no slower, for every source with three moments, with a constant that does not depend on the source beyond one ratio. That is a statement about how the limit theorem behaves, and it is true.

What it is not is a way to decide whether a particular normal approximation can be trusted at a particular sample size. For that it is too conservative by a factor that ranges from about 1.2 on a coin to more than twenty on a nearly symmetric smooth source, and it measures the error at the wrong place — the worst threshold, which for a smooth source is the centre.

The same caution applies to the rule of thumb it is usually quoted beside. “Thirty observations is enough” is a claim about the centre of the distribution, where the sum’s shape arrives quickly. At thirty exponential draws the worst error anywhere is 0.0243 and the upper 2.5% tail is really 3.40%, while three and a quarter standard deviations out the normal gives less than a quarter of the true tail. A single accuracy figure for an approximation means nothing without the threshold it is read at, and a bound that is the same at every threshold cannot supply one.

Where the half-step goes

A lattice source earns its tight bound, and the ordinary repair for a lattice removes almost all of what it earned.

The normal approximation's error in the tail, n = 100, p = 0.05. Approximate tail probability divided by the exact one. At one sigma the ratio is 0.98; at 4 sigma it is 0.104, so a rare event is understated by a factor of 10.
Fig. 5 The normal approximation to a binomial tail at a hundred trials and a proportion of 0.05, divided by the exact tail, as the threshold moves out. A continuity correction is applied throughout; the ratio still falls away from one.

The worst error on a fair coin is half a step, and half a step is what a continuity correction adds back: reading the normal at k+12k + \tfrac{1}{2} rather than at kk puts the smooth curve through the middle of each step instead of past its corner. Measured on the same two coins at a hundred flips:

source worst error, plain worst error, corrected removed
a fair coin 0.0398 0.00027 ×146
a 10% coin 0.0832 0.01747 ×4.8

On the fair coin the correction takes the error down by a factor of 146, to a quarter of a hundredth of a point. The staircase was the whole of it. The fair coin is symmetric, so once the lattice is dealt with there is nothing left for the normal to get wrong at the first order.

The 10% coin keeps a fifth of its error, and what it keeps is the skew. A proportion away from one half is a skewed source that also happens to be a lattice, and the correction repairs the second property and leaves the first. The figure above is that remainder read in the tail: continuity-corrected, at a proportion of 0.05, four standard deviations out the normal gives about a tenth of the true tail. The shape error — the one every skewed sum carries — is what decides the p-value, and it is exactly the error the Berry–Esseen constant was not sized for.

So the bound is tight on the error that a half-unit shift removes and loose on the error that nothing so simple removes. That is the precise sense in which it was written for a coin.

The rule the bound could have been

For a smooth source there is a rule that does what the bound does not, and it is already in the arithmetic above. The worst error is predicted by the skewness alone:

supzFn(z)Φ(z)    γ62πn    γ15n\sup_z \bigl|F_n(z) - \Phi(z)\bigr| \;\approx\; \frac{\gamma}{6\sqrt{2\pi n}} \;\approx\; \frac{\gamma}{15\sqrt{n}}

Checked against the exact search on every smooth source measured:

source skewness n exact worst error γ/(62πn)\gamma / (6\sqrt{2\pi n})
gamma, shape 16 0.50 100 0.00332 0.00332
gamma, shape 4 1.00 100 0.00665 0.00665
exponential 2.00 30 0.02429 0.02428
chi-square on 1 df 2.83 30 0.03437 0.03434

Agreement to four digits at a hundred draws and to three at thirty, with the small gap on the most skewed source at the smallest size being the next term of the expansion. As a constant this is 0.0665 in place of 0.4748, and it charges the source for the right property: skewness rather than third absolute moment, so a symmetric smooth source is charged nothing at this order, which is correct.

It is not a theorem in the way Berry–Esseen is. It is the leading term of an asymptotic series, and a leading term can be wrong at small samples by an amount nothing in the formula states — which is why the bound, which never fails, is the one in the textbooks. But it is the right size, and it makes the point this essay is about more sharply than the bound can: a smooth skewed source’s worst error is a property of its skewness and sits at the centre, and neither fact tells a reader anything about a tail.

What is claimed here, and what is not

The bound holds for every source at every sample size measured — six sources at thirty, a hundred and a thousand draws, eighteen comparisons, the real worst error below the bound in all of them. This is a theorem, and checking it is a check on the arithmetic of the comparison rather than on the theorem; a single failure would have meant the exact distribution functions were computed wrongly.

Every smooth source leaves at least twice the slack of every lattice. Stated as a comparison between the two groups rather than as a threshold on either, because the ratios within each group depend on the source and the claim is about which kind of source the constant is written for.

On an exponential source the worst error sits at the centre. Within a tenth of a standard deviation of it at all three sizes, and closing on zero as the sample grows.

What does not survive is the bound used as a certificate for a tail. At a hundred exponential draws it is 0.1146 and the tail is 0.025, and a certificate larger than the probability it certifies certifies nothing. It is also not claimed that the ratio of 8.62 holds for any source but the exponential, or that the Berry–Esseen constant could be much smaller: the fair coin’s ratio of 1.19 at every size is the reason it cannot.

The measurement is exact rather than simulated. Both distribution functions are evaluated in closed form — an incomplete gamma for the smooth sources, a binomial sum for the coins — so every number in the tables is a computation that could be repeated with a spreadsheet, and nothing here carries Monte Carlo error.

Still open: a bound that knows where it is being read

The bound fails as a tail certificate for a reason that is structural rather than numerical: it is uniform in the threshold, and a tail needs a guarantee that shrinks with the tail.

Such guarantees exist. Non-uniform Berry–Esseen bounds carry a factor of 1/(1+z3)1/(1 + |z|^3), which makes the allowed error fall off as the threshold moves out, and moderate-deviation results bound the ratio of the exact tail to the normal one over a range of thresholds that grows with the sample size. Either is closer to what a p-value needs. Neither has been measured here, and the obvious measurement is the one this essay made for the uniform bound — the sample size at which the certificate for a 2.5%, a 0.5% and a 0.05% tail first says something a decision could use.

The other direction is not a bound at all. The worst error on a smooth source was predicted to four digits by the first term of an expansion, which suggests that the expansion rather than the bound is the right instrument for a smooth source. A correction that goes below zero measures that term — where it repairs the normal, and the short tail where it produces a probability that is not one.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Berry–EsseenCentral limit theoremConvergence rateDiscretenessNormal approximationSkewnessTail probabilityWorst case