Concept

Tail probability — where it appears

The chance a quantity exceeds a stated value, which is what a p-value and a coverage statement are both made of. It is the part of a distribution that approximations are worst at, and the part every decision is taken in.

Named by 12 essays across 9 fields — each of them below, with the objects they name alongside it.

The normal approximation's error on a sum of 100 exponential draws, under its Berry–Esseen bound. The distance between the exact distribution function and the normal one peaks at 0.0133, at z = -0.01. The Berry–Esseen bound is 0.1146, 8.62 times the real worst error, and larger than the whole 2.5% tail a two-sided test reads.

A bound written for a coin

The Berry–Esseen theorem guarantees how far a standardised sum can be from the normal, and the guarantee is true. On an exponential source it is 8.62 times the real worst error at every sample size, the worst error sits at the centre rather than in a tail, and at a hundred draws the bound is larger than the 2.5% tail it would be asked to vouch for.

expansion · Rate
A threshold in the tail is a threshold nothing balances. The share of a coin's imbalance in an indicator 1{x > c} that survives a rule which balances the covariate itself. The smooth curve is 1 − ρ² with ρ = φ(c)/√(p(1−p)), a closed form with no trial in it; the points are counted over 500 trials of 200 units at each threshold. At the median the two agree that about a third survives — the removed share is exactly 2/π — and by two standard deviations 86.9% survives. The closed form is exact in the limit and optimistic by a few points at this many units, because the rule balances the sample's mean rather than the population's.

A threshold in the tail

How much of a threshold's imbalance a balanced covariate removes is a correlation, and the correlation is a closed form. At the median it is exactly 2/π — the same 2/π a median split throws away — and two standard deviations out it is an eighth.

shape · Blocking
Where the normal approximation converges, and where it does not. Relative error against the exact binomial. At n = 1280 the error at the median is 0.96% and three sigma out it is 25.7% — a factor of 27. The tail is where the approximation is used.

The tail converges last

The central limit theorem is usually shown as a shape arriving. What the demonstration leaves out is the rate — and the rate is wildly different in the middle and in the tail, which is where every approximation in the subject is actually read.

normal · Rate
One tail arrives; the other is still on its way at a million. The Kolmogorov distance between the exact law of a normalised maximum and its Gumbel limit, at six block sizes, for two parents that both have that same limit. Both are closed form: the exact law of a maximum is F(x)^n and no simulation is involved. The exponential parent's distance falls from 0.0280 to 2.707e-7 — a factor of a hundred thousand, which is exactly one over n. The normal parent's falls from 0.0522 only to 0.0091, a factor of 5.74, because its rate is one over log n. At a million readings a block the two differ by a factor of 33556.3.

The maximum converges slowly

The rate at which a normalised maximum reaches its limit law is computable rather than simulable, because the exact law of a maximum is always available. For a normal parent the distance falls like one over the logarithm of the block and is still 0.0091 at a million readings; for an exponential parent, with the same limit, it is 2.707×10⁻⁷.

extreme · Extremes
The upper tail of 10 exponential draws: normal, one Edgeworth term, two Edgeworth terms, each against the exact tail. Each curve is an approximation divided by the exact gamma tail, so 1 is exact. Six standard deviations out at n = 10: the normal gives ×0.0000670, one Edgeworth term ×0.00159, two ×0.0167 and the saddlepoint ×1.0007.

A correction that goes below zero

One Edgeworth term takes the normal approximation's error at two standard deviations from 38% to 8% on ten exponential draws, and stretches the range within 10% of the truth from 1.66 to 3.09 standard deviations at a hundred. It also turns negative in the short tail at every sample size — past 3.13 standard deviations at a hundred draws and 9.83 at a hundred thousand — because the region recedes only as the sixth root of n.

expansion · Rate
How the truth, the raw means, the posterior means and the constrained estimates spread, standard error 1. Beyond two population widths above the centre lie 2.28% of the true values, 7.86% of the raw means, 0.234% of the posterior means, and 2.28% of the constrained estimates.

Estimates that are too alike

Posterior means give each group its least-error estimate, and as a set they are too alike: with each group's standard error equal to the population's spread, they spread 0.707 as widely as the truth. Beyond two population widths lie 2.28% of the true effects, 7.86% of the groups' own means, and 0.234% of the posterior means — a tenth of the truth. Rescaling the estimates to the right spread counts the tail exactly and costs 17% more squared error; summing each group's posterior chance of being beyond the line counts it without changing any estimate.

borrowed · Shrinkage
Three copulas that break nothing, and a factor of two between them. The three radially symmetric copulas, at a matched Spearman correlation of 0.40, against the covariate's marginal. All three leave exactly nothing with a symmetric covariate — that is the guarantee, and it holds to twenty decimal places. What they do to a skewed covariate is not the same at all: at a skewness of 2.26 a Frank copula leaves 12.118% where a Gaussian leaves 21.539% and a t on four degrees of freedom leaves 23.640%. A factor of 2.0 between two copulas that are both symmetric, both matched on rank correlation, and both harmless on their own. So the copula matters to the marginal's leak without breaking any symmetry of its own, which is a milder version of the same finding and applies to every trial rather than to the asymmetric ones.

A copula that halves a marginal

Three copulas break nothing on their own and put a factor of two between the same skewed covariate's leaks — 12.118% under a Frank against 23.640% under a t, at the same rank correlation.

compound · Adjustment
What each criterion selects, at 50 observations. 700 series from an AR(2) with coefficients 0.6 and -0.3, every order from 0 to 8 fitted to the same 42 responses so the log-likelihoods are comparable. AIC finds the true order 55.1% of the time and lands above it 25.7%; BIC finds it 54.1% and lands above it 4.0%. The closed form for one extra lag is P(χ²₁ > 2) = 15.73% for AIC, which does not depend on n at all, and P(χ²₁ > ln n) = 4.79% for BIC at this size, which falls to zero. Under the true order is the other failure and it is BIC's: 41.9% against 19.1%.

Choosing the order

One criterion is consistent and one is not, which is the whole of what gets said about them. At two hundred observations the consistent one is right 95% of the time and the other 70%; at fifty they are both right 54% of the time and wrong in opposite directions, and consistency has not started to mean anything yet.

forecast · Order-selection
What each instrument costs to read. The number of draws each instrument needs to separate a rectangular block from a trapezoidal one at two standard errors, at a block length of 20 and 120 rows — measured from each instrument's own spread on the same draws. The implied variance needs 7.0 and the 95% point needs 20.2, a factor of 2.90 at this block length. There is a closed form beside it and it does not depend on either the scale or the size of the gap: the standard error of a p-quantile is √(p(1−p))/f(q) over √B where a standard deviation's is σ/√(2B), which at the 95% point of a nearly normal reference distribution is 3.30 times as many draws for the same statement. And the quantile route needs every one of those draws resampled, where the variance route needs none.

Measuring a variance rather than a quantile

A resample's implied long-run variance can be computed from the sample with no resampling in it at all. A critical value cannot, and the difference is a factor of three in the draws before any of the resampling is counted.

crossing · Bootstrap
The upper tail of 10 exponential draws: normal, two Edgeworth terms, saddlepoint, each against the exact tail. Each curve is an approximation divided by the exact gamma tail, so 1 is exact. Six standard deviations out at n = 10: the normal gives ×0.0000670, one Edgeworth term ×0.00159, two ×0.0167 and the saddlepoint ×1.0007.

An approximation built at the threshold

The saddlepoint approximation reads the tail of a sum of five exponential draws to within 0.19% six standard deviations out, where the normal is short by a factor of more than sixty thousand. It is within 2.2% out to ten standard deviations on a single draw, where there is nothing to average, and within 1.1% on a binomial whose expected count is one. It works because it is built where the tail is read rather than at the mean.

expansion · Rate
The normal density at sigma = 1.00. The bands hold 68.27%, 95.45%, 99.73% of the mass. Those figures are integrals of the curve drawn, not the memorised 68-95-99.7.

The shape, and where its mass is

68, 95, 99.7 is recited more often than any other set of numbers in the subject. They are integrals of a specific curve, they are worth computing rather than remembering, and the third one is the one people misuse.

normal · Bands
Estimating P(Z > 5) = 2.8665×10⁻⁷ with plain draws and with four proposals. One seed each. Plain simulation draws nothing past 5 in 100,000 and estimates zero throughout. At 100,000 draws the proposal N(5, 1) reads 1.009 of the truth, N(9, 1) 1.041, N(4.5, 0.25²) 1.039 and N(5, 0.3²) 1.008. Values above 2.2 are drawn at the top edge.

The draws aimed at the tail

The chance a standard normal exceeds 5 is 2.8665×10⁻⁷, and a plain simulation needs 349 million draws to estimate it to within ten per cent. Draws aimed at the tail and weighted back need 565. Aimed slightly too narrowly, the same method has an infinite variance, an interval that covers 86.0% and gets worse with more draws, and an effective sample size that reads healthier than a proposal that works.

method · Seeds

Named alongside it

The objects these essays reach for when they reach for this one.

Closed formCentral limit theoremConvergence rateNormal approximationSkewnessMonte CarloNormal distributionApproximation errorBerry–EsseenCorrelationCovariate balanceDegrees of freedom

All concepts