A tail the sample never saw
Worth reading first: The tail converges last.
An approximation built at the threshold found the saddlepoint approximation reading the tail of an exponential sum to within 0.19% six standard deviations out, where the normal was short by a factor of sixty thousand. The accuracy had a condition attached, stated at the end of that essay as the practical limit: the method needs the cumulant generating function of the source, , and a reader with data rather than a model does not have it.
The reader does have an estimate. With a sample the obvious substitute is the sample’s own generating function,
and the tilt can be run on it exactly as before. This essay measures what that buys, on sources whose sums have exact tails, and finds that the flat error which made the saddlepoint remarkable does not survive the substitution. What survives is the bootstrap’s error, and the bootstrap’s error grows with the threshold.
What is being estimated
The target is the tail of a standardised sum — the chance that the sum of thirty draws lands of its own standard deviations above its own mean. That is the quantity a p-value for a mean needs, and the quantity a bootstrap of a studentised statistic estimates. Standardising matters: an unstandardised tail read from a sample is dominated by the sample’s error about where the mean is, which is the same for every method and would bury the differences between them. Each reading here uses the sample standardised to mean zero and spread one, so the only thing any of them can get wrong is the tail’s shape.
Four readings are compared, each made from one sample of thirty:
- the normal, , which estimates nothing and assumes no skew;
- a fixed exponential model, the standardised tail of a gamma sum of shape 30 — correct for an exponential source and for no other, and again estimating nothing;
- a gamma fitted by the sample’s skewness, with its shape set to — the right family for every source here, with its one parameter estimated;
- the empirical saddlepoint, which assumes no family at all and tilts the sample as it stands.
The exact tail is an incomplete gamma function for every source, so each reading’s error is exact for its sample and the only randomness is the sample.
The median reading falls with the threshold
For an exponential source and samples of thirty:
| standard deviations out | exact tail | normal | empirical saddlepoint | fitted gamma |
|---|---|---|---|---|
| 1 | 0.1575 | 1.01 | 0.986 | 1.004 |
| 2 | 0.0316 | 0.719 | 0.857 | 0.910 |
| 3 | 0.00416 | 0.324 | 0.630 | 0.735 |
| 4 | 0.000386 | 0.0820 | 0.374 | 0.525 |
| 6 | 0.00000146 | 0.000675 | 0.061 | 0.185 |
Each entry after the second column is the median ratio of the reading to the exact tail. At one standard deviation every method is right. By six the normal is short by a factor of fifteen hundred, which is the story of the tail that converges last, and the empirical saddlepoint — the same Lugannani–Rice formula that was within a fifth of a per cent when given the true — is short by a factor of sixteen.
The error is no longer flat. With the source known, the saddlepoint’s error barely moved between one and six standard deviations. With the source estimated it grows at every step out, from 1.4% at one to 94% at six, which is the signature of the method the empirical saddlepoint is known to approximate: the bootstrap. Tilting the empirical distribution reads the tail of the sum of thirty resamples from the sample, and it reads that accurately; what it cannot do is make the sample’s tail into the source’s.
Why it reads low rather than merely noisily
The spread in the hero figure is wide, and it is lopsided. At six standard deviations the tenth percentile of the empirical saddlepoint’s reading is 0.005 of the truth and the ninetieth is 0.641: nine samples in ten understate the tail, and only one in ten reads it within a factor of two. The error is a bias with noise around it, not noise around the truth.
It has two causes, and both are properties of a sample of thirty from a skewed source rather than of the saddlepoint.
The sample underestimates the skewness. Thirty exponential draws have a sample skewness whose median is 1.30 against a true 2. A correction that goes below zero met the same fact when it put the estimated skewness into an Edgeworth term: the skewness is carried by rare large values, a sample of thirty usually contains too few of them, and the estimate is biased towards symmetry. A skewness that is too small is a tail that is too thin, and every reading built from the sample inherits it.
The sample has nothing past its largest value. The empirical distribution puts no mass above its maximum, so the tail of a sum of resamples is made of combinations of values that actually occurred. Six standard deviations out on a sum of thirty is reachable only by combinations dominated by the sample’s few largest draws, repeated, and if the sample did not happen to include a draw from the source’s far tail, no combination of its values can stand in for one.
The fitted gamma suffers from the first cause and not the second. Its tail is a smooth function extrapolated beyond the data on purpose, so it is not confined to what occurred, and its median reading at six standard deviations is 0.185 — three times the empirical saddlepoint’s. Its spread is wider on the upside, though: its ninetieth percentile is 1.52, because a sample that happens to contain one very large draw overestimates the skewness, and a family extrapolates that error as faithfully as it extrapolates the truth.
A wrong model against no model
The comparison that the question left open was between the empirical tilt and a fitted model that extrapolates. The sources here allow a sharper version of it, because the exponential model can be applied to data that are not exponential.
On a gamma source of shape 2 the exponential model is simply the wrong family: it assumes a skewness of 2 where the truth is 1.41. It overstates the six-standard-deviation tail by a factor of 3.98, and it does so identically on every sample, since it estimates nothing. The empirical saddlepoint, which assumes nothing, understates the same tail at a median of 0.086 — a factor of twelve — and the fitted gamma at 0.258.
On a gamma source of shape one half, skewness 2.83, the exponential model errs the other way, reading 0.233 of the truth at six standard deviations; the empirical saddlepoint reads 0.046 and the fitted gamma 0.138. The pattern holds on all three sources: a fixed family with the skewness wrong by a third still reads the far tail more closely than a sample of thirty does, and in the direction that is safer — on the gamma(2) source it overstates, where the sample understates.
That is not an argument for guessing a family. It is a measurement of how little a sample of thirty knows about its source’s tail: less, six standard deviations out, than an informed wrong guess. The ordering reverses near the centre, where the sample knows plenty — at two standard deviations on the gamma(2) source the empirical saddlepoint is within 13% at the median and the wrong model is off by 8% — and the crossover is roughly where the threshold reaches values the sample has only a few draws beyond.
Where the tilt runs out
The second limit named when the question was first asked is not a matter of accuracy. The saddlepoint equation has a solution only while lies below the sample’s largest value, because every tilt of the sample is a reweighting of its values and no reweighting can produce a mean above the largest of them.
For the sum of thirty the edge never bites here: six standard deviations out asks for a mean of about 1.1 of the sample’s standard deviations above its centre, and every sample of thirty reaches further than that. For short sums it bites hard. A sum of two, six standard deviations out, is beyond reach on 95.3% of samples of thirty; a sum of five, on 32.4%; a sum of ten, on 0.7%. The edge is where the empirical saddlepoint stops returning a small number and starts returning nothing, and in a sense that is its most honest output: the data contain no evidence about a region they never entered.
A bootstrap with a finite number of resamples has the same edge in a softer form. It returns zero for every tail smaller than one over the number of resamples, and it returns a noisy small number for tails near that — which is why the empirical saddlepoint was proposed as the bootstrap’s replacement for tail readings in the first place. It removes the resampling noise. It cannot remove the fact that the resamples are all made of the same thirty numbers, which is the subject of where the bootstrap lies.
What more data buy
The sample’s two failures both shrink as the sample grows, but slowly, and the tail is again the last place to be repaired.
| observations | sample skewness, median | empirical saddlepoint, six sd | fitted gamma, six sd |
|---|---|---|---|
| 30 | 1.31 | 0.069 | 0.189 |
| 100 | 1.70 | 0.274 | 0.513 |
| 300 | 1.85 | 0.465 | 0.721 |
| 1,000 | 1.94 | 0.655 | 0.873 |
At a thousand observations the empirical saddlepoint still reads the six-standard-deviation tail a third low at the median, and the share of samples within a factor of two of the truth is 59%. At four standard deviations it is within 12% at the median and within a factor of two on 99% of samples. The sample size a tail reading needs grows with how far out the tail is — the same shape as the rate at which the normal itself converges, transferred from the approximation to the data.
The fitted gamma improves faster because it needs only the skewness, which a thousand observations estimate to within a few per cent, whereas the empirical saddlepoint needs the whole upper tail of the source, which a thousand observations sample thinly.
The third route: a tail model fitted to the tail
Between a family chosen for the whole distribution and no family at all there is a third option, and it is the one extreme-value practice uses. Fit a model only to the largest observations — the excesses over a high threshold, which the generalised Pareto distribution describes whatever the source, under conditions most sources meet — and extrapolate from that. The threshold is a dial and a level with no data in it measure that route for a single variable’s tail, and it sits between the two compared here in exactly the way the arithmetic predicts: it assumes less than a gamma, since it commits only to a tail shape, and more than the sample, since it extrapolates past the maximum on purpose.
For the tail of a sum the route is less direct. A sum of thirty draws is far in its tail when several of its terms are large at once, or one is very large, and a model of the single-variable tail says how often each happens; the sum’s tail follows by a convolution the model makes possible and the sample does not. On a source whose tail is exponential — every gamma here — the fitted tail would have a shape parameter near zero, and its uncertainty at thirty observations, of which perhaps five are in the tail, would be large. That is the honest limit of the third route at this sample size. It does not have the empirical saddlepoint’s hard edge, and it does not escape the fact that five observations say little about a shape.
What the three routes have in common is that each makes the source’s far tail from something: the sample’s largest few values, a family, or a model of excesses. The empirical saddlepoint is distinguished only by making it from the least, and its error in the table is the price of that modesty — a price that happens to be paid entirely in one direction, towards reporting far tails as rarer than they are.
What this means for a reading from data
The saddlepoint result was a statement about approximating a known distribution: given , one root-find reads any tail to a fraction of a per cent. The result here is a statement about estimating an unknown one: given thirty numbers, the tail six standard deviations out is not in the data, and no approximation can take it out of them.
So the choice the essay on the threshold posed — an approximation built at the threshold or one built at the centre — has a second axis once the source is unknown. On that axis the question is where the information about the tail is to come from. The empirical saddlepoint takes it from the sample and gets the sample’s answer, which is too thin. A fitted family takes it from the family and gets the family’s answer, which is right if the family is and wrong by a stable, visible factor if it is not. A Berry–Esseen bound takes it from a moment and gets a guarantee too loose to use. None of the three gets a far tail right from thirty observations, and a report of a far-tail probability from a sample of that size should say which of them it is leaning on, because the three disagree by more than an order of magnitude.
The practical rule the numbers support is a modest one. Inside about three standard deviations, the empirical saddlepoint is a good default — within 37% at the median at three on thirty exponential observations, without assuming anything, and far better than the normal. Beyond four, a tail read from thirty observations is a statement about the model used to extrapolate, and it should be presented as one, with the model named and its sensitivity shown, rather than as a fact about the data.
How far thirty draws see into a tail, measured against exact tails
From one sample of thirty exponential draws the empirical saddlepoint reads the tail of a standardised sum of thirty at a median of 0.986 of the truth one standard deviation out and 0.061 six out, and nine samples in ten understate the six-standard-deviation tail. The same formula with the true generating function was within 0.19%.
A gamma fitted by the sample’s skewness reads 0.185 at six standard deviations; a fixed exponential model, wrong for a gamma(2) source, overstates it there by a factor of 3.98 while the empirical saddlepoint understates it by a factor of twelve.
The median sample skewness of thirty exponential draws is 1.30 against a true 2, and the thin tail follows from it and from the empirical distribution’s having no mass past its maximum.
The tilt cannot reach six standard deviations for a sum of two on 95.3% of samples of thirty, and for a sum of five on 32.4%.
Every ratio divides a reading by an exact incomplete-gamma tail, so the only randomness is in the samples: two thousand at each setting in the tables, five hundred for the sample-size sweep, all seeded. The empirical saddlepoint is solved by bisection and Newton steps on the sample’s own generating function, evaluated with the sample’s largest value factored out of every exponential.
Not claimed: that the empirical saddlepoint is worse than the bootstrap it approximates — it is the same estimate without resampling noise. Not claimed either that a fitted family is safe; its good showing here is on sources inside the family, or near it, and a source with a heavier tail than any gamma would defeat it in a way these three cannot show.
Still open: a statistic that is not a sum
Every tail here is the tail of a sum with its spread known after standardisation. The statistic a reader actually computes from a sample divides by a spread estimated from the same sample, and on a skewed source the coupling between the two is what makes a t interval miss low four times as often as high. Saddlepoint approximations for such a ratio exist — they tilt the joint distribution of the sum and the sum of squares — and are harder to write down.
Whether the ratio’s saddlepoint keeps the flat error that the sum’s had, when the source is known, is measurable against exact or very large simulated tails and has not been measured. It would say whether the skew of a t statistic, which is where most practical tail errors live, can be read at the threshold as accurately as the skew of a sum — and, with the result above in hand, how much of that accuracy a sample of thirty would then give back.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A copula that halves a marginal — both name skewness, tail probability
- How many subjects — both name normal approximation, sample size
- Residuals that keep their own variance — both name bootstrap, skewness
- Sums of almost anything — both name sample size, skewness
- The chance a trial succeeds — both name normal approximation, sample size
- The correction for not knowing the spread — both name sample size, skewness
Named objects
A flat tag is an object no other essay names yet.
BootstrapCumulant generating functionExponential tiltingNormal approximationSaddlepoint approximationSample sizeSkewnessTail probability