A tilt for the ratio
Worth reading first: The tail converges last.
The correction a t test would use ended with every polynomial construction on the t statistic either turning back or flattening out somewhere between a tenth of a per cent and one in a million. The one construction that had done neither for a sum was the saddlepoint, which reads the tail at the threshold rather than expanding about the centre, and so has no curvature to spend as it moves outward. For a studentised statistic a saddlepoint exists in principle. It is harder to write down, and the essay that measured the polynomials named the question it would answer: with the source known, does its error stay flat out to one in a million on both sides?
On an exponential source it does — on the samples it can be built on. That qualification is the larger finding, and it has nothing to do with the tails. It comes from where the exponential sits in its own family.
What has to be tilted
A sum’s saddlepoint tilts the law of a single draw: multiply the density by , renormalise, and choose so the tilted mean sits at the threshold. The t statistic is not a sum. It is , a ratio of the mean to a spread estimated from the same draws, and the spread is a function of the mean of the squares. So is a function of two means at once — the mean of the draws and the mean of their squares — and its saddlepoint has to tilt their joint law, multiplying the density of a draw by and choosing both exponents to put the pair of means where the tail needs them.
For a unit exponential draw the joint cumulant generating function is the logarithm of , and that integral is finite only when . A positive makes the integrand grow like and nothing in the exponential’s tail can stop it. So every tilt that exists has , or sits exactly at , where it is just another exponential.
A tilt with multiplies the exponential density by a Gaussian factor, and the result is a normal law truncated at zero. Every such law is log-concave, and a log-concave law on the half-line has a coefficient of variation of at most one. The exponential, with its standard deviation exactly equal to its mean, is the boundary case. The tilts reach every pair of means whose implied coefficient of variation is below one, and none whose coefficient of variation is above it.
That is a strong statement about samples. A sample of thirty exponential draws whose own standard deviation, with divisor n, exceeds its own mean is a point no tilt can put its mean at. At such a point the saddlepoint density does not exist: the rate function is still finite there, but its supremum is attained on the edge of the domain, and the curvature the density formula needs is the curvature of a tilt that is not inside it.
How many samples that is
The exponential’s coefficient of variation is exactly one, so a sample’s falls on either side of it, and the share on the reachable side is the first thing to measure.
It is 74.5% at ten draws, 65.8% at thirty, 59.6% at a hundred, 55.6% at three hundred and 53.2% at a thousand. On small samples the reachable side is favoured, because a sample’s standard deviation is biased low and its coefficient of variation with it; as the sample grows the bias fades and the share falls towards a half. A third of the samples of thirty that any analyst will ever see lie where the saddlepoint for their t statistic cannot be built.
The property is particular to the source. A normal source has every coefficient of variation available to its tilts, and a source with a lighter tail than the exponential sits inside its tilt family rather than on its edge. The exponential is singled out because it is the one law that is exactly as far as a log-concave tilt can go.
Built where it exists
The saddlepoint density of the two means is , with the rate at the point and the covariance of a draw and its square under the tilt that reaches it. Under a truncated normal tilt both are closed forms: the rate from the normalising constant, and the covariance from the truncated normal’s first four moments. The tail of is then the density integrated over the region of the two means where is past the critical value — restricted to the region a tilt can reach, because outside it there is nothing to integrate.
The comparison has to be like with like, so the exact tail is split the same way. The essay on Hall’s correction found that ’s tails are exact given the sample’s shape: the draws are a gamma total times a point on the simplex, independent of each other, the coefficient of variation is a function of the shape alone, and past a critical value is a gamma tail given the shape. So the exact tail separates cleanly into the part carried by shapes a tilt can reach and the part carried by the others, and the saddlepoint is measured against the first.
On thirty draws the relative error on the short side runs from +0.88% at 10% to +2.47% at one in a hundred thousand. On the long side it runs from +0.83% to +2.05% at one in a million. Nothing in the construction was fitted to these tails — it is the joint cumulant generating function of the source, a closed-form tilt, and an integral — and its error stays inside two and a half per cent at every level shown.
That is the first construction in this field whose error on a t statistic stays flat out to one in a million on both sides. The polynomials could not, because they spend their accuracy as they move from the centre. The saddlepoint recentres at every threshold, as it did for the sum, and the ratio does not change that.
An error that shrinks like a saddlepoint’s
A construction that matched to two per cent at one sample size could be a coincidence. The behaviour that identifies a saddlepoint is how its error moves with the sample.
At a level of one in a thousand the error on the short side is +9.75% at ten draws, +2.11% at thirty and +0.24% at a hundred. On the long side it is +7.22%, +1.51% and −0.16%. Each threefold increase in the sample cuts it by a factor of four or more, which is at least as fast as the relative error a saddlepoint density is known for, and faster than the of the polynomial corrections. Ten draws is the one size at which the error is large enough to matter, and even there it is a matter of about ten per cent rather than the factors of ten the other constructions reach.
The blind spot is the middle
So the saddlepoint is accurate where it exists. The question the first section raised is how much of each tail lies where it does not.
The natural expectation is that the unreachable samples are a tail problem: samples with a large spread relative to their mean sound like samples with an extreme . They are not.
On thirty draws the unreachable samples carry 20.5% of the short tail at 10%, and under a hundredth of a per cent of it by one in ten thousand. On the long side they carry 23.2% at 10% and 0.79% at one in a million. The share falls on both sides as the level falls. The samples a tilt cannot reach are the samples near the centre of the statistic.
The reason is in what makes extreme. A large positive needs a mean well above one with a small spread — a sample with a small coefficient of variation, deep inside the reachable side. A large negative needs a mean well below one with a spread that has not collapsed with it, and a sample of thirty exponential draws whose mean has fallen far is a sample that has missed its large values, which is a sample whose spread has fallen too. Extreme values of the ratio are made by samples that look less exponential than the source, not more. The samples on the far side of the edge — the ones that look more dispersed than any log-concave law — mostly produce ordinary values of .
The hundred-draw curve shows the cost of that. At a hundred draws the unreachable samples carry 33.0% of the long tail at 10% and still 3.62% at one in a million. A larger sample has more samples on the far side of the edge, because its coefficient of variation settles towards the source’s one, and more of them reach into the long tail.
Two tails that are not mirror images
The two sides behave differently enough that it is worth seeing the numbers they are read at. On thirty exponential draws the exact critical value for a one-sided level of one in a thousand is 2.571 on the short side and −5.483 on the long side, against ±3.09 for a normal statistic. At one in a million the long side’s critical value is −11.835. A t statistic on this source reaches eleven standard units below zero about as often as a normal statistic reaches nearly five.
That asymmetry is the coupling between the mean and the spread that makes a t interval on skewed data miss low several times as often as high. A sample missing its large values has a low mean and a low spread together, and dividing one by the other inflates a modest shortfall into an extreme negative ratio. The joint tilt handles that coupling for free, because it tilts the mean and the mean square together rather than correcting one for the other afterwards. That is why its error on the long side is no larger than on the short — +1.51% against +2.11% at one in a thousand — when the corrections measured earlier in this field, each built from the mean’s distribution with a term added for the spread, were all worst on the long side. The coupling is not a correction to the joint tilt; it is part of what the tilt is.
It also explains where the polynomial corrections spend their accuracy. A Cornish–Fisher or Edgeworth correction is a polynomial in the threshold, centred at zero, and on the long side the threshold it has to reach is twice as far out as on the short side for the same level. A correction that goes below zero is what a polynomial does when it is asked to be right a long way from where it was expanded, and the long side of a t statistic asks for that sooner than any other tail in this collection.
Read as an approximation to the whole tail
A practitioner does not have the exact reachable tail to divide by. Read as an approximation to the whole tail, the saddlepoint simply omits the unreachable part, so its error is the unreachable share with a sign and the small saddlepoint error on top.
On the short side the joint saddlepoint reads ×0.802 of the whole tail at 10%, ×0.974 at 1% and ×1.019 at one in a thousand, and then holds: ×1.025 at one in a hundred thousand. The normal reads ×1.33 at 10% and ×17.0 at one in a hundred thousand, Student’s t ×1.37 and ×60.6. Hall’s one-term expansion reads ×0.871 at 10%, falls to ×0.476 at three in a hundred, and is below zero from 1% outwards — the same turn the essay on the t test’s correction found in its critical value, appearing here in the tail it was inverted from.
On the long side the joint saddlepoint reads ×0.775 of the whole tail at 10%, ×0.890 at 1%, ×0.953 at one in a thousand and ×1.012 at one in a million. The normal reads ×0.462 at 10%, Student’s t ×0.515 and Hall’s expansion ×0.855, and by one in a thousand all three are under a hundredth of the truth.
So the saddlepoint’s whole-tail error is largest exactly where the other constructions are least wrong, and smallest where they fail completely. At 10% on the long side it is short by a fifth because a fifth of that tail is unreachable; Hall’s expansion is short by less there. By one in a thousand the order has reversed by two orders of magnitude.
What this settles, and what it does not
The question this essay was asked has a clean answer. With the source known, the saddlepoint for a t statistic reads both tails on thirty exponential draws within two and a half per cent of the reachable truth from 10% to one in a million, and its error falls with the sample at least as fast as . The long side of a t statistic on skewed data is not beyond the reach of a formula.
What it does not settle is the part the essay on the sample’s own tilt measured for a sum, where the source is not known and the cumulant generating function has to be read from the sample. There the failure was a sample that had missed its large values tilting towards a tail it had never seen. For the joint tilt that failure takes a sharper form, because the empirical joint cumulant generating function of thirty draws and their squares is finite for every , positive or negative — a finite sample has no tail to make it diverge — so the empirical tilt can reach points the true one cannot — including every point a sample whose coefficient of variation is past one would need.
And it leaves the centre to something else. At 10% a fifth of either tail is in samples the saddlepoint cannot see, and there Hall’s one-term expansion, short by 13% on the short side and 15% on the long, is the better reading. By three in a hundred the saddlepoint is ahead on both sides, and from 1% outwards nothing else is close. The two are not rivals: one is built at the centre and one at the threshold, and on this source they hand over to each other at a few per cent.
Still open: the empirical joint tilt
The next measurement is the empirical version. Replace the exponential’s joint cumulant generating function with the sample’s own — the logarithm of the average of over the thirty draws — and build the same saddlepoint from it, as an analyst without the source would have to. Two things are measurable with the exact-given-shape tails used here. The first is whether, on samples whose coefficient of variation is below one, the empirical tilt inherits the true tilt’s two-per-cent accuracy or the sum’s failure on the long side. The second is what it does on the third of samples past the edge, where the empirical tilt exists and the true one does not: whether it extrapolates sensibly into a region the source forbids, or reports a tail for a law that could not have produced the sample. If the second, the coefficient of variation is a diagnostic an analyst can read before trusting the long side of a t test, and a level with no data in it is the kind of model the long side would then need.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The same correction, inverted — both name cornish–fisher expansion, edgeworth expansion, skewness, tail probability
- A bound written for a coin — both name skewness, tail probability
- A copula that halves a marginal — both name skewness, tail probability
- A failure charged by its size — both name skewness, student's t
- A level set before the source is known — both name skewness, student's t
- A limit from the family it came from — both name skewness, student's t
Named objects
A flat tag is an object no other essay names yet.
Coefficient of variationCornish–Fisher expansionCumulant generating functionEdgeworth expansionExponential tiltingSaddlepoint approximationSkewnessStudent's tTail probability