Past the first term of the normal approximation

The correction a t test would use

On thirty exponential draws the one skewness term that held a sum's critical value within 10% out to five standard deviations nearly triples a t test's 1% size when the source's skewness is known, and its upper critical value turns back — at a level of 0.89% on ten draws, where a test asked for one in a thousand rejects 121 times too often. Hall's monotone cubic removes the turn and gains a floor. Built from the sample's own skewness it holds the short side within 20% from 10% down to one in a hundred thousand, and leaves the long side, the one a safety limit reads, at twice its level: the samples that fail are the ones that hide their skewness.

Worth reading first: The tail converges last.

The same correction, inverted moved a sum’s critical value by one skewness term and found the test’s actual size held within 10% of nominal out to 4.17 standard deviations at thirty exponential draws, and to 5.10 at a hundred. It ended on the version of that correction a reader actually needs. Nobody tests a sum whose spread is known; a t test divides the mean’s departure by a spread estimated from the same draws, and the statistic that results is skewed the other way, more strongly, and with its spread and its centre tied together.

That essay asked whether the inverted correction reaches as far for a studentised statistic as it did for a sum, and whether its short tail turns back at a comparable level. The answers are no, and no — it turns back far sooner, and on the other side. And the construction that repairs the turn, Hall’s monotone cubic, has a failure of its own at the opposite end. What survives is the version built from the sample’s own skewness, and it survives on one side only, for a reason that says something about which samples mislead a t test.

What a corrected critical value delivers to a t statistic: actual size over nominal, upper tail, 30 exponential drawsAt a nominal 1% the upper-tail t test's actual size is ×0.153 with Student's quantile, ×2.94 with one Cornish–Fisher term and ×1.51 with Hall's cubic, both using the source's skewness, and ×0.878 with Hall's cubic using each sample's own. At 10⁻⁶ they are under a hundredth of nominal, ×1.04e+4, ×50.0 and ×1.59.×1000×100×10×1×0.1×0.0110⁻⁶10⁻⁵10⁻⁴10⁻³10⁻²0.1nominal level of the one-sided testactual size ÷ nominalStudent's t quantileone term, skewness knownHall's cubic, skewness knownHall's cubic, sample's own skewnessexact gamma tails given each of 200,000 sample shapesfaint lines: within 10% of nominal
Fig. 1 The actual size of a one-sided t test on thirty exponential draws, over its nominal level, for levels from 10% to one in a million: Student’s quantile; one Cornish–Fisher term and Hall’s cubic with the source’s skewness; and Hall’s cubic with each sample’s own. The faint lines mark 10% either side of nominal, and a line that leaves the plotted range is cut there. The slider moves the test to the lower tail, which is the side an upper limit on the mean reads.

A statistic skewed the other way

The statistic is T=n (xˉ−μ)/sT = \sqrt n\,(\bar x - \mu)/s. On an exponential source a sample whose mean happens to be large usually got there by including one or two large draws, and those draws inflate ss as well, so a large mean is divided by a large spread and TT is held back. A sample whose mean is small has missed the large draws, and its spread is small too, so a small mean is divided by a small spread and TT is pushed further out. The source’s long upper tail becomes the statistic’s long lower tail, which is what where the two tails disagree counted for a t interval as misses on one side four times as often as on the other.

Hall’s expansion of the distribution of TT has a first term of the same form as a sum’s, with a different polynomial and the opposite sign:

P(T≤x)=Φ(x)+γ6n(2x2+1) φ(x)+O(1/n),P(T \le x) = \Phi(x) + \frac{\gamma}{6\sqrt n}(2x^2 + 1)\,\varphi(x) + O(1/n),

where γ\gamma is the source’s skewness, 2 for the exponential. Inverting it gives the one-term Cornish–Fisher critical value

c1=z−γ6n(2z2+1),c_1 = z - \frac{\gamma}{6\sqrt n}(2z^2 + 1),

with zz the normal quantile for the level. Both tails move the same way — down — because the correction pulls the whole distribution of TT towards its long lower side: upper critical values come in, lower ones go out.

Hall’s 1992 transformation is the other construction measured here. It replaces TT by

g(T)=T+aT2+13a2T3+γ6n,a=γ3n,g(T) = T + aT^2 + \tfrac13 a^2 T^3 + \frac{\gamma}{6\sqrt n}, \qquad a = \frac{\gamma}{3\sqrt n},

which agrees with the one-term correction to first order and has derivative (1+aT)2(1 + aT)^2, never negative, so its critical values g−1(z)g^{-1}(z) are guaranteed to move outward as the level falls. The side a bound is read from used it at 2.5% and found it the best of four constructions for an upper limit; this essay follows it down to one in a million.

The tails are exact given the sample’s shape. A sample of nn exponential draws is its total SS, a gamma of shape nn, times its shape D=X/SD = X/S, which is uniform on the simplex and independent of SS. The standard deviation is SS times the shape’s own, and the sample skewness is a function of the shape alone. So for a fixed shape the event T>cT > c is an event about SS alone, a gamma tail computed exactly, and a size is that exact tail averaged over 200,000 drawn shapes. Only the shape is left to chance, the averaging carries its own standard error — at thirty draws a few thousandths of the size ratio at the levels quoted below — and the method agrees with TT counted directly on three hundred thousand samples. A critical value built from the sample’s own skewness is handled the same way, since that skewness is fixed by the shape.

Worse with the truth than without it

At thirty draws, with the source’s skewness supplied exactly:

thirty draws Student’s quantile one term, γ\gamma known Hall, γ\gamma known one term, sample’s γ^\hat\gamma Hall, sample’s γ^\hat\gamma
upper, 5% 2.248% 7.322% 5.828% 5.589% 4.745%
upper, 1% 0.1531% 2.944% 1.509% 1.496% 0.8775%
lower, 5% 9.810% 6.628% 4.952% 7.695% 6.824%
lower, 1% 3.976% 1.950% 0.5512% 2.729% 2.090%

On a sum the one-term correction was the whole repair: at thirty draws it took a 0.5% upper test from twice its nominal size to within half a per cent of it. On the t statistic the same term, given the exact skewness, takes the upper 1% test from 0.1531% — Student’s quantile is far too cautious on the short side — past the target to 2.944%, nearly three times too often. On the long side it halves Student’s excess and still rejects about twice as often as it claims.

The reason is the size of what the first term leaves out. For a sum, the skewness of the standardised total is γ/n\gamma/\sqrt n and the next term is a kurtosis correction of order 1/n1/n; for TT the skewness is about twice as large, −2γ/n-2\gamma/\sqrt n, and the 1/n1/n term carries the square of the source’s skewness, its kurtosis and the covariance between the mean and the variance. With γ=2\gamma = 2 and n=30n = 30 those are not small, and a correction that removes only the first-order piece exactly has removed the part it knows about and left an error of the same size in the part it does not. The sample’s own skewness, which is biased low — its median is 1.38 at thirty draws against the true 2 — happens to correct by less, and at 5% on the upper side that is closer to right.

That is the first result, and it would not have been guessed from the sum: the one-term correction for a t statistic is better without the true skewness than with it, at least at thirty draws on this source. It is not a virtue of estimation. It is one bias partly cancelling another, and the rest of the essay is about where the cancellation holds and where it does not.

A critical value that falls as the level does

The hero figure’s upper side shows the one-term line with the source’s skewness climbing away from nominal as the level falls, to thirteen times too many rejections at one in a thousand and ten thousand times at one in a million. At ten draws the reason is visible in the critical value itself.

The upper-tail critical value for a t statistic on 10 exponential draws, exact and by three constructions. T's own upper critical value rises from 1.065 at 10% to 5.457 at 10⁻⁶. The one-term value is highest, at 1.080, at a level of 8.85e-3, and falls beyond it to −0.115 at 10⁻⁶ — a stricter test given a laxer threshold. Hall's cubic keeps rising, to 2.748; Student's quantile reaches 10.720.
Fig. 2 The upper-tail critical value for a t statistic on ten exponential draws, against the level it is built for: the statistic’s own quantile, Student’s, the one-term Cornish–Fisher value and Hall’s cubic, the last two with the source’s skewness. The dashed line marks the level at which the one-term value stops rising. Student’s quantile leaves the plotted range above.

The derivative of c1c_1 with respect to zz is 1−2γz/(3n)1 - 2\gamma z/(3\sqrt n), which is zero at z=3n/(2γ)z = 3\sqrt n/(2\gamma). For a sum the corresponding turn was on the lower side at 1.5n1.5\sqrt n standard deviations. Here it is on the upper side — the studentised statistic’s short one — at 0.75n0.75\sqrt n, half as far out. At ten draws that is z=2.372z = 2.372, a level of 0.885%, where the one-term critical value reaches its highest point, 1.080. The statistic’s own critical value at that level is 1.991, and it goes on rising; the one-term value comes back down, to 0.972 at one in a thousand, and below zero by one in a million.

So a one-sided t test on ten skewed observations, corrected by the textbook term and asked for a level of 1%, is already just above the level where its correction starts to work backwards. At 1% it rejects 9.65 times too often; at 0.5%, 19.7 times; at one in a thousand, 121 times. Even at 5% the corrected upper test rejects 12.18% of the time, where Student’s quantile rejects 1.39%.

The turn recedes with the sample, as the sum’s did, but from a much nearer start: at thirty draws it is at a level of 2.0×10−52.0 \times 10^{-5}, where the critical value tops out at 1.993, and by one in a million a thirty-draw upper test built on the one term rejects 1.040% of the time — a test announcing one in a million and delivering one in a hundred. For a sum of thirty the turn had been at about 10−1610^{-16}. The studentised correction’s polynomial, 2z2+12z^2 + 1 against the sum’s z2−1z^2 - 1, is twice as steep, and its coefficient carries the opposite sign, so the turn arrives on the side where the corrected value is being pulled in; there is no hard edge nearby to hide it.

The floor under a monotone repair

Hall’s cubic cannot turn back, and the figure above shows its upper critical value rising steadily. It pays for that on the other side.

The lower-tail critical value for a t statistic on 30 exponential draws, exact and by three constructions. T's own lower critical value falls from −1.684 at 10% to −11.810 at 10⁻⁶. Below a level of 3.71e-3 Hall's cubic sends the critical value past its flat point at −8.216, reaching −15.707 at 10⁻⁶. The one-term value reaches −7.564 and Student's quantile −5.917.
Fig. 3 The lower-tail critical value for a t statistic on thirty exponential draws: the statistic’s own quantile, Student’s, the one-term value and Hall’s cubic, the last two with the source’s skewness. The dashed line is Hall’s floor, the level below which its critical value lies beyond the cubic’s flat point.

The cubic’s derivative (1+aT)2(1 + aT)^2 is never negative, but it is zero at T=−1/a=−3n/γT = -1/a = -3\sqrt n/\gamma, and near that point gg is nearly flat. Every normal quantile below g(−1/a)=−n/γ+γ/(6n)g(-1/a) = -\sqrt n/\gamma + \gamma/(6\sqrt n) is sent past the flat point, and the inverse there is steep: a small change in the level moves the critical value a long way out. At thirty draws the flat point is at −8.216-8.216 and the floor is a level of 0.371%. Above it, Hall’s lower critical value stays within about half a unit of the statistic’s own — −4.072-4.072 against −3.590-3.590 at 1%; below it, it drops away, to −12.587-12.587 at one in a thousand where the statistic’s own quantile is −5.477-5.477, and the test all but stops rejecting.

The table’s lower 1% row already shows the approach, 0.5512% against 1%. At one in a thousand the known-skewness Hall test rejects 0.0000482%, about a two-thousandth of its level, and at one in a million about a thirty-fifth. At ten draws the floor is at a level of 7.00%, above every conventional level, and a lower 5% test built this way rejects 0.419% of the time.

The two failures are mirror images and they have one cause. A correction that is a polynomial in zz with a positive square term has to bend somewhere. The one-term correction bends by turning back, which makes a test lax and says so only if someone checks; Hall’s cubic bends by flattening, which makes a test strict and says so only by never rejecting. Neither is a defect of the particular formula, since any smooth function that agrees with the one-term correction near the centre and is monotone must spend its curvature somewhere; the cubic chooses to spend it on the long side, where the statistic’s real quantiles run out farther than any low-order polynomial can follow.

The skewness the sample supplies

A reader with thirty observations does not know γ\gamma. The version of either correction anyone can use substitutes the sample’s own skewness γ^\hat\gamma, and since γ^\hat\gamma is a function of the sample’s shape, its tails are as exact here as the others. The last line of the hero figure is that version of Hall’s cubic.

On the upper side it is the best construction measured, and by a wide margin. At thirty draws it rejects 0.8775% at a nominal 1%, 0.810 times its level at one in a thousand, 0.842 at one in ten thousand, 1.04 at one in a hundred thousand and 1.59 at one in a million — within 20% of nominal for four decades of level, on the side where the known-skewness version is fifty times too lax at one in a million and the one-term version ten thousand. The one-term version built from γ^\hat\gamma does no better than its known-skewness counterpart in the far tail, because it still turns back, now at a level that varies from sample to sample, and at one in a million it rejects 1.72×1041.72 \times 10^4 times too often.

On the lower side the ranking is the same and the numbers are not good. Hall’s cubic with γ^\hat\gamma rejects 2.090% at a nominal 1%, 5.04 times its level at one in a thousand and 246 times at one in a million. That is a real improvement on Student’s quantile — 3.976%, 12.7 times and 591 times — and it is not a correction that can be quoted as holding its level. The lower side is where a safety margin set from a t interval is read, so it is the side that matters most, and it is the side on which reading the skewness from the sample recovers least.

One point of comparison with that essay is worth making exactly. There, Hall’s cubic was inverted at Student’s quantile rather than the normal’s, and at a two-sided 95% its upper limit was exceeded 3.31% of the time on twenty thousand simulated samples. With Student’s quantile inside, the averaging here gives 3.458% for the same event; with the normal quantile, the version measured in this essay, 4.013%. The two constructions differ by about half a point at this level, and the choice of which quantile sits inside the cubic matters more than the sampling error of the earlier count.

The samples that hide their skewness

Why the sample’s skewness rescues one side and not the other is visible when the samples are sorted by it.

The size of a 1% lower-tail t test on 30 exponential draws within each tenth of samples by their own skewness. Sorted by the sample's own skewness, from a median of 0.63 in the lowest tenth to 2.66 in the highest, the lower-tail test at 1% rejects ×6.83 to ×1.76 of nominal with Student's quantile, ×6.37 to ×0.48 with one term and ×6.14 to ×0.00 with Hall's cubic, both built from that skewness. Overall: ×3.98, ×2.73 and ×2.09.
Fig. 4 The size of a 1% lower-tail t test on thirty exponential draws within each tenth of samples, sorted by their own skewness from lowest to highest, for Student’s quantile and for the one-term and Hall corrections built from that skewness. The tick labels are each tenth’s median sample skewness; the dashed line is the nominal level.

A lower-tail failure — a sample whose mean is so far below the truth that TT falls below the critical value — happens when the sample has missed the source’s large values. A sample that has missed them also looks less skewed than the source, because the large values are where the skewness lives. So the samples that make a lower-tail t test fail are disproportionately the ones whose own skewness gives the least warning.

The figure shows how strong that is. In the tenth of samples with the lowest skewness, median 0.63, Student’s quantile rejects 6.83 times its level, and a correction built from that skewness can do little: 6.37 times with one term, 6.14 with Hall’s cubic. In the tenth with the highest, median 2.66, Student’s quantile rejects 1.76 times its level, and the corrections, now given a large γ^\hat\gamma, over-correct to 0.48 and zero. Averaged over all ten, the corrections halve Student’s excess, and the averaging is the whole of the improvement: they add their correction mostly to the samples that needed it least.

The upper side runs the opposite way and that is why the estimate works there. An upper-tail failure needs a sample whose mean is far above the truth with a spread that is not, which is a sample that has caught the large values, and such a sample shows its skewness plainly. The correction then arrives on exactly the samples that need it, and within tenths the Hall construction ranges only from 1.72 to 0.39 of its level at 1%, about a level that averages out near one.

This is a general feature rather than a quirk of the exponential. Whenever the statistic’s failures on one side are produced by samples that under-represent the source’s tail, any correction that reads the tail’s weight off the sample is blind on that side, and more sampling of the same kind — a larger bootstrap, a better skewness estimator — cannot see what the sample does not contain. A tail the sample never saw found the same thing for a sum’s tail probability read through the sample’s own cumulant generating function; here it appears one level up, as a critical value that is wrong on precisely the samples it is applied to.

How fast the long side closes

The failures above are at thirty draws. Whether the long side’s excess is a small-sample problem that goes away, and how fast, is a question about the sample size rather than the level.

The actual size of a 1% lower-tail t test on exponential draws against the sample size, for four critical values. At 1% on the lower side, Student's quantile rejects ×6.28 of nominal at ten draws, ×3.98 at thirty and ×1.95 at two hundred. Hall's cubic with the sample's own skewness: ×4.99, ×2.10, ×1.12. With the source's skewness known, the one-term value gives ×1.96 at thirty and Hall's cubic ×0.553.
Fig. 5 The actual size of a 1% lower-tail t test on exponential draws, over its nominal level, against the number of draws from ten to two hundred: Student’s quantile, the one-term value and Hall’s cubic with the source’s skewness, and Hall’s cubic with the sample’s own. Values below a hundredth of nominal are drawn at the floor of the axis.

Student’s quantile starts at 6.28 times its level at ten draws and is still at 1.95 times at two hundred: an upper limit for a mean, set from a t interval on two hundred skewed observations, fails about twice as often as its level says. Hall’s cubic with the sample’s skewness starts at 4.99 and reaches 2.10 at thirty, 1.27 at a hundred and 1.12 at two hundred, so its excess over nominal shrinks about nine times while Student’s shrinks about three; the one-term version built from γ^\hat\gamma follows a little above it. Neither is within 10% of its level anywhere on the figure.

The known-skewness Hall line is the odd one. It is far too strict at ten, fifteen and twenty draws, because Hall’s floor is then above 1% and the critical value sits past the flat point; the floor passes below 1% between twenty and thirty draws, and from seventy-five on the line sits within 3% of nominal. It is the most accurate construction in the figure once the sample is large enough — and it is the one no analyst can use, since it needs the source’s skewness exactly.

What a limit on thirty skewed observations can claim

An upper confidence limit for a mean — a safety limit on a concentration, a failure time’s expected value — is read from the lower tail of TT: the limit xˉ−c s/n\bar x - c\,s/\sqrt n with cc the lower critical value is exceeded by the true mean exactly when TT falls below cc. On thirty exponential observations the measurements above say:

A limit set from Student’s quantile fails about four times as often as its level at 1%, and nearly thirteen times at one in a thousand. That is the error the side a bound is read from found at 2.5%, followed further out, and it keeps growing as the level falls: 591 times at one in a million.

A limit set from Hall’s cubic with the sample’s skewness halves that and no more: about twice its level at 1%, five times at one in a thousand. It is the best construction a reader without the source can build from a formula, and it is not good enough to state a level below about 5% with any confidence on thirty observations.

The one-term correction should not be used on the upper side at all below a few per cent on small samples, since it turns back at 0.75n0.75\sqrt n standard deviations and its rejection rate on the short side climbs past its level long before that. With the sample’s skewness in place of the source’s, the turn moves around from sample to sample and does not go away.

And with the source’s skewness known, neither correction holds a ten-draw t test within 10% of its level even at 10%, on either side, where the same term held a sum’s critical value within 10% to 3.66 standard deviations. The studentised statistic is not a harder version of the sum’s problem; the estimated spread doubles the skewness, adds a correlation the sum did not have, and pushes the error that the first term leaves into a size where the first term is no longer the main part of the repair.

What these numbers rest on is the conditioning identity: given the shape of a sample of exponential draws, the t statistic’s tail is a gamma tail in the sample total, exactly. The standard errors of the averages over shapes are printed beside the sizes in the underlying calculation and are a few thousandths of the ratio at the levels in the table; at one in a million they are a few per cent of the ratio, and no conclusion here rests on a digit they would move. Not measured: sources other than the exponential, where the shape and the total are no longer independent and the conditioning does not apply; and the second Cornish–Fisher term for TT, whose coefficients involve the source’s fourth moment and the covariance of mean and variance, and which would need its own turning-point analysis before it could be trusted further out than the first.

Still open: the saddlepoint for a ratio

The construction that did not turn back for a sum, and did not flatten either, was the saddlepoint, which is built at the threshold rather than as a polynomial about the centre and so has no curvature to spend. For a studentised statistic the saddlepoint exists — it tilts the joint distribution of the sum and the sum of squares — and it is harder to write down, which is why nobody quotes it.

Two things about it are measurable on this source with the method used here, since the conditioning on the shape gives exact tails to compare against. The first is whether, with the source known, its error stays flat out to one in a million on both sides, which would make it the first construction on the t statistic that did. The second is whether its empirical version, built from the sample’s own cumulant generating function, fails on the long side for the reason found above — a sample that has missed the large values cannot tilt towards them — and so whether the edge the empirical tilt runs into is the same edge that stops every sample-based correction for TT. If it is, the long side of a t statistic on skewed data is a limit of what a sample can say rather than of any formula, and the only remaining repair is a model for the tail, of the kind a level with no data in it fits on purpose.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Cornish–Fisher expansionCritical valueEdgeworth expansionEstimated varianceMonotone transformationSkewnessStudent's tTail probabilityUpper confidence limit