Shape, and what it does to a two-sample test

The side a bound is read from

On thirty exponential observations the upper limit of a 95% t interval is exceeded by the true mean 6.38% of the time, against the 2.5% a safety margin set from it assumes. Widen the interval until its total coverage is exactly 95% and the upper limit is still exceeded 4.69% of the time. A symmetric repair fixes the number that is reported and not the one that is used; Hall's transformation, which bends the interval, takes the same rate to 3.31%.

Worth reading first: The correction for not knowing the spread.

Where the two tails disagree split a t interval’s misses by side and found them unequal on skewed data, and it ended on a question it could not answer from coverage alone: what the imbalance costs someone who uses the interval. A decision rarely reads both ends. A safety margin reads the upper limit, a claim of superiority reads the lower one, and the literature that evaluates corrections for skew scores them on the total.

The suspicion stated there was that the corrections that most improve total coverage are not the ones that most improve the side a decision reads. This essay counts it.

Four 95% intervals for the mean of 30 exponential observations, split by the side they miss onBoth bars should read 2.5%. The t interval misses below the mean on 6.38% of samples. Widened until its total is exactly 5%, it misses below on 4.69% and above on 0.30%. Hall's transformation misses below on 3.31% and above on 1.96%.t interval — below the mean6.38%above the mean0.74%widened to exactly 95% — below the mean4.69%above the mean0.30%shifted by the skewness term — below the mean6.05%above the mean0.83%Hall's transformation — below the mean3.31%above the mean1.96%the rule marks 2.5%, where both bars belong20,000 samples of 30the low bar is what an upper bound risks
Fig. 1 Four 95% intervals for the mean of thirty exponential observations, each split by the side it misses on. The darker bar in each pair is how often the whole interval falls below the true mean — the event that makes its upper limit unsafe. The rule marks 2.5%. The slider is the sample size.

The decision that reads one end

Take thirty measurements of a right-skewed quantity — a contaminant concentration, a repair time, a cost — and set a limit from the upper end of a 95% interval for its mean. The limit claims that the true mean exceeds it no more than 2.5% of the time. That is the promise a regulator, a warranty or a budget is built on, and it is a one-sided promise — the same kind of promise a one-sided p-value makes about its own error rate.

The t interval’s upper limit breaks it on 6.38% of samples. At fifteen observations, 8.31%; at a hundred and twenty, still 4.19%. The mean sits above a limit that was supposed to contain it between one and a half and three and a third times as often as promised, and the opposite end — which nobody in the decision reads — is correspondingly too safe.

A coverage table reports this interval as covering 92.88% at thirty observations, which reads as a modest shortfall that a slightly wider interval would fix. Whether a wider interval fixes it depends on how it is widened.

Four ways to build the interval

The comparison is between four constructions, chosen because each represents a kind of repair.

The t interval, xˉ±ts/n\bar x \pm t\,s/\sqrt{n}, is the baseline.

Widened to exactly 95%. The same interval with its multiplier raised until the total miss rate on these very samples is 5.00%. No analyst can do this, since it requires knowing the population; it is here because it is the best a symmetric repair can possibly do. Any method that works by widening both ends equally — a larger quantile, a small-sample adjustment to the degrees of freedom, a variance inflation — is doing a worse version of this.

Shifted by the skewness term. The t interval moved upward by γ^s/(6n)\hat\gamma\, s / (6n), where γ^\hat\gamma is the sample skewness. This is the centring half of Johnson’s 1978 correction for skew; the other half, a quadratic term, makes the statistic non-monotone and is left out, which is how it is usually applied.

Hall’s transformation. A cubic transformation of the studentised mean, from 1992, that carries the same skewness information as Johnson’s but is monotone and so can be inverted exactly. It does not move the interval rigidly; it stretches one end and compresses the other.

What each one does to the side

thirty observations below the mean above the mean total
t interval 6.38% 0.74% 7.12%
widened to exactly 95% 4.69% 0.30% 5.00%
shifted by the skewness term 6.05% 0.83% 6.88%
Hall’s transformation 3.31% 1.96% 5.27%

The widened interval achieves the total exactly and fixes less than half of the side. Its upper limit is still exceeded on 4.69% of samples, nearly twice the promise. What it did with the extra width is visible in the second column: the high misses fell from 0.74% to 0.30%, so a large share of the widening went to the end that was already too safe.

The skewness shift barely moves anything. Moving the whole interval up by a term of order 1/n1/n is the right direction and the wrong size: at thirty observations with a sample skewness near two, it moves the centre by about a hundredth of a standard deviation.

Hall’s transformation is the only one that changes the ratio. Its upper limit is exceeded on 3.31% of samples — still above 2.5%, and about a fifth of the t interval’s excess — while its lower limit is exceeded on 1.96%, much nearer its own 2.5%. It has a worse total than the widened oracle, 5.27% against 5.00%, and on the dimension a decision uses it is the best of the four by a distance.

That is the suspicion confirmed with numbers: ranked by total coverage, the widening wins; ranked by the side a bound is read from, it comes third.

How often an upper 97.5% limit for a mean is exceeded, exponential source, four constructions. The nominal rate is 2.5%. At fifteen observations the t interval's upper limit is exceeded on 8.31% of samples, the widened interval's on 4.93%, the shifted interval's on 7.99% and Hall's on 4.49%.
Fig. 2 How often the true mean exceeds each construction’s upper limit, as the sample grows from fifteen to a hundred and twenty. The rule marks the 2.5% the limit promises.

As the sample grows, and on a more skewed source

exceeded above the upper limit 15 30 60 120
t interval 8.31% 6.38% 5.19% 4.19%
widened to exactly 95% 4.93% 4.69% 4.26% 3.79%
shifted by the skewness term 7.99% 6.05% 4.92% 3.96%
Hall’s transformation 4.49% 3.31% 2.89% 2.56%

Hall’s transformation is nearest 2.5% at every size, and at a hundred and twenty it is there to within the counting error. The widened interval converges slowly, and not because widening is failing to converge: its total is exactly 5% at every size. It converges only because the skew itself fades as the sample grows. At fifteen observations the widened interval puts 4.93% of its 5.00% total on the low side — ninety-nine misses in a hundred on one end.

On a lognormal source, more skewed than the exponential, the pattern is the same and larger. At thirty observations the t interval’s upper limit is exceeded 11.93% of the time, the widened interval’s 5.00% — every one of its misses on that side — and Hall’s 6.17%. The best symmetric construction there is puts all of its misses where they hurt.

The t interval’s own imbalance is what the other three are repairing, and on the lognormal it is extreme enough to see without a table.

Which side a 95% t interval misses on, lognormal source. Both tails should be 2.5%. At 8 observations the interval falls short of the mean on 16.88% of samples and overshoots on 0.09%. At 500 they are 4.83% and 1.35%, and the total is 6.17% — which a coverage table reports as very nearly right.
Fig. 3 Where a 95% t interval misses on a lognormal source, split by side, as the sample grows from eight observations to five hundred. The two curves should both sit on the dashed 2.5%.

The low curve starts near seventeen per cent and is still near five at five hundred observations; the high curve is almost flat on the floor. Every construction above begins from that picture. Widening lifts the whole interval’s reach on both sides, so the low curve comes down and the high curve, already near zero, has nowhere to go and the widening there is wasted. A shift moves the pair together. Only a construction that treats the two ends differently — by stretching the upper end more than the lower — can move the low curve down without spending on the high one, which is what a transformation does and a correction for an estimated spread does not.

The same shape governs the two-sample case. When two groups are compared, the imbalance is set by the skewness of the difference of their means, and a symmetric reference distribution for that difference fails the same way a symmetric interval for one mean does here.

What the margin costs

A safety limit is not free: a higher limit is a more expensive product, a larger reserve, a stricter standard. So the fair comparison is not only how often each limit is exceeded but how high each puts the limit to get there.

What an upper 97.5% limit's margin buys, 15 exponential observations. The t interval's upper limit sits 0.526 standard deviations above the mean on average and is exceeded on 8.31% of samples. A fixed one-sided multiplier tuned to 2.5% needs 0.841. Hall's transformation spends 1.044 and is exceeded on 4.49%.
Fig. 4 Each construction’s upper limit at fifteen exponential observations, placed by how far above the true mean it sits on average and how often the mean still exceeds it. The fifth point is a fixed one-sided multiplier tuned to be exceeded exactly 2.5% of the time on these samples.

The margins, in population standard deviations above the true mean, at fifteen observations:

construction mean margin exceeded
t interval 0.526 8.31%
widened to exactly 95% 0.667 4.93%
fixed one-sided multiplier, tuned 0.841 2.50%
Hall’s transformation 1.044 4.49%

The tuned fixed multiplier is an oracle, like the widening: it is the constant kk in xˉ+ks/n\bar x + k\,s/\sqrt{n} that this population needs, 3.43 at fifteen observations where the t interval uses 2.14. It marks what an honest one-sided limit costs when its shape is fixed — 60% more margin than the t interval spends.

The surprise is Hall’s row. At fifteen observations it spends more margin than the oracle — 1.044 standard deviations against 0.841 — and is exceeded more often, 4.49% against 2.50%. It pays more and gets less.

Why the correction that reads the sample overspends

The reason is the one the Edgeworth correction ran into when its skewness was estimated rather than given. Hall’s transformation reads γ^\hat\gamma off the sample, and the sample skewness of a long-tailed source is biased low and highly variable: at fifteen observations it is often a fraction of the true value of two.

Now consider which samples produce an interval that misses low. They are samples with no large values in them — a low mean, a small spread, and, for exactly the same reason, a small sample skewness. The coupling between a sample’s mean and its spread that unbalances the t interval extends to the third moment: the samples on which the correction is most needed are the samples that report the least skew. Counted at fifteen observations, the samples whose Hall upper limit the mean exceeded have a median sample skewness of 0.71; the samples whose limit held have 1.15; the population’s is 2. On those samples Hall’s transformation barely adjusts, and they miss. On the samples that do contain a large value, the sample skewness is large, the transformation stretches the upper end a long way, and it spends margin on intervals that would have covered anyway.

So the correction is adaptive in the wrong direction. It is generous where generosity is wasted and stingy where it would have helped, and its average margin overstates the protection it provides. The fixed oracle does the opposite: it spends the same multiple of ss on every sample, and since ss is itself small on the dangerous samples, a large fixed multiplier lands its margin where the misses are.

A skewness correction with the skewness estimated from the sample, 2 standard deviations out. On exponential samples the sample skewness has a median of 0.93 at ten observations and 1.85 at three hundred, against a true value of 2. The correction built on it gives a median of ×0.834 of the exact tail at ten, where the true-skewness correction gives ×1.081 and the normal ×0.617.
Fig. 5 The same bias seen from the expansion’s side: a skewness correction to a tail probability, with the skewness taken from the sample, at four sample sizes. The band is the tenth to the ninetieth percentile over samples; its centre sits well short of the correction made with the true skewness until the sample is large.

The figure measures the bias on a different quantity — a tail probability rather than an interval’s end — and the mechanism is the one at work here. The sample skewness of a long-tailed source is short of the truth on most samples and very short on the samples that contain no large value, so any correction proportional to it is short on exactly those samples. The interval version adds the conditioning: the samples that are short of large values are also the ones whose intervals sit too low.

At sixty observations the overspending has nearly gone — Hall spends 0.325 standard deviations against the oracle’s 0.316 and is exceeded on 2.89% — and at a hundred and twenty the two are indistinguishable. The failure is a small-sample failure of the skewness estimate, not of the transformation.

Pricing a limit under a stated loss

Margin and exceedance can be put on one scale once a loss is stated. Suppose each decision pays its limit’s margin, measured in standard deviations of the quantity, and pays a penalty every time the true mean turns out to exceed the limit. Call that penalty cc, in the same units. The expected cost of a construction is its mean margin plus cc times its exceedance rate, and each pair of constructions has a break-even cc above which the more cautious one is cheaper.

At thirty exponential observations:

break-even penalty from to
2.6 t interval widened to exactly 95%
6.5 t interval Hall’s transformation
11.2 widened to exactly 95% Hall’s transformation

So at thirty observations Hall’s transformation is the cheapest of the three real constructions once an exceedance costs more than about eleven standard deviations of margin, and the plain t interval is cheapest only when an exceedance costs less than about three. For a safety limit, where an exceedance is the event the limit exists to prevent, a penalty of eleven margins is modest.

At fifteen observations the arithmetic changes because of the overspending measured above. The break-even between the widened interval and Hall’s rises to 84.8: Hall spends 0.377 more standard deviations of margin to cut the exceedance rate by less than half a point. Below that penalty the symmetric widening is cheaper, if it could be had — and it cannot, because it is tuned on the population. Against the t interval, which can be had, Hall’s transformation breaks even at a penalty of 13.6.

The tuned one-sided multiplier is the cheapest of all four at a penalty of twenty or fifty at both sizes, and at a penalty of five at thirty observations; at fifteen observations and a penalty of five the widened interval edges it, 0.914 against 0.966. That is the formal version of the conclusion below: once an exceedance is expensive, knowing the shape is worth more than estimating it, and the smaller the sample the more it is worth.

Two cautions about the table itself. The penalties are in units of the quantity’s own standard deviation, so the same penalty means different things for quantities of different spread, and a reader has to translate a real cost into those units before reading off a row. And the rates and margins are counted over twenty thousand samples each, so a break-even that rests on a difference of a few tenths of a point — the 84.8 at fifteen observations is one — is itself uncertain by a large factor, and should be read as “large” rather than as that number.

What this means for setting a limit

Three conclusions, in order of how widely they apply.

Evaluate a one-sided procedure on its one side. A method compared on total coverage can win by a margin and be worse for the purpose it is used for. The widened oracle, which no real method beats on the total, is beaten on the upper side by a method whose total is worse.

A symmetric interval on a skewed quantity cannot be repaired symmetrically. Every symmetric construction’s best case is the widened row, and its best case at fifteen exponential observations still puts ninety-nine of every hundred misses on the side a safety limit reads.

At small samples, a skewness correction estimated from the sample is not a substitute for knowing the skew. If the quantity’s skewness is known from the subject — a waiting time is close to exponential, a concentration close to lognormal — a one-sided multiplier built for that shape, like the tuned row, reaches the promised rate with less margin than a correction that has to discover the shape from fifteen points. Where it is not known, Hall’s transformation is still the best of the four on the side that matters; the reader should expect it to overspend.

Four 95% intervals for the mean of 30 lognormal observations, split by the side they miss on. Both bars should read 2.5%. The t interval misses below the mean on 11.93% of samples. Widened until its total is exactly 5%, it misses below on 5.00% and above on 0.00%. Hall's transformation misses below on 6.17% and above on 1.52%.
Fig. 6 The same four intervals on thirty lognormal observations. The widened interval’s misses are all on the low side; Hall’s are split, and still too many are low.

What the counts establish, and what they do not

The widened interval misses exactly 5% in total at every size, by construction, and at every size it puts a larger share of its misses on the low side than the t interval does. Both halves are stated together because the first is what makes the second a statement about symmetric repair rather than about an interval that happens to be too narrow.

Hall’s transformation brings the low side nearer 2.5% than the t interval and than the skewness shift, at every size from fifteen to a hundred and twenty on an exponential source.

Reaching 2.5% on the upper side takes more margin than the t interval spends, and the widened interval spends less than that and misses more than 2.5% — at every size and on both sources.

What does not survive is an interval widened until its total coverage is right, read as a safe upper limit. At thirty exponential observations its upper limit is exceeded on 4.69% of samples against 2.5%.

Not claimed: that Hall’s is the best available skew correction. A studentised bootstrap and a bias-corrected and accelerated bootstrap both carry the same skewness information by resampling, and the resampling essays price them on other grounds; they have not been counted on this comparison. And the two oracles are yardsticks, not methods — each was tuned on the samples it is scored on, which is exactly what no analysis can do.

Every rate is a count over twenty thousand samples, with a counting error of about 0.11 points near 2.5% and 0.17 points near 6%, so differences of a tenth of a point between rows are not read as findings.

Still open: a limit tuned to the shape rather than to the sample

The tuned fixed multiplier did best, and it cheated: it knew the population. The honest version of it knows only the population’s family — exponential, lognormal, gamma of some shape — and estimates the one parameter that decides the multiplier. Whether a limit built that way inherits the oracle’s efficiency or Hall’s overspending depends on how badly the shape parameter is estimated at fifteen observations, and that can be counted exactly as the four limits here were, and has not been.

The other direction is a limit that is wrong on purpose. The pricing above held every construction at its nominal 97.5% and compared the costs that followed. Under a stated penalty the cheapest limit is generally not at 97.5% at all: each construction has its own best level, found by moving its quantile until the marginal margin equals the marginal exceedance saved, and the comparison that matters is between the four constructions each at its own best level. That optimisation is a one-dimensional search per construction on counts this essay already has, and it has not been run.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Conditional coverageCoverageEstimated varianceSample sizeSkewnessSkewness correctionStudent's tUpper confidence limit