Shape, and what it does to a two-sample test

A failure charged by its size

Charge an upper limit's failure by how far the true mean exceeds it rather than by a fixed amount, and the level a limit should be set at falls further than any other change in this comparison moves it: on fifteen exponential observations, a penalty of twenty per standard deviation of shortfall asks for a t limit at 97% where a penalty of twenty per failure asked for 99.95%. The failures are small whichever limit produces them — about an eighth of a standard deviation — so a charge by size is a small charge per failure. The ranking hardly moves: the t limit stays within about 1% of the fitted family, whose failures are no smaller than its own, and Hall's transformation stays last by 11%.

Worth reading first: The correction for not knowing the spread.

The level a limit should be set at priced upper limits for a skewed mean under a loss of margin plus a fixed penalty for every failure, and let each construction choose its own level. The level did most of the work: at a penalty of twenty standard deviations of margin per failure, the t limit on fifteen exponential observations belonged at 99.95%, not 97.5%, and there it was within one per cent of the best any fixed multiple of the standard error could do. The fitted gamma family came close behind, and Hall’s transformation, which reads the skewness from the sample, was last at every level.

It ended on the loss most safety decisions actually carry. A concentration just over a limit is a warning; one far over it is an incident. Under a penalty that grows with the size of the failure, the comparison is on the expected shortfall, and the essay suggested that the family limit — which tracks the tail’s shape and so might fail by less when it fails — could move ahead, and asked whether “move the level” survives once the cost of being wrong depends on how wrong. The level moves further than it did, the ranking stays where it was, and the reason is that failures of a limit on fifteen observations are all about the same size.

Expected loss of three upper limits against the level each is set at, when a failure is charged by its size, fifteen exponential observations, penalty 20A failure costs 20 times the distance, in standard deviations, by which the mean exceeds the limit. At the conventional 97.5% the t limit's loss is 0.728, Hall's 1.136 and the fitted family's 0.835. Each at its own best level: the t limit 0.726 at 97.00%, Hall's 0.805 at 90.00%, the family 0.735 at 92.50%.50%90%99%99.9%99.99%012nominal level of the upper limitexpected loss, margin plus penalty × shortfallt limitHall's transformationgamma family, shape by likelihood97.5%6,000 samples, same samples for every limita failure charged by its size
Fig. 1 Expected loss against the nominal level each upper limit is set at, on fifteen exponential observations, when a failure is charged twenty times the distance by which the mean exceeds the limit, in standard deviations: the t limit, Hall’s transformation and the gamma family’s limit with its shape by likelihood. Rings mark each construction’s best level. The slider sets the charge per standard deviation of shortfall.

A charge by size

Each decision pays its limit’s margin above the true mean, in the source’s standard deviations, as before. What changes is the failure. Instead of a fixed penalty whenever the mean exceeds the limit, the decision pays

c⋅(μ−U)+σ,c \cdot \frac{(\mu - U)^+}{\sigma},

cc times the shortfall — nothing when the limit holds, a little when the mean is just over it, a lot when it is far over. The constructions, the samples and the sources are the earlier essay’s exactly: six thousand samples of fifteen, the same samples for every construction at every level, and each construction free to choose the level that minimises its expected loss.

With a charge of twenty per standard deviation of shortfall, each construction at its own best level:

exponential, fifteen best level loss at best fails mean shortfall loss at 97.5%
t limit 97% 0.726 8.97% 0.127 0.728
Hall’s transformation 90% 0.805 12.93% 0.135 1.136
gamma family, shape by likelihood 92.5% 0.735 9.20% 0.128 0.835

The t limit is still the best of the three, the fitted family is 1.1% behind it, and Hall’s transformation is 11% behind. Under the fixed penalty of twenty per failure the same three stood at 1.263, 1.308 and 1.862, at 99.95%, 99.25% and 98.5%. Every best level has come down — the t limit’s from a nominal chance of failure of 0.05% to 3%, Hall’s from 1.5% to 10% — and the ranking is the same.

Where the level went

The nominal level at which a t limit's expected loss is smallest when a failure is charged by its size, fifteen observations. For an exponential source the best level runs from 80.00%, 90.00%, 97.00%, 99.25%, 99.75%, 99.95% as a standard deviation of shortfall costs 5, 10, 20, 50, 100, 200 standard deviations of margin; for a lognormal source, 80.00%, 92.50%, 98.50%, 99.90%, 99.98%, 99.99%.
Fig. 2 The nominal level at which a t limit on fifteen observations has its smallest expected loss when a failure is charged by its size, against the charge per standard deviation of shortfall, for three sources. The dashed line is the conventional 97.5%.

On an exponential source the t limit’s best level is 80% at a charge of five, 90% at ten, 97% at twenty, 99.25% at fifty, 99.75% at a hundred and 99.95% at two hundred. The fixed penalty that asked for 99.95% was twenty; the charge by size that asks for it is two hundred.

The reason is the size of a failure. When a limit on fifteen observations fails, the true mean is on average about an eighth of a standard deviation above it: 0.127 for the t limit at its best level. So a charge of twenty per standard deviation of shortfall is, per failure, a charge of about 2.5, and a fixed penalty of 2.5 per failure asks for the same 97% — the earlier essay’s arithmetic, at a penalty an eighth the size. A charge by size looks like a heavier penalty, since it grows without limit, and in practice it is a lighter one, because the failures it is applied to are small.

That also returns the conventional level to respectability. At a charge of twenty per standard deviation of shortfall, the t limit at 97.5% has a loss of 0.728 against 0.726 at its best: convention is within 0.3%. What the earlier essay found — that the conventional level is right only when a failure costs a few margins — is the same finding in different units, since a failure that costs twenty margins per standard deviation of shortfall costs about two and a half margins per failure.

A more skewed source still asks for more. At a charge of twenty the t limit’s best level is 97% on the exponential, 98% on a gamma of shape one half and 98.5% on a lognormal, and at a hundred, 99.75%, 99.95% and 99.98%, because on the more skewed sources the t limit fails more often at any level and its failures, though no larger, are more frequent.

The failures are the same size

The suggestion that the family limit might pull ahead rested on its failures being smaller: a limit built from the tail’s shape should, when it misses, miss by less. It does not.

How far the mean exceeds each upper limit when it does, against how often, fifteen exponential observations. At a failure rate of about 5.4%, the t limit's failures average 0.117 standard deviations, Hall's 0.125 and the fitted family's 0.119. As each limit is set stricter its failures become rarer and a little smaller, along nearly the same curve for all three.
Fig. 3 The mean distance by which the true mean exceeds each upper limit when it does, in standard deviations, against how often the limit fails, as each limit’s level is swept, on fifteen exponential observations.

Set each construction’s level so that it fails as often as the t limit does at 99% — about 5.4% of samples — and the t limit’s failures average 0.117 standard deviations, the fitted family’s 0.119 and Hall’s 0.125. On a lognormal source, matched at 10.5%, the three are 0.108, 0.108 and 0.116. As a limit is made stricter its failures become rarer and a little smaller, and all three constructions move along nearly the same curve. Hall’s failures are slightly larger, which is what would be expected if they are concentrated on the samples it misjudges — the ones that hide their skewness; the fitted family’s are no smaller than the t limit’s.

The size of a failure is a property of the sample rather than the formula. A limit fails when the sample’s mean is low and its standard error is small, and how far the true mean then sits above the limit depends on how low and how tight the sample was, which is the same for every construction that uses the same sample. The constructions differ in which samples they fail on and how often; given that they fail, they fail by about the same amount. So a charge by size reweights every construction’s failures by nearly the same factor, and the comparison between them is nearly the comparison under a fixed penalty of the equivalent size.

Where the family does move ahead

It moves ahead a little in two places, and both are places the earlier essay’s comparison already pointed to.

How much more each upper limit's best expected loss is than the t limit's, when a failure is charged by its size, by sample size and penalty. Exponential source. 15 observations at penalty 20: the family +1.1%, Hall's +10.8%; 30 observations at penalty 20: the family −0.9%, Hall's +5.7%; 60 observations at penalty 20: the family −0.5%, Hall's +2.4%; 15 observations at penalty 50: the family +2.1%, Hall's +26.4%; 30 observations at penalty 50: the family −1.4%, Hall's +13.7%; 60 observations at penalty 50: the family −1.1%, Hall's +5.4%.
Fig. 4 Each upper limit’s best expected loss as a percentage above the t limit’s, when a failure is charged by its size, on an exponential source at fifteen, thirty and sixty observations, at charges of twenty and fifty per standard deviation of shortfall.

With more observations the fitted shape settles, and the family limit edges ahead: at a charge of twenty it is 1.1% behind the t limit at fifteen observations, 0.9% ahead at thirty and 0.5% ahead at sixty; at fifty, 2.1% behind, then 1.4% and 1.1% ahead. And on a lognormal source, which is not a gamma at all, the family limit is ahead at fifteen observations, 0.782 against the t limit’s 0.793. Neither is a margin that would decide a practical choice, and both are the same small advantage a limit from the family it came from found for a fitted family once its shape is estimated from enough data.

Hall’s transformation closes on the others as samples grow — 10.8% behind at fifteen, 5.7% at thirty and 2.4% at sixty, at a charge of twenty — for the reason the side a bound is read from gave: its correction is read from the sample’s skewness, and the samples that fail are the ones that under-report it, a defect that fades as the skewness estimate improves but that no choice of level repairs at small samples.

The price of pricing a failure the wrong way

The two losses ask for levels so far apart that a limit set for one is expensive under the other, and the expense is the practical content of the choice between them.

A t limit on fifteen exponential observations set at 99.95%, where a fixed penalty of twenty per failure wants it, sits on average 1.009 standard deviations above the true mean. Set at 97%, where a charge of twenty per standard deviation of shortfall wants it, it sits 0.499 above: half the margin. If the failure is really charged by its size, the stricter limit’s expected loss is 1.038 against the best 0.726 — 43% more, almost all of it margin spent guarding against failures that would have cost little. If the failure is really charged a fixed amount, the laxer limit’s expected loss is 2.292 against the best 1.263 — 82% more, almost all of it failures that each cost the full penalty however small they were.

So the choice of loss matters more than the choice of construction by a wide margin: the gap between the t limit and the fitted family at their best levels is about one per cent, and the gap between the right and the wrong loss is between 43% and 82%. It is also a choice the data cannot make, since both losses are evaluated on the same samples and differ only in what a failure is taken to cost. A plant that treats any exceedance as a reportable breach, whatever its size, has a fixed penalty; one whose harm is proportional to how far over the limit a batch went has a charge by size. The two should set limits sixty times apart in their nominal failure rates, and a protocol that fixes a level without saying which it has in mind has made the larger of the two decisions by default.

Failures shrink as samples grow

The equivalence between a charge by size and a fixed charge rests on the typical failure’s size, and that size is not a constant. It is set by how far a sample’s mean can fall below the truth while its standard error stays small, which shrinks with the square root of the sample: at the t limit’s best level under a charge of twenty, failures average 0.127 standard deviations on fifteen exponential observations, 0.087 on thirty and 0.060 on sixty, close to the factor of 1/21/\sqrt 2 each doubling predicts.

A charge by size therefore becomes a lighter charge per failure as the sample grows, and the best level stops rising: at thirty observations and at sixty the t limit belongs at 96%, and it fails 8.20% and 7.08% of the time there. Under a fixed penalty of twenty per failure the opposite happens — a larger sample makes each unit of margin buy more protection and the level asked for climbs, from 99.95% at fifteen observations to 99.98% at thirty and at sixty, where the limit fails 0.67% and 0.35% of the time. The two losses, which already asked for different levels at fifteen observations, diverge further as the data grow, which is one more reason to write the loss down before choosing the level rather than after.

It also connects this comparison to the correction a t test would use, which measured how far the t statistic’s long side misses its nominal level on skewed data. A limit’s failure rate at a given level is exactly that miss rate, and the failures’ sizes are how far the misses go. The first shrinks slowly with the sample, as the tail converges last would predict for a statistic read in its tail, and the second quickly, which is why, where the two tails disagree, the count of misses is the more persistent problem and the size of each miss the one the sample repairs on its own.

What changes and what does not

The level changes, and by more than any other decision in the comparison. Under a charge of twenty per standard deviation of shortfall the best level for a t limit on fifteen exponential observations is 97%; under a fixed charge of twenty per failure, 99.95%. The same number in two loss functions asks for nominal failure rates sixty times apart, and which loss is right is a question about the decision, not about the data.

The ranking does not. The t limit, the fitted family and Hall’s transformation come in the same order under both losses at fifteen observations, with the same gaps to within a few per cent. A reader choosing between them does not need to know which loss applies.

A failure’s size is nearly the same for every construction, so a charge by size behaves like a fixed charge of about an eighth of its value on fifteen observations. That is what makes the comparison stable, and it is worth knowing when writing down a loss: a penalty expressed per unit of shortfall should be divided by the typical shortfall before it is compared with a penalty per failure.

Move the level, and read the move from the loss. The earlier essay’s advice survives, with its sign made explicit: moving the level is right under any loss, but whether the move is up or down from 97.5% depends on the penalty per failure the loss implies, and a loss that charges by size implies a small one. The equivalence to watch is what a p-value does not say turned round: a level is not a property of a method, it is a price.

The losses are computed on the earlier essay’s six thousand samples of fifteen, with every construction evaluated at every level of a grid from 50% to 99.99% and the fixed-penalty losses reproducing the earlier essay’s to nine decimal places on the same samples. Failure sizes at a matched rate come from sweeping each construction’s level in steps of a fiftieth of a decade of failure rate and taking the level nearest the t limit’s rate; where a matched curve rests on fewer than thirty failures, below a failure rate of half a per cent, it is not drawn. Not measured: losses that charge more than linearly in the shortfall, where a limit’s largest failures would dominate and the construction that fails on the most extreme samples would be charged most; and losses that charge the margin non-linearly too, since a limit set very high has costs of its own that are rarely proportional.

The same arithmetic, three different decisions

It is worth being explicit about what the two losses are, because the arithmetic that compares the constructions is identical under both and the decisions it supports are not. One arithmetic, three decisions found a single regression right in one causal reading and wrong in two; here a single set of simulated samples, evaluated once, gives the right level under a fixed penalty, the right level under a charge by size, and a third right level for any mixture of the two, and nothing in the samples says which applies.

A regulator’s limit that triggers the same action for any exceedance is a fixed penalty. An expected-harm calculation in which the damage grows with the excess is a charge by size. A limit whose exceedance triggers a retest, and a fine only if the retest also exceeds, is something between them, with the fixed part set by the retest’s cost and the proportional part by the fine. Each of these is a legitimate reading of “the cost of being wrong”, and at a penalty of twenty they ask a t limit for nominal failure rates from 0.05% to 3% on the same data. The measurement this essay adds is that, whichever applies, the t limit is within about a per cent of the best construction on fifteen observations, so the loss decides the level and nothing else.

Still open: a level chosen without knowing the source

Every best level here was found knowing the source. On fifteen observations the t limit belongs at 97% if the source is exponential and 98.5% if it is lognormal, at the same charge, and the samples that would tell the two apart are the fifteen being used to set the limit. A level has to be chosen before the source is known, and the choice has a worst case.

The minimax level — the one whose worst excess loss over a set of plausible sources is smallest — is computable on the same samples for the same four sources, and so is its price: how much more than the source-specific best level it costs on each. Whether a single level near 98% is within a few per cent of the best on every skewed source at a charge of twenty, or whether the sources are far enough apart that any fixed level is expensive on one of them, has not been measured, and it would say whether “set the level from the loss” can be followed by an analyst who does not know the loss’s source.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Decision theoryExpected lossExpected shortfallSample sizeSignificance levelSkewnessStudent's tUpper confidence limit