The level a limit should be set at
Worth reading first: The correction for not knowing the spread.
The side a bound is read from priced upper limits for a skewed mean under a stated loss — each decision pays its limit’s margin, and pays a penalty whenever the true mean turns out to exceed the limit — and compared them all at their nominal 97.5%. It ended on the comparison that holds fixed the wrong thing. Under a stated penalty the cheapest limit is generally not at 97.5% at all: each construction has its own best level, where the margin it would add by going stricter equals the failures it would save, and the comparison that matters is between the constructions each at its best level. A limit from the family it came from then added a fourth candidate, the gamma family’s exact limit with its shape estimated.
This essay runs the comparison. It finds that the level is doing most of the work.
Four constructions, each free to move
A limit is a formula and a level. The t limit is , with the quantile at level ; Hall’s transformation inverts its cubic at the same quantile; the gamma family limit uses the gamma quantile at ; and a fixed multiple of has a multiplier in place of a level. Moving trades margin for failures along a curve that belongs to each construction, and under a loss of margin plus times the failure rate each curve has a lowest point.
The penalty is in units of the quantity’s own standard deviation: a penalty of twenty means that one failure costs as much as twenty decisions’ worth of one standard deviation of margin. For a safety limit, where a failure is the event the limit exists to prevent, twenty is modest; the previous comparisons used five, twenty and fifty, and so does this one. A concrete reading: if a contaminant limit set one standard deviation higher costs a plant a certain amount in extra treatment every batch, a penalty of twenty says that one batch released over the true limit costs twenty such amounts — a recall, a fine, a lost contract — which is a small multiple for most of the cases where limits are set.
Each at its best level
On fifteen exponential observations, six thousand samples, the same samples for every construction:
| failure costs | t limit, best level | Hall’s, best level | gamma family, best level | fixed multiple, best | t limit at 97.5% |
|---|---|---|---|---|---|
| 5 | 0.904 at 98.5% | 1.065 at 92.5% | 0.918 at 95% | 0.901 | 0.928 |
| 20 | 1.263 at 99.95% | 1.862 at 98.5% | 1.308 at 99.25% | 1.255 | 2.143 |
| 50 | 1.532 at 99.99% | 2.426 at 99.25% | 1.650 at 99.75% | 1.529 | 4.573 |
The t limit at its own best level is within one per cent of the best fixed multiple at every penalty. That is not a coincidence: the t limit is a fixed multiple of , with the multiplier set by its level, so moving its level sweeps the whole family of fixed multiples, and the best of those was the one-sided oracle of the essay on the bound’s side. The oracle’s advantage over the t limit was never in its formula. It was entirely in its level.
At fifteen observations the fitted family is close behind the t limit and Hall’s transformation is far behind both: at a penalty of twenty its best loss, 1.862, is 47% higher than the t limit’s best, and at fifty, 58% higher. At thirty observations the family limit edges ahead — a loss of 0.812 against the t limit’s 0.839 at a penalty of twenty — as its estimated shape settles, and Hall’s is still last, at 1.093.
Where the margin went at 97.5%
Read with the level in mind, the comparison at 97.5% already contained the answer. The point that reached 2.5% was a multiple of — the t limit’s own formula — and the only thing distinguishing it from the t interval’s point was how large a multiple it used: 3.43 at fifteen observations against the t interval’s 2.14. That is the t limit at a nominal level of about 99.8%. Every other point on that figure was a different formula at the same level; the best point was the same formula at a different level.
That reading also explains why the symmetric widening came second there. Widening both ends to exactly 95% total coverage raised the t limit’s multiplier to reach a two-sided target, which moved its upper end part of the way towards the level a one-sided target needed and stopped short, because a two-sided target gives the short side the same weight as the long one.
Why no level rescues Hall’s
The comparisons on the bound’s side and on the fitted family found Hall’s transformation adaptive in the wrong direction: it reads the skewness from the sample, the samples that miss are the ones that report little skewness, and so it spends margin where it is not needed and withholds it where it is. That was a statement about the transformation at 97.5%. Moving the level scales how much it spends but not where, and the curve in the hero figure shows the consequence: Hall’s curve bottoms out higher than either of the others and turns up sooner, because pushing its level stricter mostly adds margin to samples that were never going to fail.
A fixed multiple distributes its margin in proportion to , which is small on the dangerous samples — and a small on a low-mean sample is exactly what a stricter multiplier needs in order to reach the samples that fail. A correction that scales with the sample skewness distributes its margin in proportion to a quantity that is smallest on the dangerous samples, and no rescaling fixes a distribution that points the wrong way.
The level a penalty asks for
The table’s most practical content is the second number in each cell.
At a penalty of five the t limit’s best level on exponential data is 98.5%, close to convention. At twenty it is 99.95%, and at fifty 99.99%. The conventional 97.5% is the best level only when a failure costs a few standard deviations of margin, which for a safety limit is the case where the limit hardly matters. A skewed source asks for more: the best level is higher for a squared-normal source than for an exponential, and higher still for a lognormal, because the t limit fails more often on each and the penalty pushes it further out to compensate.
The loss of staying at 97.5% is large. At a penalty of twenty the t limit’s loss at 97.5% is 2.143 against 1.263 at its best, 70% more; at fifty, 4.573 against 1.532, three times as much. Nearly all of that is recovered by a single flat choice, as the next figure shows.
A flat choice of level
The best level depends on the source, which the analyst does not know. What the analyst does know is roughly what a failure costs, and the question is whether a level chosen for the penalty alone, ignoring the source, recovers most of the gain.
Setting the t limit at 99.9% regardless of the source costs 1% more than its best level on exponential data, nothing on a gamma of shape two, 11% on a gamma of shape one half and 24% on a lognormal. Against the 97.5% it replaces, it cuts the loss by between 37% and 45% on every source. Hall’s transformation at its own best level — which requires knowing the source — is worse than the flat 99.9% t limit on all four.
So the practical rule the numbers support is simple and unglamorous: for a one-sided limit on a skewed quantity, where a failure is expensive, move the level before changing the formula. A t limit at 99.9% is a limit anyone can compute, it needs no model and no estimate of skewness, and on these sources it beats every construction built to correct the t limit in those comparisons, at their conventional level or at their best.
When a failure costs fifty
At a penalty of fifty the ordering is the same and the stakes are higher. The t limit’s best level moves to 99.99% and its loss there is 1.532, three times smaller than the 4.573 it pays at 97.5%. Hall’s best level is 99.25% and its best loss 2.426; the family’s is 99.75% and 1.650. At this penalty an analyst who keeps the conventional level is paying for about seventy-five failures in every thousand decisions that a stricter level would have prevented, at the price of about seven tenths of a standard deviation of margin on each.
The curves also show why the best level differs by construction. A limit that fails rarely at a given level reaches the loss-minimising balance at a lower level, and the family limit, which is exact at its level when its shape is right, needs less push than the t limit, which fails three times as often as its level promises at 97.5%. Hall’s needs the least push of the three and gains the least from it, because the push goes to the wrong samples.
What fifteen observations can carry
All of this is at fifteen and thirty observations, which is where the skewness corrections were meant to earn their keep: large samples make the mean nearly normal whatever the source, and at a hundred and twenty observations the side a bound is read from found every construction within a point or two of its promise. The level question does not go away with a large sample, though. A larger sample shrinks the margin every construction needs and leaves the penalty where it was, so the loss-minimising level is set by the ratio of a failure’s cost to a margin that is now smaller — and the level stays strict.
Small samples add one more reason to choose the level deliberately. At fifteen observations a single large value moves the mean, the spread and the skewness at once, and the tail the sample never saw is the region a limit is protecting. A strict level is the one protection that does not depend on the sample having seen that region; it asks every limit, on every sample, to reach further into a tail it cannot estimate, which is the only safe response to not being able to estimate it.
What the level is standing in for
A reader may find it uncomfortable that the answer is “a stricter level” rather than “a better method”. The discomfort is worth examining, because it is the same one what a p-value does not say is about: a conventional level is a convention for a symmetric, nearly normal problem in which a false alarm and a miss cost about the same, and it carries no information about the problem in hand.
The 97.5% on a one-sided limit is the half of a two-sided 95% that happens to be read. On skewed data the side that is read is the side that fails more, so the level inherited from the two-sided convention is too lenient exactly where it is used; and when a failure costs many margins, the level that balances the two errors is much stricter than any convention. Both effects point the same way. The skewness corrections tried to repair the first by changing the formula. Moving the level repairs both, and because the t limit is already the right shape of formula — a multiple of the spread — it is the level that was wrong all along.
The formula still matters once the level is right, and the family limit shows how: at thirty observations its estimated shape brings it ahead of the t limit at a penalty of twenty, by 3%, though at fifty the t limit is ahead again, 0.945 against 0.982. A family that is known is worth using. But the difference between the best formula and the worst sensible one is a few per cent at a good level, and the difference between a good level and the conventional one is a factor of two.
A threshold chosen by what the errors cost
The argument has a familiar shape, because it is the one the test is a point somebody chose made about a diagnostic threshold. A screening test’s published sensitivity and specificity are one point on a curve, and the right point depends on what a miss and a false alarm cost; a limit’s nominal level is one point on its margin-against-failure curve, and the right point depends on what a margin and a failure cost. In both cases a convention — 95% specificity, 97.5% confidence — stands in for a cost ratio that nobody stated, and in both cases the convention is right only when the costs happen to match the ones it silently assumes.
The same move appears in sizing a study. The chance a trial succeeds replaced a conventional power with an average over what the effect might be; here a conventional level is replaced by the level that minimises an expected loss. And like a threshold that jumps, the best level can move a long way for a modest change in the costs — from 98.5% at a penalty of five to 99.95% at twenty — which is a reason to state the costs rather than the level, so that a reader with different costs can recompute it.
What distinguishes the limit case is that the formula it is attached to barely matters once the level is chosen. A diagnostic threshold is chosen on a fixed test; a limit’s level is chosen on whichever construction the analyst prefers, and the comparison here says that among sensible constructions the choice of formula is worth a few per cent and the choice of level a factor of two.
What moving the level buys each construction, and what it assumes about the penalty
On fifteen exponential observations with a failure costing twenty margins, the t limit at its best level, 99.95%, has an expected loss of 1.263 against 2.143 at 97.5%, within 1% of the best fixed multiple of at 1.255; Hall’s transformation at its best level has 1.862 and the gamma family limit 1.308.
At thirty observations and a penalty of twenty the gamma family limit at its best level leads, 0.812 against 0.839 for the t limit, and Hall’s is last at every penalty and sample size computed.
A flat 99.9% for the t limit costs between nothing and 24% more than its best level across four sources, and between 37% and 45% less than 97.5%.
Every loss is a count over six thousand seeded samples (four thousand for the four-source and penalty sweeps), the same samples for every construction at each setting, and the best level is the minimum over a grid of levels from 90% to 99.999%. The best level is itself chosen on the same samples it is scored on, which flatters every construction equally; the comparison between constructions is fair, and the absolute losses at the best level are a little optimistic.
Not claimed: that a penalty is easy to state. Converting a real cost into standard deviations of margin requires knowing the quantity’s spread, which the limit is being set to protect against not knowing. Not claimed either that the rule extends to limits read at both ends, where moving the level widens both sides and the arithmetic of the two errors changes.
Still open: the penalty that is not linear
The loss here charges a fixed penalty for any failure, whatever its size. A real exceedance of a safety limit usually costs more the further the mean lies above the limit: a concentration just over a threshold is a warning, one far over it is an incident. Under a loss that grows with the size of the failure, the constructions are compared on their expected shortfall rather than their failure rate, and the family limit — whose failures, when they happen, may be smaller than the t limit’s because it tracks the tail shape — could move ahead.
Whether it does is a calculation on the same samples with a different loss, and it would say whether “move the level” survives as advice once the cost of being wrong depends on how wrong.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- Where the two tails disagree — both name coverage, sample size, skewness, student's t
- The plot is about the wrong quantity — both name coverage, skewness, student's t
- A band allowed a few misses — both name coverage, sample size
- A block size that changes — both name coverage, sample size
- A coverage table with its own error — both name coverage, sample size
- A degrees of freedom that is not a count — both name sample size, student's t
Named objects
A flat tag is an object no other essay names yet.
CoverageDecision theoryExpected lossSample sizeSignificance levelSkewnessStudent's tUpper confidence limit