A level set before the source is known
Worth reading first: The correction for not knowing the spread.
A failure charged by its size found the level an upper limit for a skewed mean should be set at when a failure costs twenty per standard deviation by which the true mean exceeds the limit, plus the limit’s own margin above the mean. On fifteen exponential observations the t limit belonged at 97%. On fifteen lognormal ones, at the same charge, it belonged higher. Each of those levels was found knowing the source, and an analyst setting a limit never knows it: the fifteen observations that would distinguish an exponential from a lognormal are the same fifteen being used to compute the limit, and fifteen is far too few to tell them apart.
The level has to be chosen first, and a choice made without the source has a worst case. The question the essay ended on is how bad that worst case is. If the sources put their best levels far apart, any single level is expensive on one of them, and “set the level from the loss” is advice an analyst cannot follow. If they put them close together, or if missing the best level is cheap, a single level serves.
Four sources and a worst case
The plausible sources are the four the earlier essays used for a mean that is skewed to the right: the exponential, a gamma of shape two which is milder, a gamma of shape one half which is harsher, and the lognormal, which is not a gamma at all and has the heaviest tail of the four. For each, six thousand samples of fifteen are drawn once, and every construction is evaluated on the same samples at every level of a grid running from 80% to 99.9995%, so the comparison across levels and across constructions is a comparison on identical data.
For each source the loss at each level is divided by that source’s own best loss. What remains is the regret: how much more the analyst pays for using that level on that source than an analyst who knew the source would have paid. The minimax level is the one whose largest regret across the four sources is smallest, and that largest regret is the price of not knowing which source the data came from.
The hero figure shows the four regret curves for the t limit under a charge by size. The sources do disagree about the best level. The gamma of shape two puts it at 96.25%, the exponential at 97%, the gamma of shape one half at 97.75% and the lognormal at 98.25% — a spread of two points, with the heaviest tail asking for the strictest limit, as it should. But every curve is shallow near its minimum. The minimax level is 97.25%, and there the regret is 0.1% on the exponential, 0.6% on the gamma of shape two, 0.2% on the gamma of shape one half and 0.8% on the lognormal. Not knowing the source costs less than one per cent of the loss, on every source.
Why the curves are flat
The shallowness has a reason, and it is the same reason the earlier essay found the failures small. Under a charge by size, moving the level up by a quarter of a point costs a little margin on every sample and saves a little shortfall on the few samples that fail, and both changes are small because the failures that the move prevents are the ones that only just failed. Near the best level the two changes cancel to first order — that is what being the best level means — and what is left is a second-order cost that grows slowly in either direction.
How slowly is visible in the conventional level. At 97.5%, a level nobody chose for these sources, the t limit’s regret is at most 1.0%: 0.3% on the exponential, 1.0% on the gamma of shape two, 0.1% on the gamma of shape one half and 0.5% on the lognormal. The textbook level is very nearly the minimax one under this loss, which is the earlier essay’s finding — that a charge by size brings the conventional level back as nearly the right one — extended from one known source to four unknown ones.
At thirty observations the picture holds. The sources’ best levels run from 96% to 98.25%, the minimax level is again 97.25%, and its worst regret is 0.9%. A larger sample shrinks every failure and every margin together, so the losses fall and the regret, which is a ratio, stays where it was.
What the minimax level’s loss is made of
At 97.25% the four sources pay for the same limit in different currencies. On the gamma of shape two, the mildest, the limit fails 6.67% of the time and sits 0.530 standard deviations above the mean on average; on the lognormal it fails 14.85% of the time and sits 0.457 above. The heavier tail fails more than twice as often at the same nominal level, which is where the two tails disagree seen from the limit’s side: a t limit’s long side misses more on a heavier-tailed source whatever level it is read at.
What keeps the regret small is the size of each failure. Given a failure, the shortfall averages 0.116 standard deviations on the gamma of shape two, 0.127 on the exponential, 0.129 on the gamma of shape one half and 0.115 on the lognormal. A heavier tail makes the t limit fail more often and not by more, so under a charge by size the extra failures are cheap ones, and a level that is slightly too lax for the lognormal costs it little. Under a fixed charge the same extra failures each cost the full twenty, which is why the next section’s curves are steep.
A minimax level is a cautious choice, and a reader may prefer an average. With the four sources weighted equally, the level that minimises the average regret is also 97.25% under a charge by size, where the average regret is 0.43%, and 99.97% under a fixed charge, where it is 3.56%. Caution and averaging agree on the level; they differ only in which number they report as its price.
The construction that was meant to adapt
The fitted gamma family’s limit estimates the source’s shape from the sample and builds the limit from the gamma pivot at that shape. It was introduced in a limit from the family it came from as the construction that knows what kind of source it is facing, and the natural expectation is that it needs the source least.
It needs it most. Its best levels are 90% for the gamma of shape one half, 92.25% for the exponential, 93.5% for the gamma of shape two and 96.5% for the lognormal — a spread of six and a half points where the t limit’s was two. Its minimax level is 93.5% and costs up to 3.4%, on the gamma of shape one half; and at the conventional 97.5% it costs up to 26.0%, because 97.5% is far above where three of the four sources want the family’s limit.
The side a bound is read from found a related asymmetry for two-sided intervals: a construction that does well on one tail by modelling the source does so by spending its accuracy where the model is right. The same is true of a level.
The reason is what the family’s level means. The family’s limit is calibrated exactly when its shape is right, so its nominal level is its true failure rate on a gamma source, and the best level then depends on how the loss trades margin against shortfall for that source’s tail. On the lognormal, which no gamma fits, the fitted shape is wrong in a way that makes the limit fail more often than its level says, and the level has to rise to compensate. The t limit’s nominal level is never its failure rate on a skewed source — on these sources it fails between 6.67% and 14.85% of the time at its minimax level — and its failures are of similar size whatever the source, so the same level serves all four. A construction that tracks the source makes its level a statement about the source; one that ignores the source makes its level a statement only about the loss.
A fixed charge puts the level out in the tail
Charging a failure a fixed amount, however small the shortfall, asks for a much stricter limit, and the sources disagree about how strict.
Under a fixed charge of twenty per failure the t limit’s best levels are 99.9% for the gamma of shape two, 99.95% for the exponential, 99.99% for the gamma of shape one half and 99.995% for the lognormal. On the failure-rate scale that is a spread of more than half a decade, and the curves are no longer shallow: a fixed charge makes every prevented failure worth the full twenty, however small it would have been, so a level slightly too low pays for many small failures at full price. The minimax level is 99.97% and costs up to 6.5%, on the lognormal, and 5.4% on the gamma of shape two at the other end. At thirty observations it is 99.99% and costs 6.2%.
What the strict level buys is visible in its failure rates. At 99.97% the t limit fails 0.37% of the time on the gamma of shape two, 0.97% on the exponential, 2.62% on the gamma of shape one half and 3.30% on the lognormal, and it pays for that with a margin of about one standard deviation — 1.107 on the mildest source and 0.957 on the heaviest, roughly twice the margin the charge by size asked for. The heavy-tailed sources still fail several times as often as the mild ones at the same level, and under a fixed charge each of those failures is a full twenty, so the minimax level is a compromise in which the lognormal pays for failures and the gamma of shape two pays for margin it did not need. Neither pays much more than six per cent for it, which is a smaller price than the shape of the curves suggests and a much larger one than under a charge by size.
The conventional 97.5% under a fixed charge is not a near miss. It costs up to 119.9% more than the best level, on the lognormal, and at least 59% on every source. The level a limit should be set at found the conventional level wrong for one known source under this charge; here it is wrong for every plausible source at once, and by more than the whole difference between any two of them.
Where each source puts the level
The four regret pictures can be condensed into one: where each source’s best level falls, for each construction and charge, with the minimax level marked.
Under a charge by size the t limit’s four dots sit within two points of each other around 97%; the fitted family’s spread from 90% to 96.5%. Under a charge per failure both constructions move out by a decade or more, and the two spreads become alike — the t limit’s from 99.9% to 99.995%, the family’s from 98.75% to 99.9%, each a little over a decade on the failure-rate scale. So a fixed charge does not make the sources disagree more about the t limit than about the family; it makes disagreement expensive, because the curves stop being flat.
What the condensed picture leaves out is how much each spread costs, and that is the next figure.
The price of not knowing the source
The t limit charged by size is the cheapest case on the chart: under one per cent at both sample sizes. The fitted family charged by size costs 3.4% at fifteen observations and 4.1% at thirty. The t limit charged per failure costs 6.5% and 6.2%. The fitted family charged per failure is the dearest, 15.4% and 17.9%, because a fixed charge makes the family’s miscalibration on the lognormal expensive at every level that suits the gammas.
Against the best construction for each source, rather than the best level for one construction, the t limit at its minimax level is within 0.1% of the better of the two constructions on the exponential, 0.6% on the gamma of shape two, 0.2% on the gamma of shape one half and 2.3% on the lognormal, where the fitted family at its own best level is slightly ahead. An analyst who does not know the source loses less by using the simplest construction at one level than by using the adaptive one at its own minimax level.
What an analyst who does not know the source can do
Under a charge by size, set the t limit at about 97.25% and stop worrying about the source. The level costs under one per cent more than the best level on any of four skewed sources at fifteen or thirty observations, and the conventional 97.5% costs at most 1.0%. The skewness of the source changes where the best level is, and barely changes what it costs to miss it.
Under a fixed charge per failure, the level is near 99.97% and the source matters more. The minimax level costs up to 6.5%, and the conventional 97.5% is more than twice as expensive as the best level on the lognormal. The choice between the two charges, which a failure charged by its size priced at 43% to 82% when made wrongly, is still the larger decision.
Expect the heavier tail to fail more often and not by more. At the minimax level the lognormal’s limit fails more than twice as often as the gamma of shape two’s, by about the same amount each time, which is a tail the sample never saw in its plainest form: fifteen observations rarely contain the values that would warn the limit, and the tail converges last of everything a sample estimates.
Do not reach for the fitted family to escape the source. It needs the source more than the t limit does, because its level is calibrated to a shape it may have wrong; its minimax regret is four times the t limit’s under a charge by size.
Every loss is counted on six thousand samples per source, the same samples at every level and for every construction, on a grid of levels from 80% to 99.9995%. The minimax level is checked to have no larger a worst-case regret than any other level on the grid, and each source’s regret to be zero at its own best level. The conventional level is refused as a safe default without the source when failures are charged a fixed amount: its worst regret there is 119.9%.
Still open: a level that reads the sample
A level fixed before the data is one answer to not knowing the source. The other is a level that reads the sample — a limit whose level rises when the sample’s skewness is large, so that a sample that looks lognormal is treated more strictly than one that looks gamma. That is a rule with the data in it, and the correction a t test would use found that the samples that fail are the ones that hide their skewness; a level that reads the skewness would be lax exactly on the samples that most need strictness.
Whether a skewness-dependent level can beat the fixed minimax level’s worst regret on these four sources — or whether it is the skewness-estimated correction’s failure under another name — is computable on the same samples, and has not been computed. The fixed level’s worst regret of 0.8% under a charge by size is a small target, which suggests the answer, but a suggestion is not a measurement.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A degrees of freedom that is not a count — both name sample size, student's t
- Sums of almost anything — both name sample size, skewness
- The plot is about the wrong quantity — both name skewness, student's t
- The spread a pilot supplies — both name sample size, upper confidence limit
- The threshold a validation study can see — both name expected loss, sample size
Named objects
A flat tag is an object no other essay names yet.
Decision theoryExpected lossExpected shortfallSample sizeSignificance levelSkewnessStudent's tUpper confidence limit