Concept

Sampling distribution — where it appears

The distribution of an estimate across repetitions of the study that produced it, which is what a standard error and a coverage rate are both statements about. Every rate in this collection is counted from it rather than derived from an approximation to it.

Named by 5 essays across 3 fields — each of them below, with the objects they name alongside it.

What each rule gives up against an oracle that is arithmetic. Expected squared error of the candidate each rule selects, minus the expected squared error of the best candidate in the table, over 500 draws of 120 rows. Both quantities are closed forms — σ_S²(1 + q/(n − q − 1)) — so the only Monte Carlo here is over which candidate got picked. The hold-out spends half its rows measuring what the criterion computes, and pays 1.8 times as much for it. Schwarz's criterion is worst because it is answering a different question: which candidate contains the truth, rather than which one forecasts best.

A criterion is a prediction of the hold-out

A rolling hold-out spends half the sample measuring what a criterion computes from all of it. Against an oracle that is arithmetic rather than an estimate, the criterion gives up 0.01701 and the hold-out 0.03200 — and the number the hold-out reports for its own winner is optimistic by more than either.

proxy · Order-selection
More blocks, and the light tail is called wrong more often. The share of records whose three-way family call is right, over 400 records at each block count, at 100 readings a block. A 95% interval for the shape is formed and the record is recorded as calling Fréchet, Weibull or Gumbel according to whether that interval sits above zero, below it, or straddles it. The two signed parents go from 51.5% and 71.0% at 20 blocks to certainty by 100. The light-tailed one goes the other way — 83.0%, 75.5%, 55.0%, 36.5%, 4.8% — because its estimate sits at about -0.1157 whatever the record length, and a longer record only shrinks the interval onto that number.

Three shapes, one limit

A normalised sum has one limit and a normalised maximum has three, indexed by a single number. Twenty blocks put the sign of that number right 97.3% of the time — and naming the family from a light-tailed record gets worse as the record grows, from 83.0% at twenty blocks to 4.8% at five hundred.

extreme · Extremes
The record stops here, and the curve does not. The level exceeded once in T blocks, against T, for a normal parent at 365 readings a block. The truth is closed form — the block maximum's own distribution function is Φ(x) raised to the 365, so the T-block level is Φ⁻¹ of the 365-th root of (1 − 1/T), with nothing fitted in it — and the fitted mean over 800 records of 50 blocks sits on top of it, 4.0186 against 4.0330 at 100 blocks. What moves is not the level but its error, which grows from 0.0956 at 10 blocks to 0.6383 at a thousand while the level itself moves only from 3.4421 to 4.5454. The rule marks the largest reading an average record contains, 4.0062: everything to the right of where it crosses is read from a fit rather than from data.

A level with no data in it

The largest of fifty block maxima is a 51-block event by its own plotting position, so a hundred-block level is read 1.96 times past the longest event the record contains — and it lands above the largest reading on 52.4% of records. The estimate stays nearly unbiased out there; what grows is its error, sixfold from ten blocks to a thousand.

extreme · Extremes
The interval every package reports first does not cover. Counted coverage of two 95% intervals for the 100-block return level of a normal parent, against the length of the record they were fitted from, over 300 records at each length. The level they are about is known in closed form, so this is coverage of a number rather than agreement between two estimates. The delta-method interval covers 80.3% at 25 blocks and reaches only 89.0% at 200; the profile-likelihood interval sits between 94.0% and 94.7% throughout. The gap is not a small-sample effect that lengthening the record removes — it narrows by 8.7 points for an eightfold longer record.

Two intervals for one return level

Two 95% intervals read off the same fits of the same records, against a level known in closed form. The symmetric one covers 80.3% at twenty-five blocks and reaches only 89.0% at two hundred — and 99.24% of its misses are the interval sitting entirely below the truth, which is not the endpoint anybody expects to fail.

extreme · Extremes
The p-value of a study with 80% power, twenty thousand times. Twenty thousand two-sided z-tests, each on 25 observations whose true mean is 0.5603 standard deviations from the null, a noncentrality of 2.802. The bars are the counted share of p-values in bins a quarter of a power of ten wide, with the leftmost bin holding everything smaller; the line is the closed form. The middle eighty per cent of the p-values runs from 4.4×10⁻⁵ to 0.13, 3.46 orders of magnitude, the median is 0.0051, and 80.0% fall below 0.05, which is what the power means.

The p-value a replication gets

Under a true null a p-value is flat. Under a real effect its distribution is closed form and wide — a study with 80% power returns anything from 4.4×10⁻⁵ to 0.13 in eight runs of ten — and the chance that an exact replication of a p = 0.05 result is significant again is exactly one half, under both of the models people use without naming them.

testing · Uniformity

Named alongside it

The objects these essays reach for when they reach for this one.

Block maximaGeneralised extreme valueShape parameterBlock sizeClosed formExtrapolationExtreme value theoryGumbel lawMaximum likelihoodMean squared errorMonte CarloReturn level

All concepts