Extremes — the series
-
Three shapes, one limit
A normalised sum has one limit and a normalised maximum has three, indexed by a single number. Twenty blocks put the sign of that number right 97.3% of the time — and naming the family from a light-tailed record gets worse as the record grows, from 83.0% at twenty blocks to 4.8% at five hundred.
-
The maximum converges slowly
The rate at which a normalised maximum reaches its limit law is computable rather than simulable, because the exact law of a maximum is always available. For a normal parent the distance falls like one over the logarithm of the block and is still 0.0091 at a million readings; for an exponential parent, with the same limit, it is 2.707×10⁻⁷.
-
The threshold is a dial
A peaks-over-threshold analysis has one knob, and raising it buys accuracy with exceedances. For a normal parent the error is smallest at the 0.925 quantile and 80.6% of it is still bias there — and both diagnostics practitioners use to set the knob lose to a fixed 0.90 rule, one by a factor of 1.590 and one by 11.881.
-
A level with no data in it
The largest of fifty block maxima is a 51-block event by its own plotting position, so a hundred-block level is read 1.96 times past the longest event the record contains — and it lands above the largest reading on 52.4% of records. The estimate stays nearly unbiased out there; what grows is its error, sixfold from ten blocks to a thousand.
-
Two intervals for one return level
Two 95% intervals read off the same fits of the same records, against a level known in closed form. The symmetric one covers 80.3% at twenty-five blocks and reaches only 89.0% at two hundred — and 99.24% of its misses are the interval sitting entirely below the truth, which is not the endpoint anybody expects to fail.
-
The clustering the tail has
Every threshold method counts exceedances as though they were independent pieces of information, and in a dependent series they arrive in clusters. Ignoring that overstates a return level by the reciprocal of the extremal index — ×3.527 counted where the mean cluster holds four — and leaves a reported standard error 2.151 times too small.
-
The run length a declustering chooses
The runs estimator of an extremal index carries a constant nobody derives. Where a cluster is a run of neighbouring exceedances the constant barely matters; where a cluster's members fall six steps apart, the estimate is 0.9069 at a run length of six and 0.3649 at seven against an index of 0.40, and a run length of four removes under a tenth of the overstatement declustering exists to remove. A rule that reads the run length off the data has the smallest worst error of the three.