the-tail-past-the-last-observation

More blocks, and the light tail is called wrong more often

The share of records whose three-way family call is right, over 400 records at each block count, at 100 readings a block. A 95% interval for the shape is formed and the record is recorded as calling Fréchet, Weibull or Gumbel according to whether that interval sits above zero, below it, or straddles it. The two signed parents go from 51.5% and 71.0% at 20 blocks to certainty by 100. The light-tailed one goes the other way — 83.0%, 75.5%, 55.0%, 36.5%, 4.8% — because its estimate sits at about -0.1157 whatever the record length, and a longer record only shrinks the interval onto that number.

The tail past the last observationwide7 views

What else it draws

The same object, drawn to answer the other questions the essays put to it.

The Kolmogorov distance between the exact law of a normalised maximum and its Gumbel limit, at six block sizes, for two parents that both have that same limit. Both are closed form: the exact law of a maximum is F(x)^n and no simulation is involved. The exponential parent's distance falls from 0.0280 to 2.707e-7 — a factor of a hundred thousand, which is exactly one over n. The normal parent's falls from 0.0522 only to 0.0091, a factor of 5.74, because its rate is one over log n. At a million readings a block the two differ by a factor of 33556.3.

The mean squared error of the estimated shape against the threshold, split into the square of its bias and its spread, over 600 records of 2000 readings from a a normal parent. At the 0.9 quantile 199 exceedances are left, the bias is -0.1708, the spread is 0.0701 and the total error is 0.0341. The bias falls as the threshold rises because the exceedances get closer to being generalised Pareto; the spread rises because there are fewer of them. The sum is smallest at the 0.925 quantile, at 0.0340, of which 80.6% is still bias — so even the best threshold on this grid is one where accuracy, not spread, is the binding constraint.

The level exceeded once in T blocks, against T, for a normal parent at 365 readings a block. The truth is closed form — the block maximum's own distribution function is Φ(x) raised to the 365, so the T-block level is Φ⁻¹ of the 365-th root of (1 − 1/T), with nothing fitted in it — and the fitted mean over 800 records of 50 blocks sits on top of it, 4.0186 against 4.0330 at 100 blocks. What moves is not the level but its error, which grows from 0.0956 at 10 blocks to 0.6383 at a thousand while the level itself moves only from 3.4421 to 4.5454. The rule marks the largest reading an average record contains, 4.0062: everything to the right of where it crosses is read from a fit rather than from data.

Counted coverage of two 95% intervals for the 100-block return level of a normal parent, against the length of the record they were fitted from, over 300 records at each length. The level they are about is known in closed form, so this is coverage of a number rather than agreement between two estimates. The delta-method interval covers 80.3% at 25 blocks and reaches only 89.0% at 200; the profile-likelihood interval sits between 94.0% and 94.7% throughout. The gap is not a small-sample effect that lengthening the record removes — it narrows by 8.7 points for an eightfold longer record.

Two estimators of the extremal index against the value the process actually has, over 300 records of 4000 steps at the 0.98 threshold. The process is a max-autoregression whose extremal index is 1 − α exactly, which is what makes a comparison of two estimators a statement about them rather than about a simulation. The runs estimator declusters with a gap of 4 and sits below the truth at 4 of the 5 dependent settings — 0.2540 against 0.25 at the strongest. The intervals estimator needs no gap at all and sits above it at all 5, 0.2701 at the same setting. Under independence the runs estimator reports 0.9396 rather than one, because two independent exceedances inside 4 steps are counted as one cluster.

The runs estimator of the extremal index at 15 run lengths from 1 to 40, each divided by the index its process actually has, over 300 records of 4000 steps at the 0.98 quantile. On the max-autoregressions, whose clusters are runs of consecutive exceedances, the estimate drifts: 0.2573 at a run length of two and 0.2121 at 40, where the truth is 0.25. On the max-moving-maximum, whose cluster members are 6 steps apart, it falls off a cliff: 0.9069 at six and 0.3649 at seven, against 0.40. At a run length of 4 it reads 0.9424, 2.36 times the truth. Without dependence the estimate is 0.9407 at four and 0.4573 at 40, because independent exceedances also land close together.

Where it is used

7 essays draw this figure, each at the numbers its own argument is about, so the same picture answers 7 different questions.

All 80 figures

FieldsThreadsSeriesConceptsAll essaysSearch