The tail past the last observation

The maximum converges slowly

The rate at which a normalised maximum reaches its limit law is computable rather than simulable, because the exact law of a maximum is always available. For a normal parent the distance falls like one over the logarithm of the block and is still 0.0091 at a million readings; for an exponential parent, with the same limit, it is 2.707×10⁻⁷.

Worth reading first: Sums of almost anything · Three shapes, one limit.

A million readings in a block. The maximum of them, shifted and stretched by the constants the theory names, sits 0.0091 away from the Gumbel law it converges to, measured as the largest gap between the two distribution functions. An exponential parent, which has the same limit and belongs to the same domain of attraction, sits 2.707×10⁻⁷ away at the same block size. The two differ by a factor of 33,556, and the normal’s distance at a million readings is still larger than the exponential’s at a hundred.

Neither number is simulated, which is the difference between this and the demonstration of the limit theorem for sums that the same question was first asked of. The exact distribution of a maximum is always available — it is F(x)nF(x)^n — so the whole of this essay is the difference between two closed forms, evaluated. That is worth having in advance, because a rate argued from a fitted slope on six simulated points is an argument about the fit as much as about the rate, and there is nothing here for a fit to be wrong about. Every important number in this collection is computed twice; here both routes are exact, and the second is what the exponential parent is for.

The essay that priced the central limit theorem’s rate found that the middle of a distribution converges fast and its tail does not — a factor of twenty-seven between the relative error at the median and three standard deviations out, at the same sample size. This is the same question asked of the limit theorem for the largest value rather than for the average, and the answer is worse in a way that has a name: membership of a domain of attraction says nothing whatever about speed.

Two parents, one limit, and no simulation anywhere

The normalising constants are not fitted. For a normal parent the theory names bn=Φ1(11/n)b_n = \Phi^{-1}(1 - 1/n) and an=1/(nϕ(bn))a_n = 1/(n\,\phi(b_n)); for an exponential parent, bn=lognb_n = \log n and an=1a_n = 1. Both are computed. That matters more than it sounds: a distance measured to a fitted limit would be a statement about the fit, and a bad choice of constants can make any parent look slow.

At a thousand readings a block the normal’s constants come out at a=0.2970a = 0.2970 and b=3.0902b = 3.0902; at a million, a=0.2021a = 0.2021 and b=4.7534b = 4.7534. The location climbs, as it must, and the scale shrinks — the maximum of a million normal readings is a more concentrated quantity than the maximum of a thousand, in absolute terms, which is the whole reason a normalisation is needed at all.

The exact law is Φ(b+ax)n\Phi(b + ax)^n, and it is not computed that way. At n=106n = 10^6 the quantity being raised to the millionth power is 11061 - 10^{-6}, and every digit of the answer lives in the part a double-precision subtraction throws away. It is computed through the complementary tail instead:

P ⁣(Mnbax)=exp{nlog ⁣(1Q(b+ax))}P\!\left(\frac{M_n - b}{a} \le x\right) = \exp\left\{n \log\!\left(1 - Q(b + ax)\right)\right\}

with the logarithm of 1Q1 - Q taken directly from the small quantity QQ rather than by forming the difference and then taking a logarithm of it. The limit it is compared with is exp(ex)\exp(-e^{-x}). Getting this wrong does not throw; it returns a smooth curve that is one everywhere and a distance that shrinks beautifully, which is the failure mode this collection keeps finding at the point where a number is small.

The exponential’s exact law is easier still — the maximum of nn exponentials normalised by logn\log n has distribution (1ex/n)n(1 - e^{-x}/n)^n — so its distance to the limit is arithmetic in one line and needs no numerics at all. Having one parent whose answer can be checked by hand is what makes the other one’s answer usable.

The distance, at six block sizes

One tail arrives; the other is still on its way at a million. The Kolmogorov distance between the exact law of a normalised maximum and its Gumbel limit, at six block sizes, for two parents that both have that same limit. Both are closed form: the exact law of a maximum is F(x)^n and no simulation is involved. The exponential parent's distance falls from 0.0280 to 2.707e-7 — a factor of a hundred thousand, which is exactly one over n. The normal parent's falls from 0.0522 only to 0.0091, a factor of 5.74, because its rate is one over log n. At a million readings a block the two differ by a factor of 33556.3.
Fig. 1 The largest gap between the exact law of a normalised maximum and its Gumbel limit, at six block sizes, for two parents that share that limit. The exponential’s falls from 0.0280 to 2.707×10⁻⁷; the normal’s falls only from 0.0522 to 0.0091.

The normal parent reads 0.0522, 0.0272, 0.0182, 0.0137, 0.0109, 0.0091 at ten through a million readings a block. Five orders of magnitude buy a factor of 5.74.

The exponential parent reads 0.0280 at ten and 2.707×10⁻⁷ at a million: a factor of a hundred thousand, which is exactly the factor by which the block grew. One of these is a limit theorem behaving the way the phrase suggests. The other is a limit theorem that is technically true and practically inert at every block size a record will ever have.

It is worth stating what the normal’s number means for a reader rather than for a theorem. A block of a million readings is more than two thousand years of daily observations. At that block size, the law an analysis fits is still wrong about the probability of some level by nearly one part in a hundred — not as a relative error on a tiny probability, which would be forgivable, but as an absolute gap between two distribution functions both of which are near 0.9 there.

Where the gap actually is

A largest gap is a supremum, and a supremum can sit somewhere nobody reads. So the place matters as much as the size.

The exact law of a maximum, and the limit it is fitted as. The exact distribution of the maximum of 1000 normal readings, normalised by the textbook constants a = 0.2970 and b = 3.0902, drawn against the Gumbel law it converges to. The exact curve is Φ(x)^n and needs no simulation. The largest gap between them is 0.0182, at x = 2.088, where the exact law puts 0.9016 of its probability below and the limit puts 0.8834. That gap closes like one over the logarithm of the block: at a million readings a block it is still 0.0091, which the exponential parent passes before its blocks reach a hundred.
Fig. 2 The exact law of the maximum of a thousand normal readings, normalised, against the Gumbel law an analysis fits it as. The largest gap is 0.0182, at x = 2.088, where the exact law puts 0.9016 of its probability below and the limit puts 0.8834.

At a thousand readings a block the gap is largest at x=2.088x = 2.088, where the exact law puts 0.9016 of its probability below and the limit puts 0.8834. That is not an obscure corner of the curve. It is the upper shoulder — the point below which nine block maxima in ten fall, which is the ten-block return level, which is one of the two or three numbers anybody fits an extreme-value law in order to read.

So the answer to “does the distance sit where it matters” is that it sits almost exactly where it matters, and this is not a coincidence about the normal. The lower tail of a maximum’s distribution is squeezed against zero by the power nn and the upper tail is squeezed against one; the only place two such curves have room to differ is the shoulder, and the shoulder is the part an extrapolation is anchored to.

The exact law of a maximum, and the limit it is fitted as. The exact distribution of the maximum of 1000000 normal readings, normalised by the textbook constants a = 0.2021 and b = 4.7534, drawn against the Gumbel law it converges to. The exact curve is Φ(x)^n and needs no simulation. The largest gap between them is 0.0091, at x = 2.159, where the exact law puts 0.9001 of its probability below and the limit puts 0.8910. That gap closes like one over the logarithm of the block: at a million readings a block it is still 0.0091, which the exponential parent passes before its blocks reach a hundred.
Fig. 3 The same two curves at a million readings a block. The gap has halved, to 0.0091 at x = 2.159, and it has not moved: the exact law puts 0.9001 of its probability below and the limit puts 0.8910.

A thousandfold larger block halves the gap and leaves it in the same place. At a million the supremum is at x=2.159x = 2.159, with 0.9001 against 0.8910 — the same shoulder, the same reading, half the error. Three orders of magnitude of extra data for a factor of two.

Which rate, read two ways

A distance falling like 1/logn1/\log n and one falling like 1/n1/n look similar on a table of six numbers and nothing alike under a fit, so each profile is regressed twice and the products are printed beside the slopes.

Two rates, and the axis each one is straight in. The slope of the logarithm of the Kolmogorov distance against two axes, for two light-tailed parents, over block sizes from 10 to 1000000. A slope of minus one names the rate. The normal's distance is straight against the logarithm of the logarithm — slope -0.9767 — and nothing like straight against the logarithm itself, at -0.1460; the exponential's is the other way round, -1.0022 and -6.3078. So the normal's maximum converges at one over log n and the exponential's at one over n. Read as products rather than slopes the same fact is arithmetic: the normal's distance times log n sits between 0.1252 and 0.1260 across five orders of magnitude.
Fig. 4 The slope of the logarithm of the distance against two axes, for both parents. A slope of minus one names the rate: the normal is straight against log log n at −0.9767 and nothing like straight against log n at −0.1460, and the exponential is the other way round.

The normal’s log-distance against loglogn\log \log n has slope −0.9767; against logn\log n it is −0.1460. The exponential’s against logn\log n is −1.0022; against loglogn\log \log n it is −6.3078. Each parent is straight in exactly one of the two axes and hopeless in the other, which is what a rate looks like when it is real rather than chosen.

The products say the same thing as arithmetic rather than as a fit, which is the reading to trust. The normal’s distance times logn\log n runs 0.12521, 0.12604, 0.12601, 0.12577, 0.12548 from a hundred readings a block up to a million — constant to within a quarter of a per cent across four orders of magnitude, with no constant fitted to make it so. The exponential’s distance times nn runs 0.27158, 0.27076, 0.27068, 0.27067, 0.27067, which stops moving in the fifth decimal place. Two rates, each identified by a product that does not trend, and neither identification depends on a regression.

The one place either product misbehaves is the smallest block. The normal’s is 0.12011 at ten readings and the exponential’s is 0.28000, both a little off their plateau, which is exactly what a leading-order rate should do at n=10n = 10 and is reported rather than dropped.

What a slow rate does to an estimate

None of this would matter if the difference between the exact law and its limit stayed inside the arithmetic. It does not: it comes out as a bias in the one number an extreme-value analysis is a claim about.

The bias is a fact about the block, not about the record. The bias in the estimated shape against the size of the block it is estimated from, at a fixed 200 blocks and 300 records apiece. The light-tailed parent's bias falls from -0.1666 at 10 readings a block to -0.0458 at 100000 — a factor of 3.64 for four orders of magnitude, which is the one-over-log-n rate showing up in an estimate rather than in a distance. The heavy-tailed parent's blocks are exactly Fréchet at every size, and its bias never exceeds 0.0258. Adding blocks does not touch either: the same light-tailed bias is -0.1133 at fifty blocks and -0.1036 at five hundred.
Fig. 5 The bias in the estimated shape against the number of readings in each block, at a fixed two hundred blocks. The light-tailed parent’s falls from −0.1666 to −0.0458 over four orders of magnitude; the heavy-tailed parent’s, whose blocks are exactly Fréchet at every size, never exceeds 0.026.

Fit a generalised extreme value law to two hundred block maxima from a normal parent and the shape comes out at −0.1666 when the blocks hold ten readings, and at −0.0458 when they hold a hundred thousand. The truth is zero at every one of those block sizes. What is being estimated is not the shape of the limit law; it is the shape of the exact law of the block that was actually formed, and that law is short-tailed relative to Gumbel by an amount falling like one over the logarithm of the block.

The control is the heavy-tailed parent in the same sweep. A Pareto’s block maxima are Fréchet at every block size, exactly, with nothing to converge to — and its bias over the same four orders of magnitude reads 0.0258, 0.0058, 0.0079, −0.0017, 0.0030, which is noise around zero with no trend in it. So the light tail’s bias is not the fitting code and not the block construction; it is the rate, showing up in an estimate instead of in a distance.

The shape estimate narrows, and one of them narrows onto the wrong number. The middle ninety per cent of the estimated shape parameter, over 400 records apiece, for three parents whose limits are a Fréchet, a Gumbel and a Weibull. Every band narrows as the record lengthens — the heavy-tailed parent's from 0.9734 wide at 20 blocks to 0.1375 at 500 — and two of the three narrow onto the shape their parent actually has. The light-tailed parent's narrows onto -0.1036 rather than onto zero, because a maximum of 100 normal readings is not yet at its limit and the fit reads the shape of what it was given.
Fig. 6 The middle ninety per cent of the estimated shape at five record lengths. Two bands narrow onto the shape their parent has; the light-tailed parent’s narrows onto −0.1036, and its spread falls from 0.2295 to 0.0289 while doing it.

And what happens next is the reason the rate is worth pricing at all. Adding blocks does nothing to a bias that is about the block, but it shrinks the interval around it: the light-tailed parent’s spread falls from 0.2295 at twenty blocks to 0.0289 at five hundred, a factor of eight, while its estimate sits between −0.15 and −0.10 throughout. The essay that counts how often the family is named correctly reads that as a call that gets worse with more data, from 83.0% to 4.8%. This is the same fact seen from underneath.

What block size would be enough

A rate stated as a rate is easy to nod at. Stated as the block it would take to reach a given accuracy, it is not.

Take the product as the constant it appears to be. The normal’s distance times the logarithm of the block sits at 0.12548 at a million readings, so a block reaching a distance of a thousandth needs logn125\log n \approx 125 — which is nn of about 1054.510^{54.5} readings. The exponential’s product is 0.27067, so the same accuracy needs 271 readings. Two hundred and seventy-one, against a number with fifty-five digits in it, for two parents that a textbook describes with the same sentence.

That is not a rhetorical flourish; it is the honest consequence of a logarithm. There is no block size available to anybody, in any subject, at which a normal parent’s maximum is close to its Gumbel limit in the sense the limit theorem promises. What is available is a block at which the discrepancy is small enough not to dominate the other errors in the analysis, and that is a different and much weaker claim — one that has to be checked against the sampling error of the fit rather than asserted.

At two hundred blocks and a hundred readings each, the shape’s bias is −0.1071 and the spread of the estimate across records is around 0.05, so the bias is twice the noise. At a hundred thousand readings a block the bias is −0.0458 and the spread is unchanged, because spread is about the number of blocks and bias is about their size. The two are independent dials, and only one of them is usually turned. The essay that counts what happens when the noise dial is turned and the bias dial is not is the arithmetic consequence.

What could have produced this without the claim being true

The constants could be badly chosen. They are the textbook ones, and a different admissible choice of ana_n and bnb_n changes the distance. This is the doubt that is not fully closed here, and it has a name in the literature: a penultimate approximation replaces the Gumbel limit with a generalised extreme value law of small negative shape and converges much faster. That is a statement in the same direction as everything measured here — the exact law is slightly short-tailed relative to Gumbel — and it would reduce the numbers without changing their sign. What it would not change is the estimate: a fit is not handed better constants, it is handed a sample.

The supremum could be an artefact of the grid. It is taken on a three-thousand-point grid and then refined by a golden-section search around the grid’s argument, so the reported number is the distance rather than the largest of three thousand samples of it. The figure at each block size independently re-derives the two curves and asserts that they are furthest apart within a tenth of a per cent of where the search put them.

The Kolmogorov distance could be the wrong metric. It weights the middle of the distribution, where both curves have mass, and an extreme-value analysis is read further out than that. This one cuts against the essay: a metric that weighted the far upper tail would make the normal parent look worse, not better, because that is where a power of a distribution function and a double exponential differ most in relative terms. The distance is reported because it is the conservative choice and because it is bounded, which lets two parents share an axis.

Six points could be too few to fit a slope to. They are — a slope fitted to six points is a weak instrument, and that is exactly why the products are printed. A product that stops moving in the fifth decimal place over four orders of magnitude is not a fit at all; it is the rate written as arithmetic, and the slopes are reported beside it only so that the reading cannot be a choice of axis. This is the same preference for a quantity that can be checked without a model that runs through the argument for enumerating a reference distribution rather than sampling it.

The exponential could be a special case chosen to flatter the comparison. It is, and that is the point of it. It was picked because its exact law is elementary and its rate is known, so it is the parent that says the machinery can detect a fast rate when there is one. A comparison in which every parent came out slow would be evidence about the method rather than about the parents.

Where this does not hold

The claim is about two parents in the Gumbel domain and it is not a claim about extreme-value theory in general. A heavy-tailed parent is a different situation entirely: a Pareto’s normalised maximum is exactly Fréchet at every block size, so its distance to its limit is zero everywhere and the whole question is empty. Every number here is about the awkward middle case, which happens to be the one almost every measured quantity is assumed to belong to.

Nor does it say what to do. A rate of one over the logarithm of the block is a statement that no achievable block size removes the discrepancy, which forecloses the obvious remedy and proposes nothing in its place. The two things that would — a penultimate law, or a bias correction derived rather than measured — are both derivations, and neither is here.

The one operational reading is negative and worth stating plainly. Membership of a domain of attraction is the qualification an extreme-value analysis checks, when it checks anything, and it is a statement about a limit and not about a distance. Two parents with the same limit differ by a factor of thirty-three thousand at the same block size. “Converges to the Gumbel law” is therefore true of both and useful about one, and the quantity that separates them is not in the statement at all.

That is a familiar shape here. A fitted autoregression reproduces its sample exactly at the lags it was fitted on and everything it says past them is the model talking; the qualification that licenses the extrapolation is about the fit’s behaviour inside the data, and the reading is taken outside it. A maximum’s limit law is the same arrangement with the boundary drawn in nn rather than in lag, and the essay that reads a level with no data in it is what happens when the two are combined.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Block maximaBlock sizeCentral limit theoremClosed formConvergence rateDomain of attractionExtreme value theoryGeneralised extreme valueGumbel lawKolmogorov distanceNormalising constantsOrder statisticReturn levelShape parameterTail probability