The maximum converges slowly
Worth reading first: Sums of almost anything · Three shapes, one limit.
A million readings in a block. The maximum of them, shifted and stretched by the constants the theory names, sits 0.0091 away from the Gumbel law it converges to, measured as the largest gap between the two distribution functions. An exponential parent, which has the same limit and belongs to the same domain of attraction, sits 2.707×10⁻⁷ away at the same block size. The two differ by a factor of 33,556, and the normal’s distance at a million readings is still larger than the exponential’s at a hundred.
Neither number is simulated, which is the difference between this and the demonstration of the limit theorem for sums that the same question was first asked of. The exact distribution of a maximum is always available — it is — so the whole of this essay is the difference between two closed forms, evaluated. That is worth having in advance, because a rate argued from a fitted slope on six simulated points is an argument about the fit as much as about the rate, and there is nothing here for a fit to be wrong about. Every important number in this collection is computed twice; here both routes are exact, and the second is what the exponential parent is for.
The essay that priced the central limit theorem’s rate found that the middle of a distribution converges fast and its tail does not — a factor of twenty-seven between the relative error at the median and three standard deviations out, at the same sample size. This is the same question asked of the limit theorem for the largest value rather than for the average, and the answer is worse in a way that has a name: membership of a domain of attraction says nothing whatever about speed.
Two parents, one limit, and no simulation anywhere
The normalising constants are not fitted. For a normal parent the theory names and ; for an exponential parent, and . Both are computed. That matters more than it sounds: a distance measured to a fitted limit would be a statement about the fit, and a bad choice of constants can make any parent look slow.
At a thousand readings a block the normal’s constants come out at and ; at a million, and . The location climbs, as it must, and the scale shrinks — the maximum of a million normal readings is a more concentrated quantity than the maximum of a thousand, in absolute terms, which is the whole reason a normalisation is needed at all.
The exact law is , and it is not computed that way. At the quantity being raised to the millionth power is , and every digit of the answer lives in the part a double-precision subtraction throws away. It is computed through the complementary tail instead:
with the logarithm of taken directly from the small quantity rather than by forming the difference and then taking a logarithm of it. The limit it is compared with is . Getting this wrong does not throw; it returns a smooth curve that is one everywhere and a distance that shrinks beautifully, which is the failure mode this collection keeps finding at the point where a number is small.
The exponential’s exact law is easier still — the maximum of exponentials normalised by has distribution — so its distance to the limit is arithmetic in one line and needs no numerics at all. Having one parent whose answer can be checked by hand is what makes the other one’s answer usable.
The distance, at six block sizes
The normal parent reads 0.0522, 0.0272, 0.0182, 0.0137, 0.0109, 0.0091 at ten through a million readings a block. Five orders of magnitude buy a factor of 5.74.
The exponential parent reads 0.0280 at ten and 2.707×10⁻⁷ at a million: a factor of a hundred thousand, which is exactly the factor by which the block grew. One of these is a limit theorem behaving the way the phrase suggests. The other is a limit theorem that is technically true and practically inert at every block size a record will ever have.
It is worth stating what the normal’s number means for a reader rather than for a theorem. A block of a million readings is more than two thousand years of daily observations. At that block size, the law an analysis fits is still wrong about the probability of some level by nearly one part in a hundred — not as a relative error on a tiny probability, which would be forgivable, but as an absolute gap between two distribution functions both of which are near 0.9 there.
Where the gap actually is
A largest gap is a supremum, and a supremum can sit somewhere nobody reads. So the place matters as much as the size.
At a thousand readings a block the gap is largest at , where the exact law puts 0.9016 of its probability below and the limit puts 0.8834. That is not an obscure corner of the curve. It is the upper shoulder — the point below which nine block maxima in ten fall, which is the ten-block return level, which is one of the two or three numbers anybody fits an extreme-value law in order to read.
So the answer to “does the distance sit where it matters” is that it sits almost exactly where it matters, and this is not a coincidence about the normal. The lower tail of a maximum’s distribution is squeezed against zero by the power and the upper tail is squeezed against one; the only place two such curves have room to differ is the shoulder, and the shoulder is the part an extrapolation is anchored to.
A thousandfold larger block halves the gap and leaves it in the same place. At a million the supremum is at , with 0.9001 against 0.8910 — the same shoulder, the same reading, half the error. Three orders of magnitude of extra data for a factor of two.
Which rate, read two ways
A distance falling like and one falling like look similar on a table of six numbers and nothing alike under a fit, so each profile is regressed twice and the products are printed beside the slopes.
The normal’s log-distance against has slope −0.9767; against it is −0.1460. The exponential’s against is −1.0022; against it is −6.3078. Each parent is straight in exactly one of the two axes and hopeless in the other, which is what a rate looks like when it is real rather than chosen.
The products say the same thing as arithmetic rather than as a fit, which is the reading to trust. The normal’s distance times runs 0.12521, 0.12604, 0.12601, 0.12577, 0.12548 from a hundred readings a block up to a million — constant to within a quarter of a per cent across four orders of magnitude, with no constant fitted to make it so. The exponential’s distance times runs 0.27158, 0.27076, 0.27068, 0.27067, 0.27067, which stops moving in the fifth decimal place. Two rates, each identified by a product that does not trend, and neither identification depends on a regression.
The one place either product misbehaves is the smallest block. The normal’s is 0.12011 at ten readings and the exponential’s is 0.28000, both a little off their plateau, which is exactly what a leading-order rate should do at and is reported rather than dropped.
What a slow rate does to an estimate
None of this would matter if the difference between the exact law and its limit stayed inside the arithmetic. It does not: it comes out as a bias in the one number an extreme-value analysis is a claim about.
Fit a generalised extreme value law to two hundred block maxima from a normal parent and the shape comes out at −0.1666 when the blocks hold ten readings, and at −0.0458 when they hold a hundred thousand. The truth is zero at every one of those block sizes. What is being estimated is not the shape of the limit law; it is the shape of the exact law of the block that was actually formed, and that law is short-tailed relative to Gumbel by an amount falling like one over the logarithm of the block.
The control is the heavy-tailed parent in the same sweep. A Pareto’s block maxima are Fréchet at every block size, exactly, with nothing to converge to — and its bias over the same four orders of magnitude reads 0.0258, 0.0058, 0.0079, −0.0017, 0.0030, which is noise around zero with no trend in it. So the light tail’s bias is not the fitting code and not the block construction; it is the rate, showing up in an estimate instead of in a distance.
And what happens next is the reason the rate is worth pricing at all. Adding blocks does nothing to a bias that is about the block, but it shrinks the interval around it: the light-tailed parent’s spread falls from 0.2295 at twenty blocks to 0.0289 at five hundred, a factor of eight, while its estimate sits between −0.15 and −0.10 throughout. The essay that counts how often the family is named correctly reads that as a call that gets worse with more data, from 83.0% to 4.8%. This is the same fact seen from underneath.
What block size would be enough
A rate stated as a rate is easy to nod at. Stated as the block it would take to reach a given accuracy, it is not.
Take the product as the constant it appears to be. The normal’s distance times the logarithm of the block sits at 0.12548 at a million readings, so a block reaching a distance of a thousandth needs — which is of about readings. The exponential’s product is 0.27067, so the same accuracy needs 271 readings. Two hundred and seventy-one, against a number with fifty-five digits in it, for two parents that a textbook describes with the same sentence.
That is not a rhetorical flourish; it is the honest consequence of a logarithm. There is no block size available to anybody, in any subject, at which a normal parent’s maximum is close to its Gumbel limit in the sense the limit theorem promises. What is available is a block at which the discrepancy is small enough not to dominate the other errors in the analysis, and that is a different and much weaker claim — one that has to be checked against the sampling error of the fit rather than asserted.
At two hundred blocks and a hundred readings each, the shape’s bias is −0.1071 and the spread of the estimate across records is around 0.05, so the bias is twice the noise. At a hundred thousand readings a block the bias is −0.0458 and the spread is unchanged, because spread is about the number of blocks and bias is about their size. The two are independent dials, and only one of them is usually turned. The essay that counts what happens when the noise dial is turned and the bias dial is not is the arithmetic consequence.
What could have produced this without the claim being true
The constants could be badly chosen. They are the textbook ones, and a different admissible choice of and changes the distance. This is the doubt that is not fully closed here, and it has a name in the literature: a penultimate approximation replaces the Gumbel limit with a generalised extreme value law of small negative shape and converges much faster. That is a statement in the same direction as everything measured here — the exact law is slightly short-tailed relative to Gumbel — and it would reduce the numbers without changing their sign. What it would not change is the estimate: a fit is not handed better constants, it is handed a sample.
The supremum could be an artefact of the grid. It is taken on a three-thousand-point grid and then refined by a golden-section search around the grid’s argument, so the reported number is the distance rather than the largest of three thousand samples of it. The figure at each block size independently re-derives the two curves and asserts that they are furthest apart within a tenth of a per cent of where the search put them.
The Kolmogorov distance could be the wrong metric. It weights the middle of the distribution, where both curves have mass, and an extreme-value analysis is read further out than that. This one cuts against the essay: a metric that weighted the far upper tail would make the normal parent look worse, not better, because that is where a power of a distribution function and a double exponential differ most in relative terms. The distance is reported because it is the conservative choice and because it is bounded, which lets two parents share an axis.
Six points could be too few to fit a slope to. They are — a slope fitted to six points is a weak instrument, and that is exactly why the products are printed. A product that stops moving in the fifth decimal place over four orders of magnitude is not a fit at all; it is the rate written as arithmetic, and the slopes are reported beside it only so that the reading cannot be a choice of axis. This is the same preference for a quantity that can be checked without a model that runs through the argument for enumerating a reference distribution rather than sampling it.
The exponential could be a special case chosen to flatter the comparison. It is, and that is the point of it. It was picked because its exact law is elementary and its rate is known, so it is the parent that says the machinery can detect a fast rate when there is one. A comparison in which every parent came out slow would be evidence about the method rather than about the parents.
Where this does not hold
The claim is about two parents in the Gumbel domain and it is not a claim about extreme-value theory in general. A heavy-tailed parent is a different situation entirely: a Pareto’s normalised maximum is exactly Fréchet at every block size, so its distance to its limit is zero everywhere and the whole question is empty. Every number here is about the awkward middle case, which happens to be the one almost every measured quantity is assumed to belong to.
Nor does it say what to do. A rate of one over the logarithm of the block is a statement that no achievable block size removes the discrepancy, which forecloses the obvious remedy and proposes nothing in its place. The two things that would — a penultimate law, or a bias correction derived rather than measured — are both derivations, and neither is here.
The one operational reading is negative and worth stating plainly. Membership of a domain of attraction is the qualification an extreme-value analysis checks, when it checks anything, and it is a statement about a limit and not about a distance. Two parents with the same limit differ by a factor of thirty-three thousand at the same block size. “Converges to the Gumbel law” is therefore true of both and useful about one, and the quantity that separates them is not in the statement at all.
That is a familiar shape here. A fitted autoregression reproduces its sample exactly at the lags it was fitted on and everything it says past them is the model talking; the qualification that licenses the extrapolation is about the fit’s behaviour inside the data, and the reading is taken outside it. A maximum’s limit law is the same arrangement with the boundary drawn in rather than in lag, and the essay that reads a level with no data in it is what happens when the two are combined.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- Two intervals for one return level — both name block maxima, closed form, generalised extreme value, return level, shape parameter
- A bound written for a coin — both name central limit theorem, convergence rate, tail probability
- A correction that goes below zero — both name central limit theorem, convergence rate, tail probability
- The clustering the tail has — both name closed form, return level, shape parameter
- The run length a declustering chooses — both name closed form, return level, shape parameter
- Where the derivative is zero — both name central limit theorem, closed form, convergence rate
Named objects
A flat tag is an object no other essay names yet.
Block maximaBlock sizeCentral limit theoremClosed formConvergence rateDomain of attractionExtreme value theoryGeneralised extreme valueGumbel lawKolmogorov distanceNormalising constantsOrder statisticReturn levelShape parameterTail probability