Concept

Mean squared error — where it appears

The average squared distance between an estimate or forecast and the truth, which is a variance plus the square of a bias. Splitting it into those two parts is what says whether a procedure is wrong on average or merely noisy, and the repairs are different.

Named by 37 essays across 17 fields — each of them below, with the objects they name alongside it.

What each rule gives up against an oracle that is arithmetic. Expected squared error of the candidate each rule selects, minus the expected squared error of the best candidate in the table, over 500 draws of 120 rows. Both quantities are closed forms — σ_S²(1 + q/(n − q − 1)) — so the only Monte Carlo here is over which candidate got picked. The hold-out spends half its rows measuring what the criterion computes, and pays 1.8 times as much for it. Schwarz's criterion is worst because it is answering a different question: which candidate contains the truth, rather than which one forecasts best.

A criterion is a prediction of the hold-out

A rolling hold-out spends half the sample measuring what a criterion computes from all of it. Against an oracle that is arithmetic rather than an estimate, the criterion gives up 0.01701 and the hold-out 0.03200 — and the number the hold-out reports for its own winner is optimistic by more than either.

proxy · Order-selection
The correction is not a property of the sample. tr(HΩ)/q for each of fifteen candidates, at ρ = 0.7. Two candidates that fit the same number of coefficients need corrections that differ by as much as 1.49, because one of them is fitting the persistent predictors and the other is not — so no single number can be right for both, and the scalar n/n_eff = 5.537 is above every one of them. The four predictors carry persistences 0.9, 0.6, 0.3, 0; at one persistence for every column the whole spread collapses and a scalar looks exactly as good as the trace.

A penalty is a trace

Akaike's 2q is not a count of coefficients. It is the answer a trace collapses to when the rows are independent — and once they are not, the trace is still the right object and is no longer the count.

effective · Order-selection
What each variant loses before anything has been searched for. The mean loss differential of each of the eight variants against the benchmark, over 600 tables of 60 origins, with every fit given 71 rows. The series is an AR(1) and every variant adds a lag whose coefficient is zero, so in population the two forecasts are the same forecast and the difference drawn here is estimation noise and nothing else. The marked line is σ²(q₁ − q₀)/n = -0.01408, which is an expression in how many coefficients each model has and how many rows it was fitted on — it knows nothing about the series, the persistence or which lag the variant added, and every bar is within a fifth of it. This is the amount a reference distribution recentred at each column's own sample mean believes the candidates are already behind by.

A table of nested models

A benchmark and eight variants of it, each adding one thing. Every variant is behind before the search begins, by an amount that can be written down before the data exists — and the two most natural ways of reading the table are wrong in opposite directions.

search · Forecast
What a sample shows, and what the algebra does. The difference between a rectangular block's implied long-run variance and a trapezoidal one's, as a share of the truth. Above the axis the rectangle is less biased and below it the trapezoid is. The heavy line is exact — computed from the law's own autocovariances — and it crosses at 19.2. The others are what samples of 120, 240, 480, 960 rows report, and every one of them exaggerates whichever window is ahead: at ℓ = 20, where the exact difference is 0.28 points, a sample of 120 rows shows 4.31 points — 15 times larger. That is the number the earlier reading of this comparison was missing: three tenths of a point is what the algebra says and not what a hundred and twenty rows report.

The gap a sample shows

The exact difference between two block windows at a block length of twenty is three tenths of a point. What a hundred and twenty rows report is four and a third, because the autocovariances the window is applied to are attenuated too.

crossing · Bootstrap
Three rules and a target none of them is aimed at. Which block length each rule picks, over 400 samples of 120 rows, for the tapered window. Two of the rules are points: a length written into a protocol is 8.00 on every draw and the rule of thumb is 4.00, because n to the one third does not read the data at all. The plug-in reads the sample's own persistence and lands at 14.36 with a standard deviation of 2.93. The length that would actually have been best on that draw averages 24.57 with a standard deviation of 16.23 and runs from 10 to 48 between its tenth and ninetieth percentiles. The target moves five times as much as the best estimate of it does, which is why no rule can be close to it and why the two that do not try are not merely worse — they are somewhere else.

The length nobody has

Every comparison of block windows in this collection is made at each window's own best block length. That length has a standard deviation of sixteen across draws and averages twenty-five. No rule is aimed at it.

feasible · Bootstrap
One forecast, and the band the arithmetic puts round it. An AR(1) with φ = 0.75, 60 observations, fitted by least squares and forecast 14 steps ahead. The point forecast decays towards the fitted mean at φ̂^h; the band is ±1.96 standard errors from σ̂²Σψ̂², which grows with the horizon and stops at the unconditional spread 1.72. The dashed pair is the same band computed at the true parameters, which nobody has. The marks past zero are what actually arrived: 12 of 14 inside the band this once, which is one draw and settles nothing.

What the model says next

The usual account of a time series stops at estimation. A forecast asks the other question — not what the parameter is but what the next observation will be — and the band round it is a closed form that grows with the horizon and then stops growing, at a value the series was going to reach anyway.

forecast · Forecast
Three quantities, and only one of them crosses zero. Two forecasts of an AR(1) — the last value carried forward and the mean of the last 60 observations — at 1 step ahead. The curve through zero is σ₁² − σ₂², the difference in expected squared error that a comparison of accuracy tests; it changes sign at φ = 0.4922. The two curves above it are σ₁² − σ₁₂ and σ₂² − σ₁₂, the quantities the two encompassing tests are about, and neither of them comes near zero anywhere: the smallest value either takes across the range is 0.008 times the variance of the series. All three are closed forms in φ, R and h with no simulation in them. Equal accuracy is one hypothesis about this picture and encompassing is another, and a set of numbers can satisfy either without the other.

What the other forecast adds

Two forecasters, one series, and two different questions about them. Which is more accurate has an answer that changes with the persistence of the series; whether either is redundant has an answer that never changes at all.

ranking · Forecast
One true null, one table, five readings. every subset of four, fifteen models, at a null where nothing any candidate holds is worth anything, over 500 draws. Each bar is the share of draws on which that reading declares a difference at a nominal 5%. The reading is the whole of the difference between the bars: the data is identical. An open search over all 210 ordered pairs rejects 76.2%; the table's own 5% point is 3.163 against the 1.671 a single comparison uses. Bonferroni takes the open reading to 0.6% — and on the nested ladder the same correction does not reach the nominal level at all, because there the excess is a shift in the mean rather than a maximum over many.

When the benchmark is a candidate

A specification search with a benchmark nailed down is the case with a closed form. Take the nail out — let the model that would have been reported be one of sixteen, chosen by the same data as its rivals — and the same true null is read three ways, at 2.0%, 7.8% and 76.2%.

select · Forecast
One comparison, and the two error bars it can be given. 60 rolling origins, a window of 60 observations, forecasts 4 steps ahead, at the persistence φ = 0.8256 where the two benchmarks have exactly equal population mean squared error. Each mark is one origin's difference in squared error; the horizontal line is their mean, 0.6522. The two vertical bars at the right are ±1.96 standard errors round that mean computed two ways — 0.5337 treating the differences as independent, 0.6880 allowing for the overlap between neighbouring forecasts. The null is true here by construction, so an interval that excludes zero is a mistake, and the narrow one does it far more often than the wide one.

Which forecast is better

Two forecasters, one series, and a difference in mean squared error. Whether that difference is real is a hypothesis test, its terms are not independent, and the standard error it needs is not the one a t-test computes.

evaluation · Forecast
Which groups partial pooling serves, standard error 1 population width. Pooling's expected squared error for a group, divided by its own mean's, against how far the group truly sits from the centre. It is ×0.25 at the centre and crosses ×1 at 1.732 population widths, beyond which 8.33% of a normal population lies; capping the shift at one standard error holds every group under ×2.

A group from the population's own tail

Partial pooling halves the total squared error when a group's own standard error equals the spread between groups. Every group whose true effect sits more than 1.73 population widths from the centre — 8.33% of a perfectly normal population — does worse than it would have with its own mean, and its loss grows without bound. Among eight groups with the spread estimated, the most extreme is worse off in 61.6% of datasets. Capping the shift at one standard error keeps the total at 0.528 of the unpooled error and holds every group under twice it.

borrowed · Shrinkage
Which window is better depends on who chose the block length. The margin between a rectangular block and a tapered one, on 400 samples of 120 rows, under four rules for choosing the block length. At the length that would actually have been best on each draw the taper is ahead by 2.12 points of a 35.8% error, at 22.1 paired standard errors; at a length estimated from the sample's own persistence it is ahead by 1.89. At the length this field's own figures use — eight — the rectangle is ahead by 2.04, and at the rule of thumb by 4.54. Every rule sees the same draws. What separates them is the length: the two rules that lose to the rectangle pick 4.00 and 8.00 where the best available is 24.57, and a tapered window at a quarter of the right length has thrown away most of what it was weighting.

An ordering that depends on the rule

The tapered block beats the rectangular one at the best available block length and at one estimated from the data. At a length written into a protocol, and at the rule of thumb, the rectangle wins — at every sample size measured.

feasible · Bootstrap
What a longer block buys and what it costs. A trapezoidal block at 120 rows, with the error split into the two things it is made of. The bias falls with the block length, because a longer block attenuates less, and it flattens at 23.2% because the sample's own autocovariances are short whatever window is applied to them. The spread rises with it, because a longer block means fewer of them. Their sum in quadrature has a minimum at ℓ = 16, which is not where either of the two has one. The faint line is the rectangle's total error, for scale: it is above the trapezoid's from ℓ = 12 onwards.

Bias is not the whole of it

A window that reaches zero at its ends attenuates less and uses less of each block. The block length that minimises its bias is not the one that minimises its error, and comparing two windows at one length compares one of them mis-tuned.

crossing · Bootstrap
Every candidate is behind by what its parameter count says. Each dot is one of the fifteen subsets of four predictors, fitted on a rolling window of 80 rows and scored against the benchmark out of sample over 60 origins, at a null where every one of them contains the truth. The line is σ²(q₀/(R − q₀ − 1) − q/(R − q − 1)), which is arithmetic on two integers and a window length. Most of these pairs are not nested — a subset of two predictors and a different subset of two share neither model — and the closed form does not care: the displacement is a statement about how many coefficients each side estimates. The candidates of the benchmark's own dimension sit at zero.

The displacement is a parameter count

A nested variant is behind its benchmark out of sample before anything is searched for. The closed form for how far turns out to have nothing about nesting in it — only two integers and a window length — and it prices a table where no candidate contains any other.

select · Multiplicity
A test between nested models, under a null that is true. 1000 comparisons: an AR(1) truth, forecast by a fitted AR(1) and by a fitted AR4 whose extra coefficients are zero. In population the two forecasts are identical, so every rejection is false. The larger model's mean squared error is 1.1663 against 1.0583 — worse, by exactly the noise in estimating coefficients that are not there — and the ordinary test therefore declares the smaller model significantly better 67.2% of the time. Read one-sided in the direction anybody asks about, it finds the larger model better 0.0% of the time. Adding the squared difference between the two forecasts back into the loss differential puts the level at 4.9%.

When one model contains the other

The comparison a forecaster most often wants is between a model and the same model with one more term. That is exactly the comparison the standard test cannot make — and it fails by declaring the smaller model significantly better, more confidently the more data it is given.

evaluation · Forecast
The crossing is in the dependence, not in the split. Regret of each rule as the design and the errors are made persistent at the same coefficient, scored on fresh rows because the closed form assumes exactly what is being taken away. An optimism theorem counts rows; when the rows repeat each other there are fewer of them than there are rows, the penalty is too small for the fit it is correcting, and the criterion starts buying coefficients it should not — its average winner grows from 3.31 coefficients to 3.90. The hold-out never used the theorem and overtakes at ρ ≈ 0.81. Schwarz's criterion, worst of the three on independent rows, is best on repeating ones — its heavier penalty is right for the wrong reason.

Where the two searches cross

The obvious dial between a criterion and a hold-out is how much of the sample to hold out, and moving it never changes the answer. The dial that does is one nobody chooses — how much each row repeats the one before it — and the two rules change places at about 0.81.

proxy · Forecast
How the truth, the raw means, the posterior means and the constrained estimates spread, standard error 1. Beyond two population widths above the centre lie 2.28% of the true values, 7.86% of the raw means, 0.234% of the posterior means, and 2.28% of the constrained estimates.

Estimates that are too alike

Posterior means give each group its least-error estimate, and as a set they are too alike: with each group's standard error equal to the population's spread, they spread 0.707 as widely as the truth. Beyond two population widths lie 2.28% of the true effects, 7.86% of the groups' own means, and 0.234% of the posterior means — a tenth of the truth. Rescaling the estimates to the right spread counts the tail exactly and costs 17% more squared error; summing each group's posterior chance of being beyond the line counts it without changing any estimate.

borrowed · Shrinkage
Cut the charge and the width follows it. The band width each charge picks, averaged over 400 draws of 120 rows under AR(1) at 0.8, with the standard deviation across draws beside it. Schwarz's charge — half a log n a lag, which is 2.39 here — picks 3.67. Akaike's picks 6.02. Charging the numbers the window actually leaves free, which is half the width, picks 10.12; charging what the optimism measures, 0.767 of that, picks 14.15. A charge and the width it buys are very nearly reciprocal, which is what a likelihood rising at a fixed rate a lag implies and is why the four answers span a factor of 3.86. The width that was actually best on the draw averages 13.90 and moves by 10.30 from draw to draw — three times as much as any rule's answer does.

A width that moves and an error that does not

Four charges give four widths a factor of four apart and four errors half a per cent apart. The derived charge wins, significantly, by a quarter of what was on offer — and none of the four is an estimate of anything.

dimension · Criterion
Least squares estimates persistence low, by an amount with a formula. 3000 series of 50 observations at each persistence. The lower curve is the counted bias of the least-squares estimate of φ, and the open marks on it are −(1 + 3φ)/n, computed rather than fitted. The upper curve is the bias left after adding that quantity back, evaluated at the estimate rather than at the truth nobody has: -0.0020 at φ = 0.3, -0.0023 at φ = 0.5, -0.0039 at φ = 0.7, -0.0059 at φ = 0.8, -0.0108 at φ = 0.9, -0.0165 at φ = 0.95. The formula is a leading-order expression and it understates the bias where the persistence is nearest one — -0.0882 counted against -0.0770 predicted at φ = 0.95, which is the corner of the parameter space every one of these approximations is worst in.

Correcting the persistence

Least squares estimates how much a series remembers of itself as smaller than it is, at every value it can take, by an amount with a closed form. Subtracting that amount back is one line of arithmetic, and what the line costs is variance.

evaluation · Bias
The row count entered twice, and a penalty is one place. Regret against the best available model as the errors are made persistent. Counting rows more than quadruples; both penalty repairs — the trace, and the scalar effective sample size — are worse than it at every persistence measured; and whitening the sample and keeping the ordinary penalty falls, recovering 86.9% of what counting rows gives up at ρ = 0.85. Doing it at an estimated ρ recovers 80.6%, so having to estimate the dependence from the rows being selected on costs 7.3% of what knowing it is worth. Mallows' forms are drawn beside the logarithmic ones and behave the same, which is what rules the linearisation out.

The repair that was exact and made it worse

A penalty computed from the trace is exactly the optimism it estimates, and selecting with it gives up a fifth more than not correcting anything. The row count entered the criterion twice, and a penalty is the second place.

effective · Forecast
The ranking on the left, the weights on the right. Eight moving-average forecasts of an AR(1) at φ = 0.4895, the persistence at which the best of them exactly ties the 60-observation benchmark. On the left, each candidate's expected squared error in units of the series' own variance: the smallest belongs to L = 2, at 1.0156. On the right, the weight each carries in the variance-minimising combination of all eight — and the best of them carries 0.00000. The two ends of the family carry 1.0172 of the weight between them, and the combination they make is worth 0.7817, which is 23.0% below the best single forecast. Both columns are closed forms in φ. Which forecast to keep and which forecasts to use are different questions, and this is a set where the answers share nothing.

The weight that is a vector

Two forecasts have a best combination and one number describes it. Eight have a best combination too, and the vector describing it puts nothing at all on the forecast with the smallest mean squared error.

search · Rank
The argument is a third of the size of the thing it is inside. Three quantities on one scale, in points of the error in a block resample's implied long-run variance, at 120 rows. The gap between the two windows at the best available block length — the whole subject of the comparison this field inherited — is 2.12 points. What the best rule a practitioner could actually run gives up against that same best length is 7.26, a factor of 3.42. What the rule of thumb gives up is 26.01. So the ordering between windows is worth establishing and is not worth arguing about, and the sentence that follows from it is not use the taper but estimate the block length, because that is where the points are.

What choosing the length costs

The gap between two block windows at the best available length is 2.12 points. What the best rule a practitioner could run gives up against that same length is 7.26. The argument is a third of the size of the thing it is inside.

feasible · Bootstrap
Two structures in three are made worse. What adjusting for every covariate measured does to the bias in the treatment's estimated effect, against adjusting for none, over 4000 randomly drawn structures of 6 covariates each. Each covariate is independently a common cause with probability 0.25, a cause of the treatment only, a cause of the outcome only, a cause of neither, a step on the causal path, or a common effect. The rule leaves a larger bias on 65.5% of structures, a smaller one on 33.8%, and the same on 0.7%. The share is a property of that population of structures rather than of adjustment, which is why the weights are stated; what does not depend on them is that the rule has no direction — it is not a conservative default that occasionally overcorrects, it is a rule whose error is whatever the structure happens to be.

Adjusting for everything

"Control for every covariate that was measured" leaves a larger bias than controlling for nothing on 65.5% of four thousand randomly drawn structures and a smaller one on 33.8%. Its squared error is 4.110 times that of using no covariate at all, and half of it sits in its worst tenth of structures.

collider · Conditioning
The threshold buys accuracy and spends exceedances. The mean squared error of the estimated shape against the threshold, split into the square of its bias and its spread, over 600 records of 2000 readings from a a normal parent. At the 0.9 quantile 199 exceedances are left, the bias is -0.1708, the spread is 0.0701 and the total error is 0.0341. The bias falls as the threshold rises because the exceedances get closer to being generalised Pareto; the spread rises because there are fewer of them. The sum is smallest at the 0.925 quantile, at 0.0340, of which 80.6% is still bias — so even the best threshold on this grid is one where accuracy, not spread, is the binding constraint.

The threshold is a dial

A peaks-over-threshold analysis has one knob, and raising it buys accuracy with exceedances. For a normal parent the error is smallest at the 0.925 quantile and 80.6% of it is still bias there — and both diagnostics practitioners use to set the knob lose to a fixed 0.90 rule, one by a factor of 1.590 and one by 11.881.

extreme · Extremes
What the family-wise correction does to the effect it lets through. At two standard errors the estimate that clears an uncorrected 5% threshold averages 1.35 times the truth, and the one that clears the family-wise threshold averages 1.69 times it. The correction fixes the error rate by demanding a larger estimate, and a larger estimate is a more selected one.

The correction that makes the estimate worse

Correcting for twenty analyses repairs the p-value by demanding a larger statistic, and a larger statistic is a more selected one. At two standard errors the surviving estimate averages 1.35 times the truth before the correction and 1.69 times it after — so the honest error rate is bought with a more inflated effect.

paths · Forking
Where the taper's case begins, and it is not where the algebra says. The block length at which a trapezoidal block's implied variance stops being more biased than a rectangular one's, against the length of the sample. Computed exactly — from the law's own autocovariances, with no sampling in it — the answer is 19.2 and does not depend on the sample at all. What a sample of 120 rows reports is 13.3, and the reported crossing walks out towards the exact one as the sample grows: 13.3, 15.0, 16.4, 18.0. The mechanism is that the autocovariances the window is applied to are themselves attenuated, worst at the longest lags, and the window that discards those lags loses less of them.

The error no window repairs

Every block window's best estimate of a long-run variance is wrong by about forty per cent at a hundred and twenty rows, and the largest part of that is not a bias at all. Choosing the window moves a twentieth of it.

crossing · Bootstrap
Three estimates of the same effect, and the honest one is the worst. 8,000 trials in which every arm has an effect of exactly 0. The arm carried forward is reported by its first stage at +0.187 above the truth — it was chosen for being ahead — and by its second stage at -0.0032, which is unbiased by construction because the selection could not see it. The combination that gets published is at 0.123, exactly the share of the first stage's bias the first stage contributes. Root mean squared error: 0.240, 0.261, 0.182 — the unbiased estimate is the least accurate of the three.

The estimate after the choice

An arm chosen for being ahead is ahead by more than it should be, and the trial then publishes the average of the stage that chose it and the stage that did not. The unbiased estimate is the one built from a third of the data — and it is the least accurate of the three.

adaptive · Curse
The correction does not arrive at the truth, it passes it. The average decay factor a forecast applies to the last observation, at φ = 0.85 and 50 observations, 3000 series per horizon. The middle curve is φʰ, what the model actually does. Below it is the uncorrected forecast, which uses φ̂ʰ and reverts too fast — 24.8% short at h = 4, 30.0% short at h = 6, 32.7% short at h = 8. Above it is the forecast built on the corrected estimate, which overshoots, and the reason is arithmetic rather than a bad correction: raising an unbiased estimate to a power does not give an unbiased estimate of the power, and the higher the power the more the spread of φ̂ is converted into overshoot.

The repair that moves the wrong number

Correcting the bias in a persistence parameter is one line of arithmetic that works. Feeding the corrected estimate into a forecast repairs the number everybody looks at, makes the forecast worse by squared error at moderate persistence, and improves the interval for a reason that has nothing to do with bias.

evaluation · Bias
The scale moves the width; the curve does not. The band width each charge picks, averaged over 150 draws of 120 rows. The two conventions — a unit a lag and half a log n a lag — pick 6.08 and 3.65 lags. The four charges derived from the measured optimism pick 15.05, 15.60, 15.47 and 15.44, against a best width on the draw of 14.13. So the scale a charge is levied on moves the width by a factor of 4.27 and the shape of the charge moves it by 3.6%. None of the six is an estimate of the draw's own best width: the correlations are -0.006, -0.003, 0.017, -0.001, 0.016, -0.004.

What a better charge buys

Four charges derived from the same measurements pick band widths within six per cent of each other and deliver errors within two per cent of the gap any of them leaves. The scale a charge is levied on decides the width; the shape of the charge decides nothing.

curve · Criterion
Eight encompassing nulls, all of them true at once. The forecast under test is the variance-minimising combination of the eight candidates, and the first-order condition that defines its weights is cov(e_c, e_j) = var(e_c) for every j — so every one of the eight nulls is exactly true simultaneously and every rejection counted here is false. 800 draws of 60 origins. The largest of the eight statistics rejects 25.0% of the time at a nominal 5% and one comparison stated in advance rejects 3.1%, which makes the set worth about 9.1 independent comparisons — nearly the eight it has. Bonferroni, which is far inside its level on a search over accuracy comparisons of the same eight forecasters, is at 4.8% here.

When every null is true

A reality check assumes that every candidate in the set is exactly as good as the benchmark, which is a configuration nobody's data is ever in. Test a combination against its own parts and that configuration is not assumed — it is what the arithmetic makes true.

search · Multiplicity
The record stops here, and the curve does not. The level exceeded once in T blocks, against T, for a normal parent at 365 readings a block. The truth is closed form — the block maximum's own distribution function is Φ(x) raised to the 365, so the T-block level is Φ⁻¹ of the 365-th root of (1 − 1/T), with nothing fitted in it — and the fitted mean over 800 records of 50 blocks sits on top of it, 4.0186 against 4.0330 at 100 blocks. What moves is not the level but its error, which grows from 0.0956 at 10 blocks to 0.6383 at a thousand while the level itself moves only from 3.4421 to 4.5454. The rule marks the largest reading an average record contains, 4.0062: everything to the right of where it crosses is read from a fit rather than from data.

A level with no data in it

The largest of fifty block maxima is a 51-block event by its own plotting position, so a hundred-block level is read 1.96 times past the longest event the record contains — and it lands above the largest reading on 52.4% of records. The estimate stays nearly unbiased out there; what grows is its error, sixfold from ten blocks to a thousand.

extreme · Extremes
What each estimator costs, against how many groups there are. Each point is 4,000 datasets, with the population spread estimated from the data rather than supplied. Partial pooling first beats BOTH of the estimators it sits between at 5 groups; below that, complete pooling — which estimates nothing at all — is the better answer. Its own cost falls from 1.751 at 2 groups to 0.795 at 40.

The fewest groups that can borrow

At three groups the estimator that shrinks towards its own data's mean returns the group means untouched, on every dataset, because its constant is J − 3. At two it expands instead of shrinking. And the number of groups at which partial pooling starts to be worth doing is five, or two, or never — it depends on how far apart the groups are.

hierarchical · Pooling
The average decay factor each route produces, φ = 0.85, 6 steps ahead. The truth is φ^6 = 0.3771. no correction averages 0.2616 with a spread of 0.1646 and a squared forecast error of 3.2516; the formula, on the persistence averages 0.4213 with a spread of 0.2528 and a squared forecast error of 3.4827; the bootstrap, on the persistence averages 0.4355 with a spread of 0.2655 and a squared forecast error of 3.5120; the bootstrap, on the decay factor averages 0.3375 with a spread of 0.2278 and a squared forecast error of 3.4132. 800 series, 100 bootstrap refits each.

Correcting the forecast instead

The complaint against the usual repair is that a correction aimed at the persistence lands on the wrong quantity. Aiming it at the decay factor the forecast actually uses fixes exactly that — the error stops compounding with the horizon, 69.7% becomes 9.5% at twelve steps — and the forecast still gets worse.

evaluation · Bias
Reading the draw changes what is charged, not what is tracked. The correlation between the band width each rule picks and the best band width on the same draw, over 400 draws. The three fixed charges read -0.069, -0.066, 0.012. The three that read the sample read -0.012, -0.019, -0.041. None of the six is distinguishable from nothing. The statistic the first plug-in reads does vary — the draw's own summed squared autocorrelation runs from 2.06 to 10.19 with a mean of 4.36 — so the failure is not that the charge stopped moving. It is that what it moves with carries no information about which width this draw wanted.

A charge that reads the draw

Three charges built to read the sample track the best band width on their own draw at −0.012, −0.019 and −0.041, deliver more error than the fixed rule they are calibrated to, and pick a width half again as variable. The statistic moves; the answer does not.

curve · Criterion
What each group gains from being pooled, τ = 1. Eight groups whose sizes span a factor of 13.3. The smallest gains 2.076 of squared error, which is 36.8% of the total reduction; the largest gains 0.030, which is 0.5%. 2 of the eight account for half of everything pooling buys.

Where the borrowing goes

Pooling cuts the total squared error across eight groups by 56%. Two of the eight take 61% of that reduction, the four best-measured groups share 11% between them, and the largest group gets 1.5% of what the smallest does. The headline is a fact about the groups nobody was asking about.

hierarchical · Pooling
Five treatments of an estimate above one, φ = 0.95, n = 25. The correction exceeds one on 31.1% of series at this setting. left where it lands: squared forecast error 12.828, average decay factor 0.7974 against a true 0.7351; capped at 0.995: squared forecast error 5.680, average decay factor 0.5950 against a true 0.7351; capped at 1 − 1/n: squared forecast error 5.535, average decay factor 0.5256 against a true 0.7351; correction scaled to fit: squared forecast error 5.535, average decay factor 0.5256 against a true 0.7351; correction refused where it leaves: squared forecast error 5.868, average decay factor 0.4423 against a true 0.7351.

The correction that leaves the region

The bias correction adds (1 + 3φ̂)/n whatever φ̂ is, so it pushes the estimate above one whenever φ̂ exceeds (n − 1)/(n + 3) — on 31.1% of series at φ = 0.95 and twenty-five observations. Five obvious things to do about it differ by a factor of 2.3 in squared forecast error, and none of them is documented as a choice.

evaluation · Bias
A run length of 4 makes every exceedance its own cluster. 200 steps of a max-moving-maximum, X(t) = max(0.4·Z(t), 0.3·Z(t−6), 0.3·Z(t−12)) with unit Fréchet innovations Z, whose extremal index is exactly 0.40: one large innovation can put three readings above a threshold, 6 steps apart. The rule marks the 0.9 quantile and 19 readings clear it; 12 of the gaps between consecutive exceedances are exactly 6 steps. With a run length of 4, so that two exceedances 4 or more steps apart start separate clusters, they form 19 clusters, shaded, the largest holding 1. The runs estimator reads 1.000 against 0.40.

The run length a declustering chooses

The runs estimator of an extremal index carries a constant nobody derives. Where a cluster is a run of neighbouring exceedances the constant barely matters; where a cluster's members fall six steps apart, the estimate is 0.9069 at a run length of six and 0.3649 at seven against an index of 0.40, and a run length of four removes under a tenth of the overstatement declustering exists to remove. A rule that reads the run length off the data has the smallest worst error of the three.

extreme · Extremes
What the forecast interval is short by, φ = 0.85, 6 steps ahead. The plug-in interval covers 88.42% against a claimed 95%. Correcting the variance recovers 0.56 points, propagating the persistence's own standard error recovers 0.40, correcting the persistence recovers 2.66, and all three together recover 4.20 — leaving 2.38 points unaccounted for.

What the interval is short by

The forecast interval covers 88.42% where it claims 95%. Correcting the persistence recovers 2.66 points, correcting the innovation variance 0.56, propagating the persistence's own standard error 0.40 — and all three together recover 4.20 of the 6.58, leaving a residual none of the standard repairs reaches.

evaluation · Bias

Named alongside it

The objects these essays reach for when they reach for this one.

Monte CarloModel selectionAutocorrelationOut of sampleBenchmark forecastClosed formDependenceLong-run varianceTaperingBias-varianceEstimation errorInformation criterion

All concepts