Concept

Benchmark forecast — where it appears

A forecast that estimates no dynamic parameter — the last value carried forward, or the long-run average — against which a fitted model has to justify itself. It is the row a specification search compares everything against, and whether it was named in advance or selected from the same table changes the reference distribution for every comparison in it.

Named by 21 essays across 10 fields — each of them below, with the objects they name alongside it.

What each variant loses before anything has been searched for. The mean loss differential of each of the eight variants against the benchmark, over 600 tables of 60 origins, with every fit given 71 rows. The series is an AR(1) and every variant adds a lag whose coefficient is zero, so in population the two forecasts are the same forecast and the difference drawn here is estimation noise and nothing else. The marked line is σ²(q₁ − q₀)/n = -0.01408, which is an expression in how many coefficients each model has and how many rows it was fitted on — it knows nothing about the series, the persistence or which lag the variant added, and every bar is within a fifth of it. This is the amount a reference distribution recentred at each column's own sample mean believes the candidates are already behind by.

A table of nested models

A benchmark and eight variants of it, each adding one thing. Every variant is behind before the search begins, by an amount that can be written down before the data exists — and the two most natural ways of reading the table are wrong in opposite directions.

search · Forecast
The eighth was not a constant. The probability that a per-candidate tuning list changes which candidate the table selects, at each of 5 separations between the candidates, over 800 draws apiece. The earlier field reports this flat at about an eighth across list length, on a table and a world it never varies. Vary how much the omitted coefficients are worth — one multiplier, with the table, the list, the law and the sample size all held — and it runs from 17.6% to 1.5%, a factor of 11.75. The world in which every candidate is true is the world in which the tuning list decides most; the world in which one candidate dominates is the world in which it decides nothing.

The eighth that was not a constant

How often a per-candidate tuning list changes which candidate wins is reported flat at about an eighth across list length. Vary how far apart the candidates are instead and it runs from 17.6% to 1.5%.

turnover · Order-selection
One forecast, and the band the arithmetic puts round it. An AR(1) with φ = 0.75, 60 observations, fitted by least squares and forecast 14 steps ahead. The point forecast decays towards the fitted mean at φ̂^h; the band is ±1.96 standard errors from σ̂²Σψ̂², which grows with the horizon and stops at the unconditional spread 1.72. The dashed pair is the same band computed at the true parameters, which nobody has. The marks past zero are what actually arrived: 12 of 14 inside the band this once, which is one draw and settles nothing.

What the model says next

The usual account of a time series stops at estimation. A forecast asks the other question — not what the parameter is but what the next observation will be — and the band round it is a closed form that grows with the horizon and then stops growing, at a value the series was going to reach anyway.

forecast · Forecast
Three quantities, and only one of them crosses zero. Two forecasts of an AR(1) — the last value carried forward and the mean of the last 60 observations — at 1 step ahead. The curve through zero is σ₁² − σ₂², the difference in expected squared error that a comparison of accuracy tests; it changes sign at φ = 0.4922. The two curves above it are σ₁² − σ₁₂ and σ₂² − σ₁₂, the quantities the two encompassing tests are about, and neither of them comes near zero anywhere: the smallest value either takes across the range is 0.008 times the variance of the series. All three are closed forms in φ, R and h with no simulation in them. Equal accuracy is one hypothesis about this picture and encompassing is another, and a set of numbers can satisfy either without the other.

What the other forecast adds

Two forecasters, one series, and two different questions about them. Which is more accurate has an answer that changes with the persistence of the series; whether either is redundant has an answer that never changes at all.

ranking · Forecast
One true null, one table, five readings. every subset of four, fifteen models, at a null where nothing any candidate holds is worth anything, over 500 draws. Each bar is the share of draws on which that reading declares a difference at a nominal 5%. The reading is the whole of the difference between the bars: the data is identical. An open search over all 210 ordered pairs rejects 76.2%; the table's own 5% point is 3.163 against the 1.671 a single comparison uses. Bonferroni takes the open reading to 0.6% — and on the nested ladder the same correction does not reach the nominal level at all, because there the excess is a shift in the mean rather than a maximum over many.

When the benchmark is a candidate

A specification search with a benchmark nailed down is the case with a closed form. Take the nail out — let the model that would have been reported be one of sixteen, chosen by the same data as its rivals — and the same true null is read three ways, at 2.0%, 7.8% and 76.2%.

select · Forecast
One comparison, and the two error bars it can be given. 60 rolling origins, a window of 60 observations, forecasts 4 steps ahead, at the persistence φ = 0.8256 where the two benchmarks have exactly equal population mean squared error. Each mark is one origin's difference in squared error; the horizontal line is their mean, 0.6522. The two vertical bars at the right are ±1.96 standard errors round that mean computed two ways — 0.5337 treating the differences as independent, 0.6880 allowing for the overlap between neighbouring forecasts. The null is true here by construction, so an interval that excludes zero is a mistake, and the narrow one does it far more often than the wide one.

Which forecast is better

Two forecasters, one series, and a difference in mean squared error. Whether that difference is real is a hypothesis test, its terms are not independent, and the standard error it needs is not the one a t-test computes.

evaluation · Forecast
Eight candidates, one of them exactly as good as the benchmark. The candidate set: moving averages of the last 1, 2, 3, 5, 8, 13, 21 and 34 observations, each drawn as its expected squared error divided by the benchmark's — the mean of all 60. The persistence is not chosen, it is solved for: at φ = 0.4895 the best candidate in the set, the average of 2, has exactly the benchmark's expected squared error, and every other candidate is worse by between 0.5% and 5.1%. So the null that no candidate beats the benchmark is true, with one candidate on its boundary. Everything a set comparison claims about its own error rate has to be measured here, because anywhere further inside the null every procedure flatters itself.

Eight forecasters and one benchmark

A set of forecasters is a multiplicity problem on top of a dependence problem, and the two do not separate. Eight windows of one series carry the multiplicity of two and a half independent comparisons; eight separate problems carry eight.

ranking · Multiplicity
Every candidate is behind by what its parameter count says. Each dot is one of the fifteen subsets of four predictors, fitted on a rolling window of 80 rows and scored against the benchmark out of sample over 60 origins, at a null where every one of them contains the truth. The line is σ²(q₀/(R − q₀ − 1) − q/(R − q − 1)), which is arithmetic on two integers and a window length. Most of these pairs are not nested — a subset of two predictors and a different subset of two share neither model — and the closed form does not care: the displacement is a statement about how many coefficients each side estimates. The candidates of the benchmark's own dimension sit at zero.

The displacement is a parameter count

A nested variant is behind its benchmark out of sample before anything is searched for. The closed form for how far turns out to have nothing about nesting in it — only two integers and a window length — and it prices a table where no candidate contains any other.

select · Multiplicity
What a disagreement costs, split on whether it decided anything. The regret from choosing the tuning parameter per candidate, on the draws where the candidates disagreed, split on whether the disagreement changed which candidate the table selects. Over 1200 draws at each list length: when the winner changes the regret is 0.02215, 0.03029, 0.03145; when it does not it is -0.00243, -0.00069, -0.00065 — negative, and small enough that it is inside two standard errors of nothing at every length. The whole of the cost lives in the first column, and the second column is not merely small but slightly the wrong sign: when the table's answer is unaffected, letting each candidate use its own window is a very slightly better rule than making them share one. So a disagreement about the tuning parameter is not a cost. A disagreement that changes the winner is.

The quarrel that changes the winner

A disagreement about the tuning parameter costs 0.031 when it changes which candidate the table selects and −0.0007 when it does not. The distance between the values disagreed about has nothing to do with it.

apiece · Order-selection
Two factors, opposite directions. The two factors the cost of a per-candidate tuning parameter is a product of, as the candidates are pulled apart, over 800 draws at each of 5 separations. How often the candidates disagree about the tuning parameter rises from 31.8% to 88.8%; the share of those disagreements that change which candidate the table selects falls from 51.6% to 1.7%. So the setting where the candidates quarrel most about the tuning parameter is the setting where the quarrel matters least, and a sweep that reads the rate and stops has read the factor pointing the wrong way.

Two factors pointing opposite ways

As the candidates on a table are pulled apart, they quarrel about the tuning parameter three times as often and the quarrel decides the winner thirty times less often. A sweep that reads the first factor has read the one pointing the wrong way.

turnover · Order-selection
Both halves grow; the difference does not. The control pair's two components and their difference, against how much each of its two searches can find, over 1200 draws at each dictionary size. Two disjoint sets of independent columns are additive at every size — the excess stays inside a standard error or two of zero throughout — and it is not because there is nothing there. The overlap grows from 0.000112 at two columns to 0.000870 at ten, a factor of 7.76, and the interaction grows with it, staying within a factor of two of the overlap at every size. Two searches competing for one residual sum share ground and find configurations neither has alone, in almost equal measure, and their difference is what the earlier field's scale calls zero.

What a zero is made of

Two disjoint dictionaries of independent columns read an excess of 0.000116 and are made of an overlap of 0.000583 and an interaction of 0.000467. The control the whole scale is anchored on reads zero because two effects cancel.

separate · Break point
Flat along a row, apart between them. The probability that a per-candidate tuning list changes the winner, at three list lengths on three candidate tables, over 800 draws in each of the nine cells. Along a row — the reading the earlier field takes — it moves by a factor of at most 1.21, so that field's invariant survives on every table. Down a column it moves by up to 1.98. The list length is the dial that does not move this number and the table is one that does, and the earlier field varied only the first.

A table and a list

A nested ladder of candidates differing by one coefficient was predicted to turn over more often at every list length. It turns over less at every one, and its list changes the winner half as often.

turnover · Order-selection
Sixteen candidates nobody would have run, and what they cost. At φ = 0.65 one candidate in the original set is genuinely better than the benchmark, and the question is how often each procedure finds it. The added candidates are stale copies of the last value — read two, four, six … steps late — every one of them worse than the benchmark by at least 43%, and not one of them is ever the best candidate in a sample. The reality check goes from 35.5% to 0.0% as they are added, because its reference distribution has to assume every candidate is exactly as good as the benchmark and sixteen such assumptions is a critical value nothing reaches. The recentred version, which drops from the recentring the candidates the data has already ruled out — 15.1 of 24 of them — goes from 25.5% to 24.5%. The top line never moves: reporting the winner's own p-value cannot notice a change to a set it never looks at.

The models that were never in the running

A reference distribution for a set has to assume something about every candidate in it. Assuming that all of them are as good as the benchmark is what makes the reality check honest, and it is what sixteen hopeless candidates use to destroy it.

ranking · Multiplicity
The ranking on the left, the weights on the right. Eight moving-average forecasts of an AR(1) at φ = 0.4895, the persistence at which the best of them exactly ties the 60-observation benchmark. On the left, each candidate's expected squared error in units of the series' own variance: the smallest belongs to L = 2, at 1.0156. On the right, the weight each carries in the variance-minimising combination of all eight — and the best of them carries 0.00000. The two ends of the family carry 1.0172 of the weight between them, and the combination they make is worth 0.7817, which is 23.0% below the best single forecast. Both columns are closed forms in φ. Which forecast to keep and which forecasts to use are different questions, and this is a set where the answers share nothing.

The weight that is a vector

Two forecasts have a best combination and one number describes it. Eight have a best combination too, and the vector describing it puts nothing at all on the forecast with the smallest mean squared error.

search · Rank
The distribution the table does not have. 599 series simulated from the smaller model fitted to one comparison's own data, the whole rolling comparison re-run on each, and the ordinary statistic recorded. Under this null the two forecasts are the same forecast in population, so what is left in a sample is the larger model's estimation error and the statistic is centred at -1.134 rather than at zero. Its 95% point is 0.264; the standard normal drawn behind it puts that point at 1.645. Reading this statistic against that curve is not a poor approximation, it is a different distribution: the share of this one above 1.645 is 0.2%.

A distribution drawn from the null

Between nested models the ordinary comparison statistic has a null distribution centred at minus one and a 95% point of a quarter. A correction to its mean repairs the centre and leaves the shape; simulating the null repairs both.

ranking · Bootstrap
What the corner costs when the table is full of hopeless candidates. The benchmark holds two predictors, one of which is worth 1; a third predictor, worth the amount on the horizontal axis, is held only by candidates the benchmark does not contain. At the left the null is true and both procedures hold their level. To the right there is a genuinely better candidate, and the uncorrected reality check finds it 2.7% of the time while the same test with the clearly bad columns recentred finds it 51.0% of the time. The columns doing the damage are the ones nobody would have looked at twice: they are so far behind that they cannot win, and calibrating as though they might is what makes the test blind.

The corner the test is calibrated at

"No candidate is better than the benchmark" is not a null but a face of a region, and a reality check is calibrated at one corner of it. Fill the table with candidates that are hopeless rather than equal and the test finds a genuine improvement 0.0% of the time.

select · Reference
What the interval covers once the order is chosen as well. 1200 series of 40 observations from an AR(1) with φ = 0.7, at each horizon, on one set of seeds. The upper line is the interval computed at the true parameters — it covers 95.5% on average, which is the check that σ²Σψ² is the right formula rather than a claim about anything a forecaster can do. The lower line is the same formula fed σ̂² and φ̂: 93.6% at one step and 90.8% at 6. The third line chooses the order by AIC from the same data before computing the interval, which costs a further 0.8 points at h = 6.

The interval after the choice

Estimating the coefficients of a known model costs a 95% forecast interval about two points of coverage. Choosing which coefficients to estimate, from the same forty observations, costs another four and a half — so the step nobody records in the output is the more expensive of the two.

forecast · Order-selection
The scale moves the width; the curve does not. The band width each charge picks, averaged over 150 draws of 120 rows. The two conventions — a unit a lag and half a log n a lag — pick 6.08 and 3.65 lags. The four charges derived from the measured optimism pick 15.05, 15.60, 15.47 and 15.44, against a best width on the draw of 14.13. So the scale a charge is levied on moves the width by a factor of 4.27 and the shape of the charge moves it by 3.6%. None of the six is an estimate of the draw's own best width: the correlations are -0.006, -0.003, 0.017, -0.001, 0.016, -0.004.

What a better charge buys

Four charges derived from the same measurements pick band widths within six per cent of each other and deliver errors within two per cent of the gap any of them leaves. The scale a charge is levied on decides the width; the shape of the charge decides nothing.

curve · Criterion
Eight encompassing nulls, all of them true at once. The forecast under test is the variance-minimising combination of the eight candidates, and the first-order condition that defines its weights is cov(e_c, e_j) = var(e_c) for every j — so every one of the eight nulls is exactly true simultaneously and every rejection counted here is false. 800 draws of 60 origins. The largest of the eight statistics rejects 25.0% of the time at a nominal 5% and one comparison stated in advance rejects 3.1%, which makes the set worth about 9.1 independent comparisons — nearly the eight it has. Bonferroni, which is far inside its level on a search over accuracy comparisons of the same eight forecasters, is at 4.8% here.

When every null is true

A reality check assumes that every candidate in the set is exactly as good as the benchmark, which is a configuration nobody's data is ever in. Test a combination against its own parts and that configuration is not assumed — it is what the arithmetic makes true.

search · Multiplicity
Where each rule says to put the number. The expected score of reporting each probability on the axis when the event's true probability is 0.25, for four scoring rules. The Brier, logarithmic and spherical scores each bottom out at 0.25 — found by search rather than assumed, to 8 decimal places — which is what makes them proper: a forecaster with a genuine belief cannot improve its expected score by reporting anything else. The absolute-error score is a straight line in the reported value, p + r(1 − 2p), so it has no interior minimum at all; its optimum is 0.0, a distance of 0.250 from the truth, and taking it saves 0.125.

A score that rewards lying

An absolute-error score pays a forecaster exactly ⅛ of a point to replace a true quarter with a zero, and over two hundred records a liar beats a truthful forecaster on 200 of 200. A skill score against the forecaster's own average buys 0.012633 of reported skill for 0.002035 of real score.

calibrate · Calibration
Reading the draw changes what is charged, not what is tracked. The correlation between the band width each rule picks and the best band width on the same draw, over 400 draws. The three fixed charges read -0.069, -0.066, 0.012. The three that read the sample read -0.012, -0.019, -0.041. None of the six is distinguishable from nothing. The statistic the first plug-in reads does vary — the draw's own summed squared autocorrelation runs from 2.06 to 10.19 with a mean of 4.36 — so the failure is not that the charge stopped moving. It is that what it moves with carries no information about which width this draw wanted.

A charge that reads the draw

Three charges built to read the sample track the best band width on their own draw at −0.012, −0.019 and −0.041, deliver more error than the fixed rule they are calibrated to, and pick a width half again as variable. The statistic moves; the answer does not.

curve · Criterion

Named alongside it

The objects these essays reach for when they reach for this one.

Monte CarloModel selectionMean squared errorOverfittingInformation criterionOut of sampleError rateNested modelsNull hypothesisSelection effectBandwidth selectionDependence

All concepts