Concept

Whitening — where it appears

A linear transform applied to a sample so that its errors become uncorrelated with equal variances. Least squares on the transformed sample is then generalised least squares on the original, so a dependence is removed before a procedure that assumes independence reads the data rather than corrected for afterwards.

Named by 23 essays across 8 fields — each of them below, with the objects they name alongside it.

The profile a break point is chosen from. One sample of 120 rows under a break in the persistence, fitted as two first-order regimes at every admissible break point. The maximum is at row 78, where the true break is at 60. The shaded band is every break point within two log-likelihood units of the best one — 8 of the 73 positions searched, which is 11% of the range. The horizontal line is the one-regime fit the search is compared against; the whole profile is above it, at every position, which is the point: a maximum over 73 candidates is above the null by construction and not by evidence.

A break that was looked for

A two-regime whitening finds its change point by maximising a profile, and then reads a criterion that counts parameters. Under no break there is no parameter to count, because every position describes the same model.

charged · Break point
The dependence, at four removes. Under AR(1) at 0.8, four different sequences all called the dependence. The top line is the law. The middle line is what a sample of 120 errors reports on average — computable exactly, because the expectation of a sample autocovariance is arithmetic once the covariance is known. The lower line is what a candidate's residuals report, which is what every two-step rule in this collection actually reads: a fit removes variance, and it removes more of the persistent part than of the rest. At the first lag the three are 0.800, 0.7773 and 0.7338. The dots are counted from draws and share no arithmetic with the line they sit on; the worst departure is 1.2 standard errors.

A dependence fitted with the line

Every whitening in this collection reads the dependence off a set of residuals, and residuals are not errors. Fitting the two together recovers most of what that costs, and changes almost nothing about the decision it feeds.

together · Dependence
Four dependences a single parameter cannot tell apart. Every law here is standardised to a lag-one autocorrelation of 0.8, so a rule told the errors are a first-order autoregression finds the same number in all four and has no way of seeing what separates them. The geometric decay is the world in which estimating a covariance rather than naming it was priced, and found to cost. The five-period moving average has 0.200 at the fourth lag and exactly nothing past it, where the geometric law says 0.328 at the fifth. Long memory at d = 4/9 is still at 0.576 by the twentieth lag, where the geometric law has reached 0.012. The break has no autocorrelation function at all: what is drawn for it is the average over the pairs at each gap, which is what a stationary estimate converges to.

A dependence with a shape

Four ways for errors to repeat, all with the same first lag and nothing else in common. A rule told the errors are a first-order autoregression finds the same number in all four, and is right about one of them.

general · Dependence
Where the general fit becomes the parametric one. The band family's objective at the autoregression's own geometric sequence, cut off at each width, on one sample of 60 rows. The horizontal line is the profile likelihood the parametric fit maximises, written independently through a different whitening. At the full width the two are the same number to 3e-14, which is what says the general construction contains the parametric one rather than resembling it. Below 10 lags there is no line at all: the geometric sequence cut off short is not a covariance matrix, so the objective has nothing to evaluate. Between the two the truncation is briefly above the parametric likelihood — a wrong covariance can fit one sample better than the right one, which is the whole reason a width has to be charged for rather than chosen.

A family before a fit

A regression's coefficients and one correlation can be maximised together. Replace the correlation with an estimated covariance and there is nothing left for "jointly" to mean — until a set of covariances is named, and the set turns out not to contain the truth.

family · Dependence
One factor moves and the other does not. The two factors of the same average, each drawn against its own largest value so that they share an axis. The rate at which the five candidates disagree about the tuning parameter rises from 28.6% at 4 values on the list to 43.3% at 8, a factor of 1.52. What a disagreement costs, given that there was one, is 0.00975 ± 0.00224 and 0.00848 ± 0.00113 at the same two points — 0.5 standard errors apart, and the paired comparison on the draws that disagree under both lists puts it the other way. The guess this field was written to test was that a longer list makes disagreements commoner and each one smaller. The first half is right and there is no second half.

A rate times a size

A sweep reported what it costs to let every candidate choose its own tuning parameter and found it flat across the list. It was reporting a product, and the two things multiplied together do not behave the same way at all.

apiece · Order-selection
The number the comparison was missing. What it costs to choose the tuning parameter for every candidate separately rather than once for the table, under AR(1) at 0.8, paired on the draw. The window's figure is the one the earlier field reported; the order's is the one it named and did not make. They are the same size — 0.00401 against 0.00360, at 2.30 and 1.72 paired standard errors — and matching the lists at eight values leaves them the same size again. The prediction that the longer list would make the order's cost the larger of the two is not what happens; what happens is that the two rules cost the same once they are scored by the same criterion, which took a missing term to arrange.

The comparison that was not made

Choosing a whitening's window separately for every candidate costs 0.00401 of regret. The same question about an order was named and left, because the two lists are different lengths. The order's answer is 0.00360, and matching the lists changes almost nothing.

lists · Order-selection
Two searches find some of the same luck. What each search reports on a sample with no break in it, and what the two report together, on four dependences. The dashed line is the sum of the two — what a rule charging each search separately would levy — and the two together always come in below it: 19.30, 15.44, 14.24, 26.64 short, on 100%, 99%, 99%, 100% of draws. The shortfall is not a rounding. Under AR(1) at 0.8 it is 19.30 of the 34.70 the break search manufactures on its own, which is more than half of it. Two searches over one sample are looking at the same noise, and the second one has less left to find.

Two searches, one sample

A searched break in a regression manufactures 34.7 of likelihood ratio where a count of coefficients says 11.1. A searched window manufactures 84.0. The two together manufacture 99.4, not 118.7.

twice · Break point
How much memory a fit takes out, candidate by candidate. Under AR(1) at 0.8, the lag-one autocorrelation a candidate's residuals report, computed exactly for each candidate on 200 draws. The upper line is the law at 0.8000. A candidate that is an intercept alone reports 0.7773 — which is exactly what a sample of 120 errors reports, because an intercept annihilates the sample mean and nothing else, and the two arithmetics agree to the last bit. Every predictor after that takes more out, down to 0.7341 at the fullest candidate. That is the collision this field is about: the rule every whitening here uses estimates its nuisance once, from the fullest candidate, so that the criteria stay comparable — and the fullest candidate is the one whose residuals report the least.

The fit that takes the memory out

A candidate's residuals report less dependence than its errors do, and how much less is arithmetic rather than noise. The rule used for a good reason reads the series that has lost the most.

together · Dependence
Most of the rise is the optimiser's, and under one law it is not. The rise in log-likelihood from the tapered plug-in to the maximum over the same eight-lag band, beside what the same optimiser produces on a sample generated from the plug-in's own covariance — where the family is correctly specified by construction and there is nothing to find. Under AR(1) at 0.8 the raw rise is 5.72 and the manufactured baseline is 4.79, leaving 0.93 at 1.8 standard errors; under long memory the excess is 0.14, at 0.2. Under the moving average it is 11.87 at 19.4 standard errors, on every draw. The taper is a shrinkage, and it costs nothing where the sequence decays smoothly and a great deal where it stops dead.

The plug-in and the maximum

A tapered covariance estimate sits five and a half log-likelihood units below the maximum of the likelihood it is substituted into. Four fifths of that is what the optimiser would have found if nothing were missing.

family · Dependence
What a disagreement costs, split on whether it decided anything. The regret from choosing the tuning parameter per candidate, on the draws where the candidates disagreed, split on whether the disagreement changed which candidate the table selects. Over 1200 draws at each list length: when the winner changes the regret is 0.02215, 0.03029, 0.03145; when it does not it is -0.00243, -0.00069, -0.00065 — negative, and small enough that it is inside two standard errors of nothing at every length. The whole of the cost lives in the first column, and the second column is not merely small but slightly the wrong sign: when the table's answer is unaffected, letting each candidate use its own window is a very slightly better rule than making them share one. So a disagreement about the tuning parameter is not a cost. A disagreement that changes the winner is.

The quarrel that changes the winner

A disagreement about the tuning parameter costs 0.031 when it changes which candidate the table selects and −0.0007 when it does not. The distance between the values disagreed about has nothing to do with it.

apiece · Order-selection
The term that cancels, and the term that does not. The volume each candidate's whitening moves — log|Ω̂| — for a sieve of order 4 on one sample of 120 rows. Estimated once from the fullest candidate and used for the whole table, it is the same number for every candidate, so it drops out of every difference the criterion reads: that is why nothing in this collection has ever needed to carry it. Estimated from each candidate's own residuals it ranges over 23.87, which is more than a parameter is worth, and the criteria being compared are then fits made under different error models with no term saying so. The window's rule has carried this term since the estimated-covariance field and the sieve's never had it.

The volume a whitening moves

A sieve's whitening has a determinant and this collection's criterion for it never carried one. Shared across a table the term cancels exactly, which is why nothing ever noticed; used per candidate it is worth more than a parameter and the whole comparison turns on it.

lists · Criterion
Generality in the wrong direction buys nothing. Regret on a sample whose persistence changes from 0.95 to 0.65 at row 60, over 200 draws. The three stationary rules — told one number, told a window, told an order — are within 0.4 standard errors of each other, and all three stop in the same place: they are general in the lag direction, and the departure is in the other one. Letting the model change once, at a point estimated from the same residuals, is worth 0.05021 more at 4.5 paired standard errors — about as much again as the whole of the first repair. Being told where the break is adds 0.01926, and being told the entire covariance adds 0.02465.

Where the generality runs out

A covariance that changes half way through a sample is not one a window can estimate. One number, a window and an order are worth the same as each other on it — and letting the model change once, at a point nobody can locate, is worth as much again as all three.

general · Dependence
What a search costs is not a property of that search. The likelihood ratio a searched break in the regression reports, two ways, on every law. On its own — the whole rule being a split of the sample, at no whitening — it averages 34.70 under AR(1) at 0.8, against the 11.07 a chi-square on the five coefficients a split adds would use as a threshold. Inside a rule that also chooses a window from a list of eight, the same search adds only 15.40 — less than half. Most of what a break search finds under correlated errors is the correlation, and a whitening chosen from the same sample has taken it already. A charge measured for one search, carried into a rule that makes two, is not conservative in some harmless direction: it is measuring a different quantity.

A charge that depends on the rule

The break search's charge is 34.7 on its own and 15.4 once a window has been chosen from the same sample. Most of what a break search finds under correlated errors is the correlation, and a whitening has taken it already.

twice · Break point
What a longer list actually changes. How often the five candidates choose different tuning parameters, at a true null where every one of them contains the truth, so a disagreement is manufactured rather than discovered. The order's list is an interval of integers, and thinning it moves the rate smoothly from 0% at two values to 39% at thirteen. The window's is not an interval — it runs 0, 1, 2, 4, 8, 12, 20, 30 — so a thinned window list jumps depending on whether it happens to keep the width the criterion wants, between 0% and 42% with no order to it. So "the same length" was never quite the same thing for the two rules, and it is a smaller effect than the field it was invoked to explain.

A list is not a rule

How often five candidates disagree about a tuning parameter runs from nothing at two values on the list to two draws in five at thirteen. What the disagreement costs does not move at all.

lists · Order-selection
How often the split is taken, and by which rule. Over 400 draws on each of five laws. The first two rows have no break in them at all, the last two have one at row 60, and the middle one is a moving average. A criterion that counts a fitted two-regime model's parameters and nothing else takes the split on 99% of draws where there is no break. Counting the break point as one more parameter brings that to 67%. Charging what the search actually manufactures — 5.16 units, measured on a law with no break — brings it to 16%, and still takes the split on 61% of draws where there is one.

Choosing whether to break

Charging what the search manufactures takes a rule from splitting a stationary sample on 99% of draws to 16%. It also costs regret, because the two mistakes a rule can make are not the same size.

charged · Break point
The quantity that does not depend on the list. The probability that letting each candidate choose its own tuning parameter changes which candidate the table selects — the product of the two moving shares — against the length of the list, over 1200 draws apiece. It is 14.2%, 11.9%, 12.3%: a spread of 2.2% across a list length that moves the disagreement rate by a factor of 1.52. This is the invariant the whole field turns on. Everything downstream of the winner — the coefficients, the regret, whatever a reader is going to quote — is a function of whether the winner changed, and how often that happens is not something the list controls. A longer list changes how often the candidates quarrel and not how often the quarrel matters.

How often it matters

The disagreement rate rises by half across the list and the share of disagreements that decide anything falls by nearly the same factor. Their product — how often the tuning list changes which candidate wins — sits at an eighth and does not move.

apiece · Order-selection
One likelihood, three answers. The concentrated Gaussian log-likelihood of one sample of 120 rows under AR(1) at 0.8, as a function of the correlation the errors are whitened at. Three rules put three different numbers on this curve. The two-step rule reads the least-squares residuals and lands at 0.7616, giving up 0.304 of log-likelihood. Iterating moves it to 0.8080 and gives up 0.002. The maximum is at 0.8044. The curve is not flat between them: what a fixed point of the residual update finds is a solution of a different equation, and the difference is the Jacobian term ½log(1 − ρ²), which grows as the correlation does.

Iterating is not maximising

Re-reading a correlation from the generalised residuals and refitting converges in seven steps. What it converges to solves the first-order condition of a sum of squares, and the likelihood has one term more than that.

together · Dependence
A wider band is always a better fit. The likelihood maximised over the band, at five widths, averaged over 30 samples. A band at L lags is a band at L + 1 with the last entry held at zero, so the families are nested and the maximised likelihood cannot fall — it does not, on any draw. What it does is rise at 0.984 of log-likelihood a lag. A parameter that is doing nothing buys half a unit in expectation and Akaike's criterion charges one, so this is a criterion very nearly indifferent between every width on offer. The dashed line is what a charge of one unit a lag would exactly cancel. Nothing in the fit chooses a width, and what does choose one is a charge somebody has to pick.

Nothing in the fit picks the width

A wider band is always a better fit, and it is better by about one unit of log-likelihood a lag — which is the order of what a criterion charges for a parameter. Three defensible rules choose widths a factor of three apart.

family · Order-selection
The window a whitening wants is not the memory of the errors. Regret under a five-period moving average as the tapered estimate is given more lags, over 120 draws at n = 120. The best window is L = 30; the automatic bandwidth is 4 and the error model's own likelihood chooses 9.7 on average. Both land in the same place and both are short, and the reason is the taper: a Bartlett weight at lag k is 1 − k/(L + 1), so a window of 8 keeps 0.556 of whatever the fourth lag carries and a window of 30 keeps 0.871. A window has to be several times the memory before it stops removing the memory. The dashed line is the rule told the errors are a first-order autoregression, which needs no window at all.

The window a whitening wants

Every law here is best whitened by a window several times longer than its own memory, including the one whose memory ends at the fourth lag. The three ways of choosing it from the sample all land in the same place, and it is the wrong one.

general · Order-selection
The ladder is the same ladder under every law. Each pair's overlap under each of the four laws, over 400 draws apiece. The scale is fixed at both ends by construction: two searches over disjoint sets of independent columns read -0.001, 0.013, -0.002, 0.004, and a break search paired with a step column it already contains reads exactly one under every law. Between them the pair that reads one residual series twice runs 0.763, 0.805, 0.577, 0.752 — lowest under long memory, where a whitening has most to do and the break search has least left to find that the whitening has not taken. The rung that moves most is the pair of step dictionaries, from -0.148 under a moving average to 0.419 under a break; and the pair that is negative is negative under all four.

Three quarters of the way to one search

The pair that started this reads 0.762 on a scale whose one is containment. And the pair that shares nothing but its response reads −0.306, so the sign the earlier field found does not transport at all.

apart · Criterion
One window for the table, or one each. The regret of the same fifteen-candidate table under AR(1) at 0.8 over 150 draws, with the window attached three ways. Chosen once from the fullest candidate's residuals it gives up 0.02518. Chosen from each candidate's own residuals, with the covariance estimate still shared, it gives up 0.02799 — a paired cost of 0.00281 at 2.0 standard errors for the tuning parameter alone. Estimating the covariance per candidate as well costs 0.01087, so the objection already on record is about 3.9 times the size of the one that was not.

A window for every candidate

The window and the order a whitening needs are chosen once, from the fullest candidate, on an argument that was made about an estimated covariance. A tuning parameter is not a covariance, and the two cost different amounts.

together · Order-selection
Four sequences, and the rule only ever sees the last one. Under long memory at d = 4/9, four things that are all called the dependence. The law itself is the top line. What a sample of 120 rows reports on average is the second, computed exactly: subtracting a sample mean takes the first lag from 0.800 to 0.538. What a candidate's residuals report is the third, lower again at 0.472, because a fit removes dependence along with signal. The autoregressions are fitted to that third sequence and reproduce it exactly out to their own order — the Yule–Walker equations are solved to make it so — so everything they say past that is extrapolation. At the twentieth lag the law has 0.576, the residuals report 0.006, and an AR(8) extrapolates 0.028.

The order the tail is drawn at

A fitted autoregression reproduces the sample exactly at the lags it was fitted on, so everything it says past them is extrapolation — and the order is the dial that decides how much of it there is.

general · Order-selection
Fitting them together is worth something under one law. Four fits of the same regression under four dependences: least squares, the two-step plug-in every whitened rule in this collection runs, the coefficients and the band maximised together, and a whitening at the law's own covariance that nobody has. Under the moving average — the one law the band family contains — the joint fit beats the two-step by 0.0077 at 3.3 paired standard errors. Under the autoregression, long memory and the break it is a tie: 0.4, 1.0, 0.3 standard errors. That is the same ordering the likelihood gap gave, arrived at through the coefficients rather than through the objective.

What fitting them together buys

Maximising over the coefficients and the covariance together beats the two-step under one of four dependences and ties under the other three. It is the one the band family contains, and the likelihood said so before any coefficient was compared.

family · Dependence

Named alongside it

The objects these essays reach for when they reach for this one.

Model selectionInformation criterionNuisance parameterRegretSelection effectCovariance matrixLong memoryDependenceStructural breakAutocorrelationGeneralised least squaresMonte Carlo

All concepts