When the observations repeat each other

The cliff that is a slope

A regression between two independent series is called significant 4.9% of the time at no persistence, 52.4% at a lag-one correlation of 0.9, and 83.4% at a unit root. The rule the field offers asks whether the last of those holds, and at 0.9 the unit-root test correctly refuses one 87.2% of the time.

Worth reading first: Two walks and a finding · Three series and a count.

The field opened with two independent random walks and a slope that was called significant three times in four, and it has treated that as a statement about random walks ever since. The repair it offers is a rule with a switch in it: test each series for a unit root, and if the test finds one, difference before regressing.

The switch asks whether one number is exactly one. What the damage depends on is how close that number is, and the two are not the same question. Swept across the persistence of the two series, the rate at which a regression between them is called significant reads 4.85% at no persistence, 13.80% at a lag-one correlation of 0.5, 34.20% at 0.8, 52.45% at 0.9 and 83.40% at a unit root. There is no value at which the curve turns. The rule placed on it has a step in it because a rule has to decide something, and the quantity it decides about does not.

The damage and the warning, against the same dialTwo readings at each persistence. In the darker colour, how often a regression between two independent series of 200 steps is called significant at 5%: 4.9% at φ = 0, 34.2% at 0.8, 52.4% at 0.9, 83.4% at a unit root. In the lighter, how often the standard unit-root test refuses a unit root on one of those series — the chance the analyst is told the series is stationary and may be regressed: 87.2% at φ = 0.9 and 31.9% at 0.95. At φ = 0.9 both are high at once, which is a correct diagnostic licensing a regression that is wrong half the time.00.2000.4000.6000.800100.2000.4000.6000.8001φ, how much of yesterday survives into todayrate5%independent series called relateda unit root refused2,000 independent pairs at each φ, 200 steps, critical value -2.8852.4% at φ = 0.9
Fig. 1 Two readings at each persistence, over two hundred observations. The darker curve is how often a regression between two independently generated series is called significant at 5%. The lighter one is how often the standard unit-root test refuses a unit root on one of those series — the chance the analyst is told the series is well behaved. At φ = 0.9 the first reads 52.45% and the second 87.15%.

The second curve is the one that makes this more than a remark about continuity. At a persistence of 0.9 the unit-root test is doing its job: the series genuinely is stationary, it genuinely has a level to return to, and the test says so 87.15% of the time. An analyst who runs the check, reads a clean result, and regresses the levels is following the procedure correctly at every step, and the regression they then run between two unrelated series is significant more often than not.

A dial nobody put a mark on

The processes being swept are first-order autoregressions. Each observation is φ times the last one plus a fresh disturbance, so φ is how much of yesterday survives into today, and a disturbance decays by a factor of φ every step. At φ = 0.5 half of it is gone in one step; at φ = 0.9 it takes 6.58 steps to halve; at φ = 0.99, 68.97. At φ = 1 it never goes, which is what a random walk is — and a series that never returns to a level is not the same thing as a series that has not converged yet, though the two are routinely described in the same words.

Four series, one of which never comes back. One realisation of 200 steps at each of four persistences, each panel scaled to its own series. Three of them are stationary and return to a level they have; the fourth is a random walk and has no level to return to. The half-lives of a disturbance are 3.1, 6.6, 13.5 steps and then infinite. Nothing in the shapes says which is which, and the difference between the third panel and the fourth is what every rule in this field is asked to decide.
Fig. 2 One realisation of two hundred steps at each of four persistences, each panel scaled to its own series. Three are stationary and one is a walk. Nothing in the shapes says which is which, and the difference between the third panel and the fourth is what every rule in this field is asked to decide.

Read that way the dial is a length of time rather than a number between zero and one, and the question the rule asks becomes visibly strange. It asks whether a disturbance’s half-life is infinite. A half-life of sixty-nine steps in a series two hundred long is, for every practical purpose the regression cares about, the same as an infinite one: the series does not return to its level inside the sample, so whatever the theory says about its long-run behaviour, the data contains no return to see.

The same curve, on the scale a reader has. The false-positive rate against the half-life of a disturbance rather than against φ, on a log scale because the half-life runs from 0.6 steps to 69. A series whose disturbances halve in about seven steps — the middle of this axis — already produces a spurious finding 52.4% of the time. The unit root is the point at infinity on this axis, and everything a practitioner has is to the left of it.
Fig. 3 The same false-positive rate against the half-life of a disturbance rather than against φ, on a log scale because the half-life runs from half a step to sixty-nine. The unit root is the point at infinity on this axis, and everything a practitioner has is to the left of it.

That is the whole of the argument in one sentence: the rule is a statement about the point at infinity, and the data is always somewhere to the left of it.

Where the damage is and where the warning is

Putting the two curves in one figure was not a presentational choice; the region they define together is the finding, and neither curve alone shows it.

Below φ = 0.8 the test is essentially certain — it calls the series stationary 100.00% of the time at every persistence up to and including 0.8 — and the damage is real but moderate: 34.20% at 0.8, against a promised 5%. That region is bad, and it is bad in a way that has already been priced. Two stationary series are not independent observations, the standard error divides by a square root of a count that overstates the information in them, and fifty correlated observations worth about six is the same defect measured on a mean instead of on a slope. The repair there is not differencing. It is counting the observations correctly.

Above φ = 0.95 the test has mostly stopped certifying anything — 31.85% at 0.95 and 10.50% at 0.98 — and the damage is severe. That region is also bad and is not dangerous in the same way, because the analyst is being told something. A check that comes back inconclusive is a check that has done its job.

The band between them is the problem. At φ = 0.9 the test is right nearly nine times in ten and the regression it licenses is wrong more than half the time. The correct answer to the question asked and the correct answer to the question meant point in opposite directions, and nothing in either output records that two different questions were involved.

The question the test answers is whether φ is exactly one. How often the standard unit-root test refuses a unit root, against the true persistence, on series of 200 steps. The test is calibrated at a unit root and holds it: 5.8% at φ = 1, against the 5% it promises. Everywhere else it is a power curve, and it is near certain down to φ = 0.9 — where the series it certifies as stationary produce a spurious finding 52.4% of the time.
Fig. 4 How often the unit-root test refuses a unit root, against the true persistence. It holds its own size where it is calibrated — 5.80% at a genuine unit root against the 5% it promises — and everywhere else it is a power curve. Its certainty at φ = 0.9 is real and is about a different question from the one the analysis needs answered.

The arithmetic underneath, and it has a closed form

The rate is not the primitive. What the persistence does is to the statistic, and the statistic’s behaviour can be written down rather than measured.

A regression between two series divides an estimated slope by an estimated standard error, and that standard error is computed on the assumption that the residuals carry no information about each other. When both series are autoregressions they do. The variance of the resulting t statistic — which is supposed to be about one, because that is what makes ±1.96 mean 5% — is inflated by (1+φ2)/(1φ2)(1 + \varphi^2)/(1 - \varphi^2). It is φ2\varphi^2 rather than φ because both series contribute, each once.

The inflation has a closed form, and the simulation agrees with it. The variance of the t statistic for the slope between two independent AR(1) series of 200 observations, measured over 3,000 pairs at each persistence, against the curve (1 + φ²)/(1 − φ²). A t statistic is supposed to have a variance of about one; at φ = 0.9 it has 9.71, and the closed form says 9.53. The worst disagreement anywhere on the sweep is 2.9%. Two routes to the same number, and the second one is the reason the correction that follows is not fitted to the data it repairs.
Fig. 5 The measured variance of the t statistic against that curve, over three thousand independent pairs at each persistence. At φ = 0.9 the statistic has a variance of 9.71 where a t has about one, and the closed form says 9.53. The worst disagreement anywhere on the sweep is 2.7%.

The agreement is what turns the previous figures from a set of readings into an account. A rejection rate is a count; a variance with a formula that predicts it is a mechanism, and the mechanism says exactly why the curve has no step. (1+φ2)/(1φ2)(1 + \varphi^2)/(1 - \varphi^2) is a smooth increasing function on the whole of [0, 1). It rises without bound as φ approaches one, and it has no value there at all — which is the precise sense in which a unit root is a different object rather than the last point of this scale.

The correction the closed form licenses

If the damage has a formula, so does the repair: divide the statistic by the square root of the inflation before comparing it with 1.96. That is not a new procedure, and the reason for running it here is to see whether the dichotomy is doing any work at all.

It is not. With φ known, the corrected rate reads 5.23% at no persistence, 5.33% at 0.8, 4.20% at 0.9 and 4.90% at 0.95 — against raw rates of 5.23%, 34.03%, 51.50% and 65.73%. One continuous correction, applied to the continuum, holds the promise across the whole of it, and no rule anywhere in it has to decide whether a number is exactly one.

A correction with no step in it either. The same false-positive rate, uncorrected and after dividing the statistic by the square root of (1 + φ²)/(1 − φ²). With φ known the rate is held at 4.2% at a persistence where the raw rate is 51.5%, and it is held all the way to φ = 0.95 without any rule deciding anything. With φ estimated from the two series themselves — the geometric mean of their lag-one correlations — the rate reads 6.1% at 0.9 and 13.1% at 0.98. The gap between the two corrected curves is the cost of not knowing the dial.
Fig. 6 The same rate before and after the correction, with φ known and with φ estimated from the two series themselves. The known-φ curve sits on 5% across the range. The estimated-φ curve reads 6.10% at a persistence of 0.9 and 13.07% at 0.98, which is the cost of having to find the dial in the data.

Two things about that figure are worth separating, because they point in opposite directions.

The known-φ column is a proof of concept and not a method. Nobody knows φ. What it establishes is that the continuum can be handled continuously — that the two-case structure the field uses is a consequence of the instrument rather than of the problem.

The estimated-φ column is what is actually available, and it leaks. Correcting with the geometric mean of the two series’ own lag-one correlations takes the rate at φ = 0.9 from 51.50% to 6.10%, which is most of the repair, and at φ = 0.98 from 73.83% to 13.07%, which is not. The leak has a cause: a sample autocorrelation is biased downwards, badly so when φ is near one, and a correction built from an understated φ understates itself. So the practical correction degrades exactly where the persistence is hardest to measure, which is the same place the unit-root test loses its certainty — the two instruments fail in the same region for the same underlying reason.

And it stops being available at all at a unit root, because the factor it divides by has no value there. That is the narrow job the dichotomy genuinely has. It is not a rule for deciding whether to worry; it is a rule for deciding whether a correction exists.

The size of the finding moves too, and gives nothing away

A reader who has been told that significance is unreliable will reach for the effect size instead, and here that is the R2R^2. It moves the same way and offers no threshold either: the median R2R^2 between two independent series reads 0.0022 at no persistence, 0.0214 at φ = 0.9, 0.0710 at 0.98 and 0.1740 at a unit root.

The size of a finding that is not there. The median R² of a regression between two independent series, against how persistent each one is. Every pair here is generated independently, so the honest answer at every φ is zero. At φ = 0 the median reads 0.0022, at φ = 0.9 it reads 0.021, and at a unit root 0.174. The quantity moves smoothly across the range, which is what makes it useless as a warning: there is no value of it that says the relationship is manufactured.
Fig. 7 The median R2R^2 of a regression between two independent series, against how persistent each one is. Every pair is generated independently, so the honest answer at every setting is zero. A seventeen-per-cent R2R^2 between two series with nothing to do with each other is not an outlier of the distribution; it is the middle of it.

An R2R^2 of 0.17 is the kind of number that gets called modest but real. It is neither: the truth is exactly zero at every point on that curve, and the number is a function of the persistence alone. There is no value of R2R^2 that means “this cannot have been manufactured”, in the same way there is no value of p, because both statistics are computed on the assumption the persistence destroys.

The same is true of any diagnostic read off the fitted line. The residuals of a spurious regression between persistent series are themselves persistent, so a Durbin–Watson statistic reads low; but it reads low at every φ above about a half, and its reading at 0.9 is not qualitatively different from its reading at 0.7. Everything that could serve as a warning is a continuous function of the same dial, so nothing among them contains a threshold that the dial itself does not.

Why the test cannot be blamed for this

The unit-root test is not defective and it is not being used wrongly, which is what makes this worth an essay rather than a footnote. It is calibrated at φ = 1, it holds its size there — 5.80%, against a nominal 5%, over two thousand series of two hundred observations — and away from there it is a power curve. A power curve rising towards certainty as φ falls is exactly what a well-behaved test looks like. Its certainty at φ = 0.9 is the test working.

The critical value it is read against is −2.88, simulated here rather than looked up, because the statistic under a unit root is not a t and no table of t values contains its distribution. That is the same construction the next test in this field needs and the same reason there is no table for it: when the regressor is not stationary the usual asymptotics do not apply, and a critical value has to be generated from the null it is supposed to describe.

The mismatch is between what the test is about and what the analysis needs. The test asks a question whose answer is a point — is φ exactly one — and the analysis needs a question whose answer is a magnitude. A test of a point null can be exactly right about the point and carry no information about the magnitude, and this is the case where that gap is expensive rather than academic.

What a defensible reading looks like

Three things follow from the sweep, and the first is the one that takes discipline.

A clean unit-root test is not a licence to regress levels. It is the absence of one specific catastrophe. What the sweep says is that at a persistence the test happily certifies, a regression between unrelated series is significant more often than not; so the check that was run has ruled out the worst case and not the common one.

The quantity to report is the persistence, not the verdict. A lag-one correlation of 0.9 with two hundred observations is a statement a reader can price from this figure, in the same way a run length is a statement and a declustered count is a decision. “The series passed an augmented Dickey–Fuller test” is not, because the same sentence covers φ = 0.3 and φ = 0.9, whose false-positive rates differ by a factor of seven. This is the same complaint the threshold essay makes about a knob reported as a decision: a number that was continuous before somebody decided about it is more informative than the decision.

And the repair is a correction, not a difference. Two stationary series with a persistence of 0.9 carry far less independent information than their count suggests — the same arithmetic that turns fifty observations into about six — and the statistic can be deflated for that directly, without discarding the levels. It recovers most of what is lost even with the dial estimated from the data. Differencing is the repair for the other case, and applying it here is applying a remedy to a series that did not need one, which doubles the variance and installs a correlation the data never had.

What is claimed here and what is not

The sweep is of one process family. Every series above is a first-order autoregression, so “persistence” means one number. Real series carry seasonal structure, moving-average components and longer memory than any AR(1) reaches, and the mapping from those to a single φ is not exact. What the sweep establishes is that the false-positive rate is a continuous function of how long a disturbance survives; it does not establish that φ is the right summary of how long a disturbance survives in a series that is not an AR(1).

The unit-root test used is the simplest one. No lags, no trend term, and a critical value simulated at the sample size used. Augmenting it with lagged differences would change the numbers — generally lowering the power, because each lag costs degrees of freedom — and would not change their ordering, because the augmentation does not alter what the test is a test of. The claim being made is about the target of the hypothesis, not about the efficiency of any particular version of it.

The crossing region depends on the sample length. At two hundred observations the test is near-certain down to φ = 0.9. At fifty it is much less so, and the dangerous band sits lower; at a thousand it sits higher, because a longer sample can distinguish 0.98 from 1. The slider on the first figure moves it. What does not move is the existence of a band where the test is confident and the damage is large, because the test’s power and the regression’s failure rate are driven by different features of the same process and there is no reason for their transitions to coincide.

The closed form is asymptotic, and the correction inherits that. (1+φ2)/(1φ2)(1 + \varphi^2)/(1 - \varphi^2) is the limiting inflation, and at two hundred observations with φ above 0.95 the statistic’s actual variance is a little smaller than the limit — so the known-φ correction becomes mildly conservative rather than mildly liberal, reading 2.97% at φ = 0.98. That is the safe direction and it is reported rather than smoothed, because a reader entitled to know the correction works is also entitled to know where it stops being exact.

The estimated-φ correction is one of several. The lag-one correlation is the simplest estimate of the dial and not the best; a Yule–Walker fit at a higher order, or a long-run variance estimated with a taper, would each do better at high persistence and each brings its own tuning decision. Nothing here says the leak at φ = 0.98 is irreducible. What it says is that the leak exists for the obvious estimator, and that the direction of the error is predictable from how a sample autocorrelation behaves.

And the 5% at φ = 0 is the calibration. If the arithmetic produced anything other than about one in twenty at no persistence, every number above would be measuring a bug in the regression rather than an effect of the dependence. It reads 4.85% over two thousand pairs, which is within a standard error of nominal, and that reading is what licenses the rest of the column.

Still open: which repair the middle of the dial wants

The sweep says where the damage is. It does not say what to do in the band it identifies, because the two repairs the field has are each built for an end of the dial: differencing is for the unit root and an effective-sample-size correction is for a mildly dependent stationary series. At φ = 0.9 neither is obviously right, and the question of which loses less — a correction that assumes stationarity when the persistence is nearly total, or a difference that discards a level the series genuinely has — is a measurement nobody has made here.

The correction above sharpens that question rather than settling it. It shows a continuous repair is possible with the dial in hand, and it shows the estimate of the dial degrading in the band that matters. Whether a better estimator of the persistence closes the gap — and how much of the leak at φ = 0.98 is the estimator rather than the approximation — is one measurement, and it is the one that would decide whether the dichotomy has any job left beyond the unit root itself.

There is a second question underneath it, and it is the one that decides whether the first is worth asking. The two worlds this field has been treating as one continuum are not actually one: a series can trend because it accumulates its own disturbances, or because it is a straight line with a wobble around it, and those are different objects that produce similar pictures. The repairs are different too, and applying either to the wrong world does its own damage — which is what each repair costs in the world it was not built for, and where this argument goes next.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Augmented dickey fullerAutocorrelationClosed formCritical valueDeterministic trendEffective sample sizeFalse positiveHalf-lifeLong-run varianceMonte CarloNear-unit rootRandom walkSpurious regressionStationarityStatistical powerTrend-stationaryUnit rootVariance inflation