The cliff that is a slope
Worth reading first: Two walks and a finding · Three series and a count.
The field opened with two independent random walks and a slope that was called significant three times in four, and it has treated that as a statement about random walks ever since. The repair it offers is a rule with a switch in it: test each series for a unit root, and if the test finds one, difference before regressing.
The switch asks whether one number is exactly one. What the damage depends on is how close that number is, and the two are not the same question. Swept across the persistence of the two series, the rate at which a regression between them is called significant reads 4.85% at no persistence, 13.80% at a lag-one correlation of 0.5, 34.20% at 0.8, 52.45% at 0.9 and 83.40% at a unit root. There is no value at which the curve turns. The rule placed on it has a step in it because a rule has to decide something, and the quantity it decides about does not.
The second curve is the one that makes this more than a remark about continuity. At a persistence of 0.9 the unit-root test is doing its job: the series genuinely is stationary, it genuinely has a level to return to, and the test says so 87.15% of the time. An analyst who runs the check, reads a clean result, and regresses the levels is following the procedure correctly at every step, and the regression they then run between two unrelated series is significant more often than not.
A dial nobody put a mark on
The processes being swept are first-order autoregressions. Each observation is φ times the last one plus a fresh disturbance, so φ is how much of yesterday survives into today, and a disturbance decays by a factor of φ every step. At φ = 0.5 half of it is gone in one step; at φ = 0.9 it takes 6.58 steps to halve; at φ = 0.99, 68.97. At φ = 1 it never goes, which is what a random walk is — and a series that never returns to a level is not the same thing as a series that has not converged yet, though the two are routinely described in the same words.
Read that way the dial is a length of time rather than a number between zero and one, and the question the rule asks becomes visibly strange. It asks whether a disturbance’s half-life is infinite. A half-life of sixty-nine steps in a series two hundred long is, for every practical purpose the regression cares about, the same as an infinite one: the series does not return to its level inside the sample, so whatever the theory says about its long-run behaviour, the data contains no return to see.
That is the whole of the argument in one sentence: the rule is a statement about the point at infinity, and the data is always somewhere to the left of it.
Where the damage is and where the warning is
Putting the two curves in one figure was not a presentational choice; the region they define together is the finding, and neither curve alone shows it.
Below φ = 0.8 the test is essentially certain — it calls the series stationary 100.00% of the time at every persistence up to and including 0.8 — and the damage is real but moderate: 34.20% at 0.8, against a promised 5%. That region is bad, and it is bad in a way that has already been priced. Two stationary series are not independent observations, the standard error divides by a square root of a count that overstates the information in them, and fifty correlated observations worth about six is the same defect measured on a mean instead of on a slope. The repair there is not differencing. It is counting the observations correctly.
Above φ = 0.95 the test has mostly stopped certifying anything — 31.85% at 0.95 and 10.50% at 0.98 — and the damage is severe. That region is also bad and is not dangerous in the same way, because the analyst is being told something. A check that comes back inconclusive is a check that has done its job.
The band between them is the problem. At φ = 0.9 the test is right nearly nine times in ten and the regression it licenses is wrong more than half the time. The correct answer to the question asked and the correct answer to the question meant point in opposite directions, and nothing in either output records that two different questions were involved.
The arithmetic underneath, and it has a closed form
The rate is not the primitive. What the persistence does is to the statistic, and the statistic’s behaviour can be written down rather than measured.
A regression between two series divides an estimated slope by an estimated standard error, and that standard error is computed on the assumption that the residuals carry no information about each other. When both series are autoregressions they do. The variance of the resulting t statistic — which is supposed to be about one, because that is what makes ±1.96 mean 5% — is inflated by . It is rather than φ because both series contribute, each once.
The agreement is what turns the previous figures from a set of readings into an account. A rejection rate is a count; a variance with a formula that predicts it is a mechanism, and the mechanism says exactly why the curve has no step. is a smooth increasing function on the whole of [0, 1). It rises without bound as φ approaches one, and it has no value there at all — which is the precise sense in which a unit root is a different object rather than the last point of this scale.
The correction the closed form licenses
If the damage has a formula, so does the repair: divide the statistic by the square root of the inflation before comparing it with 1.96. That is not a new procedure, and the reason for running it here is to see whether the dichotomy is doing any work at all.
It is not. With φ known, the corrected rate reads 5.23% at no persistence, 5.33% at 0.8, 4.20% at 0.9 and 4.90% at 0.95 — against raw rates of 5.23%, 34.03%, 51.50% and 65.73%. One continuous correction, applied to the continuum, holds the promise across the whole of it, and no rule anywhere in it has to decide whether a number is exactly one.
Two things about that figure are worth separating, because they point in opposite directions.
The known-φ column is a proof of concept and not a method. Nobody knows φ. What it establishes is that the continuum can be handled continuously — that the two-case structure the field uses is a consequence of the instrument rather than of the problem.
The estimated-φ column is what is actually available, and it leaks. Correcting with the geometric mean of the two series’ own lag-one correlations takes the rate at φ = 0.9 from 51.50% to 6.10%, which is most of the repair, and at φ = 0.98 from 73.83% to 13.07%, which is not. The leak has a cause: a sample autocorrelation is biased downwards, badly so when φ is near one, and a correction built from an understated φ understates itself. So the practical correction degrades exactly where the persistence is hardest to measure, which is the same place the unit-root test loses its certainty — the two instruments fail in the same region for the same underlying reason.
And it stops being available at all at a unit root, because the factor it divides by has no value there. That is the narrow job the dichotomy genuinely has. It is not a rule for deciding whether to worry; it is a rule for deciding whether a correction exists.
The size of the finding moves too, and gives nothing away
A reader who has been told that significance is unreliable will reach for the effect size instead, and here that is the . It moves the same way and offers no threshold either: the median between two independent series reads 0.0022 at no persistence, 0.0214 at φ = 0.9, 0.0710 at 0.98 and 0.1740 at a unit root.
An of 0.17 is the kind of number that gets called modest but real. It is neither: the truth is exactly zero at every point on that curve, and the number is a function of the persistence alone. There is no value of that means “this cannot have been manufactured”, in the same way there is no value of p, because both statistics are computed on the assumption the persistence destroys.
The same is true of any diagnostic read off the fitted line. The residuals of a spurious regression between persistent series are themselves persistent, so a Durbin–Watson statistic reads low; but it reads low at every φ above about a half, and its reading at 0.9 is not qualitatively different from its reading at 0.7. Everything that could serve as a warning is a continuous function of the same dial, so nothing among them contains a threshold that the dial itself does not.
Why the test cannot be blamed for this
The unit-root test is not defective and it is not being used wrongly, which is what makes this worth an essay rather than a footnote. It is calibrated at φ = 1, it holds its size there — 5.80%, against a nominal 5%, over two thousand series of two hundred observations — and away from there it is a power curve. A power curve rising towards certainty as φ falls is exactly what a well-behaved test looks like. Its certainty at φ = 0.9 is the test working.
The critical value it is read against is −2.88, simulated here rather than looked up, because the statistic under a unit root is not a t and no table of t values contains its distribution. That is the same construction the next test in this field needs and the same reason there is no table for it: when the regressor is not stationary the usual asymptotics do not apply, and a critical value has to be generated from the null it is supposed to describe.
The mismatch is between what the test is about and what the analysis needs. The test asks a question whose answer is a point — is φ exactly one — and the analysis needs a question whose answer is a magnitude. A test of a point null can be exactly right about the point and carry no information about the magnitude, and this is the case where that gap is expensive rather than academic.
What a defensible reading looks like
Three things follow from the sweep, and the first is the one that takes discipline.
A clean unit-root test is not a licence to regress levels. It is the absence of one specific catastrophe. What the sweep says is that at a persistence the test happily certifies, a regression between unrelated series is significant more often than not; so the check that was run has ruled out the worst case and not the common one.
The quantity to report is the persistence, not the verdict. A lag-one correlation of 0.9 with two hundred observations is a statement a reader can price from this figure, in the same way a run length is a statement and a declustered count is a decision. “The series passed an augmented Dickey–Fuller test” is not, because the same sentence covers φ = 0.3 and φ = 0.9, whose false-positive rates differ by a factor of seven. This is the same complaint the threshold essay makes about a knob reported as a decision: a number that was continuous before somebody decided about it is more informative than the decision.
And the repair is a correction, not a difference. Two stationary series with a persistence of 0.9 carry far less independent information than their count suggests — the same arithmetic that turns fifty observations into about six — and the statistic can be deflated for that directly, without discarding the levels. It recovers most of what is lost even with the dial estimated from the data. Differencing is the repair for the other case, and applying it here is applying a remedy to a series that did not need one, which doubles the variance and installs a correlation the data never had.
What is claimed here and what is not
The sweep is of one process family. Every series above is a first-order autoregression, so “persistence” means one number. Real series carry seasonal structure, moving-average components and longer memory than any AR(1) reaches, and the mapping from those to a single φ is not exact. What the sweep establishes is that the false-positive rate is a continuous function of how long a disturbance survives; it does not establish that φ is the right summary of how long a disturbance survives in a series that is not an AR(1).
The unit-root test used is the simplest one. No lags, no trend term, and a critical value simulated at the sample size used. Augmenting it with lagged differences would change the numbers — generally lowering the power, because each lag costs degrees of freedom — and would not change their ordering, because the augmentation does not alter what the test is a test of. The claim being made is about the target of the hypothesis, not about the efficiency of any particular version of it.
The crossing region depends on the sample length. At two hundred observations the test is near-certain down to φ = 0.9. At fifty it is much less so, and the dangerous band sits lower; at a thousand it sits higher, because a longer sample can distinguish 0.98 from 1. The slider on the first figure moves it. What does not move is the existence of a band where the test is confident and the damage is large, because the test’s power and the regression’s failure rate are driven by different features of the same process and there is no reason for their transitions to coincide.
The closed form is asymptotic, and the correction inherits that. is the limiting inflation, and at two hundred observations with φ above 0.95 the statistic’s actual variance is a little smaller than the limit — so the known-φ correction becomes mildly conservative rather than mildly liberal, reading 2.97% at φ = 0.98. That is the safe direction and it is reported rather than smoothed, because a reader entitled to know the correction works is also entitled to know where it stops being exact.
The estimated-φ correction is one of several. The lag-one correlation is the simplest estimate of the dial and not the best; a Yule–Walker fit at a higher order, or a long-run variance estimated with a taper, would each do better at high persistence and each brings its own tuning decision. Nothing here says the leak at φ = 0.98 is irreducible. What it says is that the leak exists for the obvious estimator, and that the direction of the error is predictable from how a sample autocorrelation behaves.
And the 5% at φ = 0 is the calibration. If the arithmetic produced anything other than about one in twenty at no persistence, every number above would be measuring a bug in the regression rather than an effect of the dependence. It reads 4.85% over two thousand pairs, which is within a standard error of nominal, and that reading is what licenses the rest of the column.
Still open: which repair the middle of the dial wants
The sweep says where the damage is. It does not say what to do in the band it identifies, because the two repairs the field has are each built for an end of the dial: differencing is for the unit root and an effective-sample-size correction is for a mildly dependent stationary series. At φ = 0.9 neither is obviously right, and the question of which loses less — a correction that assumes stationarity when the persistence is nearly total, or a difference that discards a level the series genuinely has — is a measurement nobody has made here.
The correction above sharpens that question rather than settling it. It shows a continuous repair is possible with the dial in hand, and it shows the estimate of the dial degrading in the band that matters. Whether a better estimator of the persistence closes the gap — and how much of the leak at φ = 0.98 is the estimator rather than the approximation — is one measurement, and it is the one that would decide whether the dichotomy has any job left beyond the unit root itself.
There is a second question underneath it, and it is the one that decides whether the first is worth asking. The two worlds this field has been treating as one continuum are not actually one: a series can trend because it accumulates its own disturbances, or because it is a straight line with a wobble around it, and those are different objects that produce similar pictures. The repairs are different too, and applying either to the wrong world does its own damage — which is what each repair costs in the world it was not built for, and where this argument goes next.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- How slow a return a sample can see — both name critical value, half-life, monte carlo, random walk, spurious regression, statistical power
- Which series goes on the left — both name critical value, random walk, spurious regression, stationarity, unit root
- A detector built for the ordering — both name autocorrelation, critical value, monte carlo, statistical power
- An order that spends the error rate — both name closed form, false positive, monte carlo, statistical power
- Draws that repeat each other — both name autocorrelation, critical value, effective sample size, monte carlo
- How long a block a multiplier shares — both name autocorrelation, critical value, effective sample size, monte carlo
Named objects
A flat tag is an object no other essay names yet.
Augmented dickey fullerAutocorrelationClosed formCritical valueDeterministic trendEffective sample sizeFalse positiveHalf-lifeLong-run varianceMonte CarloNear-unit rootR²Random walkSpurious regressionStationarityStatistical powerTrend-stationaryUnit rootVariance inflation