The repair that keeps the question
Worth reading first: Two walks and a finding · Three series and a count.
A series that goes up can go up for two reasons, and the distinction is not a technicality about models. It can accumulate its own disturbances, so that whatever happened to it in March is still in its level in December and always will be. Or it can be a straight line with a wobble around it, so that March’s disturbance has faded and the level is where the line says it should be.
The two produce pictures that are frequently indistinguishable, and the treatments they want are opposite. The first wants differencing; the second wants the line subtracted. The literature’s standard advice is to find out which world a series is in and then apply the matching repair, and everything in that sentence is reasonable.
Run it as a table and the advice comes apart. Six cells — two worlds by three treatments — in which every pair of series is generated independently, so 5% is the correct answer in all six:
The treatment that controls the rate does so in both worlds, and the treatment built for one of the worlds controls it in neither. That is the opposite shape to the advice.
What the two worlds look like
The reason the advice is hard to follow is not that the two worlds are subtle. It is that they are subtle in a sample.
Nothing visible separates them, and nothing should: the claim that distinguishes them is a claim about the limit. A random walk’s variance grows without bound and a trend-stationary series’ does not, but “without bound” is a statement about a horizon no sample reaches. Two hundred observations of each contain a rising line and some wobble in both cases.
This is the same shape as the continuum the previous essay measured, arriving from a different direction. There the dichotomy was φ = 1 against φ near 1; here it is accumulation against a line. In both cases the rule is a claim about the point at infinity and the evidence is a finite stretch of series.
Why a line does not remove the problem
Subtracting a fitted line is the natural repair and the intuition behind it is correct as far as it goes: two series that both trend upward will correlate because they both trend upward, so remove the trends. What that intuition misses is that the trend is not the only thing making them correlate.
In the trend-stationary world the line is genuinely there and removing it is genuinely right, and what is left over is an AR(1) at φ = 0.8. A regression between two such residual series is a regression between two persistent series, which is the case the previous essay priced at 34.20%. The table’s cell reads 33.45%. The repair did exactly what it promised and the false-positive rate stayed at a third, because removing the trend removed the trend and not the persistence.
In the random-walk world it is worse, and for a different reason. A random walk has no trend to remove; a line fitted to one is fitting noise, and the residuals of that fit are a detrended walk, which is still very nearly a walk — lag-one correlation 0.94, and 57.01% of the original variance still present. The cell reads 74.20% against 82.90% untreated. The line took eight points off a seventy-eight-point problem.
A line fitted to a walk is not a neutral operation
The 74.20% cell has a cause worth separating out, because it is the reason “detrend first” is not a harmless default.
A random walk has no trend. Its expected level at every horizon is where it started, and its rise in any particular realisation is accumulated disturbance rather than direction. Fit a least-squares line in time to one anyway and the line is called significant 90.60% of the time at two hundred observations — and the rate rises with the sample, to 95.93% at eight hundred.
The pair of readings is the diagnosis. A trend genuinely coming into focus would show a rising rate because the was rising — more data, a clearer signal, a larger share of variance explained. Here the share is flat at 0.44 at every length. The certainty is rising and the apparent size of the effect is not, which is what a wrong reference distribution looks like from the outside.
So an analyst who detrends because the series “obviously has a trend” has, in the random-walk world, been told so by a test that says so about nine walks in ten. The operation is not applied to a series that turned out not to need it; it is applied to a series that was diagnosed as needing it by an instrument that would have said the same of almost any walk.
That also explains the other end of the table. Two independent trend-stationary series share a deterministic trend of the same sign, and the median between their levels is 0.866 — not a tail event, the middle of the distribution. The 100.00% cell is not the regression failing; it is the regression reading a real common trend correctly and the analyst reading it as a relation between the series.
What each treatment leaves behind
A rate that returns to 5% says a treatment removed whatever was inflating it. It says nothing about what the treatment left. Both are worth measuring, and measuring them on the series rather than on a test is what separates a repair from a deletion.
The two differenced rows are the only ones that leave something close to independent observations, and the variance column says what that cost. Differencing a trend-stationary series keeps 3.13% of its variance. What was discarded is not error: the trend was the largest thing in that series, and it was the thing a study of a trending series was about.
The slight negative correlation in those rows — −0.0971 and −0.0055 — is the signature of differencing a series that did not need it, showing up here as a much milder version of the −0.5 an already-stationary series acquires, because these series were nowhere near stationary to begin with.
The difference is a change of estimand, not a loss of precision
This is the sentence the table is really about, and it is easy to slide past.
Differencing does not estimate the same quantity less precisely. It estimates a different quantity. The regression of Δy on Δx asks whether a change in one series accompanies a change in the other within the same period. The regression of y on x asks whether the two series stand in a relation of levels. Those are different questions with different answers, and a study that differenced to control its false-positive rate has answered the second question by not asking it.
The clearest way to see that is the mean. The average of a differenced series of length n is identically (last − first)/(n − 1) — every interior observation cancels — so a study working in differences has, for the purposes of the level, two observations rather than two hundred. That identity holds exactly on every realisation and is not an approximation that improves with the sample.
So the table has no cell in it that both controls the rate and keeps the question. A reader who wanted to know whether two trending quantities stand in a relation has three options, and one of them is not on the table:
- regress the levels and report a rate between a third and all, which is not an option;
- difference, and answer a different question honestly, which is what the field’s advice amounts to in practice;
- regress the levels and correct the statistic for the persistence, which is where the previous essay ended and which needs the persistence to be estimable.
There is a fourth possibility the table does not contain and that is worth naming, because a reader will reach for it: regress the levels and include time as a covariate, so that the common trend is removed inside the regression rather than before it. Arithmetically that is the detrended cell — the coefficient on x is identical to the one from regressing detrended y on detrended x — so it inherits the same 74.20% and 33.45%. What changes is only where the operation is written down, and a false-positive rate is not affected by notation. The appeal of doing it in one step is that the standard errors are computed with the right degrees of freedom; the reason it does not help is that the degrees of freedom were never the problem.
The test that is supposed to choose
All of the advice rests on being able to tell which world a series is in, so it is worth asking how well that can be done. The instrument is the same unit-root test with a trend term added, read against its own simulated critical value of −3.45.
At two hundred observations with a disturbance that halves in three steps, the test is essentially perfect. That is a genuine result and it is the easy case: a wobble with a three-step memory around a visible line is not anyone’s idea of a random walk.
The power collapses as the disturbance’s own memory lengthens, and it collapses into the region where the choice matters most. Two features of that curve are worth reading separately. It is not a gradual decline: between φ = 0.8 and φ = 0.95 the test goes from 99.75% right to 17.85% right, which is most of its useful range spent inside a band of persistence that a reader would describe with the same two words at both ends. And the floor it reaches is its own size — 6.85% at φ = 0.98 — meaning the test is no longer responding to the difference between the worlds at all, rather than responding weakly.
A disturbance with a half-life of 13.51 steps is not a random walk: it returns, the return is a real feature of the process, and a study that differenced it away has thrown out a trend that was genuinely there. The test identifies it correctly 17.85% of the time.
That is the same failure the previous essay found in the no-trend version of the test, one step further on: there the test was confident and the analysis it licensed was wrong, and here the test simply stops answering. Adding the trend term is what costs it — the same series without the trend term is called stationary 31.85% of the time at φ = 0.95, and the extra parameter has to be paid for out of the same two hundred observations.
What a defensible procedure looks like
Do not choose the treatment by a test that cannot see the difference. Where the disturbance is mildly persistent the test is near-perfect and the choice is easy; where it is strongly persistent the test is near-useless and the choice is exactly the one being made badly. A procedure whose reliability is highest where it is least needed should not be the procedure the conclusion rests on.
Say which quantity is being reported. A coefficient from a differenced regression and a coefficient from a levels regression are different numbers about different things, and the practice of reporting either as “the relationship between x and y” is what allows the choice of treatment to be invisible in the result. This is the same discipline an interval that answers a different question needs, arriving as a question about a regression rather than about a bound.
Treat a diagnosed trend as a finding about the instrument, not only about the series. A significant fitted line is evidence of a trend only against a null the data may not satisfy, and on a walk it arrives nine times in ten. The useful companion reading is the share of variance it claims and whether that share moves with the sample, because a real trend and a manufactured one separate there and not in the p-value. It is the same distinction a searched break has to be charged for: the statistic is fine and the distribution it is read against is not.
And prefer the treatment that keeps the estimand where the rate can be handled another way. The correction from the previous essay is available in the trend-stationary world, because a detrended series there is a stationary AR(1) whose persistence can be estimated. It is not available in the random-walk world, where the factor it divides by has no value. So the honest statement of the advice is narrower than the usual one: the world matters, but it matters because it decides whether a correction exists — not because each world has its own treatment.
What is claimed here and what is not
One slope and one persistence. The trend-stationary world here has a slope of 0.1 against a disturbance standard deviation of 1, so over two hundred observations the trend spans twenty and the wobble spans about three. A weaker trend is harder to find and pushes the chooser’s power down at every persistence; a stronger one pushes it up. The 33.45% cell does not move with the slope, because the line is removed before the regression either way, but the power figures do, and they should be read as “at a trend this visible” rather than as a property of the test.
The critical values are simulated at each specification. The no-trend test is read against −2.88 and the trend version against −3.45, both generated under the null at the sample size used rather than looked up. Reading either against the other, or against a t table, would change every power figure here — which is the failure the next test in this field was built to avoid.
Six cells is not a proof that no fourth treatment exists. The table prices the three treatments the field actually uses. Detrending with a nonlinear trend, filtering, and regressing in levels with a correction are all absent from it, and the third is the one the previous essay measured separately. The claim is about the three treatments on the table, not about the space of everything one could do.
The trend-hunting figure uses the ordinary t test on the slope. That is deliberate — it is the test a reader’s software prints — and it is not the state of the art. A trend test with a reference distribution built for an integrated series exists and behaves properly, which is exactly the fix this field keeps arriving at: the statistic is not wrong, the table it is read against is. Nothing above says a well-calibrated trend test is unavailable; it says that the one in general use manufactures the finding that licenses the treatment.
And the 100.00% is not a rounding of 99.5%. Two independent trend-stationary series with slopes of the same sign are called significantly related in every one of two thousand replications, which is what a deterministic trend common to both does: it is not noise, it does not average out, and the regression is reading it correctly. Removing it is genuinely necessary; the essay’s point is that it is not sufficient.
Still open: how slow a return can be and still be seen
Both essays so far have been about pairs with nothing between them, and the rate at which a procedure invents a relationship. The mirror question is what happens to a pair that genuinely has one, and it is the question the field’s one positive result depends on — the case where a levels regression is right because the two series are tied to each other and return to their relation after a shock.
The measurements above suggest what will be found and do not make it. Everything in this field degrades as a return gets slower, and a pair tied together by a mechanism that undoes a small fraction of a disagreement each period is a pair whose return is slow by construction. So there should be a speed below which the mechanism is there and the sample cannot see it — and the sensible way to state it is not as a coefficient but as a half-life, which is a length of time a reader can compare with the length of their data. What that boundary is, and how it moves with the sample, is the next thing this field has to measure.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The model that corrects its error — both name autocorrelation, differencing, spurious regression, stationarity, unit root
- Which series goes on the left — both name random walk, spurious regression, stationarity, unit root
- An order that spends the error rate — both name closed form, false positive, statistical power
- Counting what is still wandering — both name differencing, random walk, stationarity
- The check before the standard error — both name autocorrelation, differencing, statistical power
- The clustering the tail has — both name autocorrelation, closed form, stationarity
Named objects
A flat tag is an object no other essay names yet.
Augmented dickey fullerAutocorrelationClosed formDeterministic trendDetrendingDifferencingEstimandFalse positiveNear-unit rootOver-differencingR²Random walkSpurious regressionStationarityStatistical powerTrend-stationaryUnit root