The check before the standard error
Worth reading first: The observations that repeat each other · Two walks and a finding.
Three essays of this field end with a recommendation: check the dependence before trusting a standard error. A recommendation whose error rate has not been counted is exactly what this site does not do, so this essay counts it.
The check
Compute the correlation between the series and itself shifted by one step. Under independence, that sample autocorrelation is approximately normal with standard deviation 1/√n, so the test is
|r₁| × √n > 1.96
and the band drawn on every correlogram in this field is that threshold, at ±1.96/√n.
It is one line, needs no model, and is available for any ordered data. What follows is what it is worth.
Two details of the definition matter and are easy to get wrong. The correlation is computed against the series’ own mean rather than against separate means for the two shifted copies, which keeps the estimate bounded between −1 and 1 and biases it slightly towards zero. And the divisor is n rather than n − 1 at every lag, so the correlations at long lags — computed from few pairs — are pulled towards zero rather than allowed to wander wildly. Both conventions trade a little bias for a great deal of stability, and both are why a sample correlogram of an independent series looks tidier than a naive calculation would make it.
Its size and its power
Run on four thousand independent series of fifty observations, the check flags 4.25% of them — close to the 5% it should, slightly conservative, which is the calibration that makes everything else here meaningful.
Run on four thousand series with a lag-one correlation of 0.5, it flags 90.0%. Strong dependence is caught almost always.
Run at φ = 0.2, it flags 20.8%. Weak dependence goes undetected four times out of five.
What the check finds, and at what
The test’s power has a closed form worth writing down, because it says which dependences this recommendation actually protects against.
Under a first-order process at φ, the sample autocorrelation is centred near φ with a standard deviation of about , so the non-centrality is .
At fifty observations:
- φ = 0.5: power 98%
- φ = 0.3: power 60%
- φ = 0.2: power 30%
At two hundred observations the middle row rises to 99%.
Which is the wrong way round
Set those against what each φ is doing to the standard error the check is protecting.
A lag-one correlation of 0.2 inflates the variance of a mean by , so the standard error is understated by 22% — and the check misses it seven times in ten.
At 0.3 the standard error is understated by 36%, and the check misses it two times in five.
At 0.5 the understatement is 73% and the check almost never misses.
So the check is reliable exactly where the damage is obvious and unreliable where it is merely serious. A twenty-two per cent understatement of every standard error in an analysis is a real problem — it turns a 95% interval into an 89% one — and at fifty observations this test finds it less than a third of the time.
The estimate’s own bias makes it slightly worse. The sample autocorrelation is pulled towards zero by roughly , which at φ = 0.2 and n = 50 is about 0.032 — a sixth of the quantity being estimated, removed before the test sees it.
None of that is an argument against running the check. It is an argument for reading a quiet result as what it is: at these sample sizes, a non-detection rules out a large dependence and says almost nothing about a moderate one.
The gap that matters
The 20.8% would be unremarkable if weak dependence were harmless. It is not.
At φ = 0.2 and fifty observations, the variance of the mean is inflated by 1.49, the fifty observations are worth 33, and a 95% interval covers 88.6%.
So there is a band of dependence — roughly φ between 0.15 and 0.35 at this sample size — that the recommended check misses most of the time and that costs six or seven points of coverage. That is worse than most of the failures this site measures elsewhere, and it is invisible to the diagnostic that is supposed to catch it.
Two things follow, and neither is “use a better test”.
A passing check is weak evidence of independence. It rules out strong dependence and says little about moderate dependence, in the same way that a non-significant result at low power says little about a moderate effect. This is the same argument as reading a null result, applied to a diagnostic instead of to a finding.
And the decision should not rest on the check alone. Whether the observations could plausibly be dependent is a question about how the data was collected, answerable before any test is run. Data arriving in time order, from the same unit, or from neighbouring locations should be assumed dependent unless there is a reason it is not, and the check is a way of discovering how bad it is rather than a way of deciding whether it exists.
The power depends on the series length, sharply
The 90.0% above is at fifty observations, and the check’s power moves quickly with n.
At twenty observations it flags a correlation of 0.5 only 36.2% of the time. At fifty, 90.0%. At a hundred, 99.7%.
The twenty-observation number is the uncomfortable one, because short series are exactly where dependence does the most damage to an interval and where the check is least able to find it. A twenty-point series with φ = 0.5 has an interval that badly undercovers and a diagnostic that will say nothing about it two times in three.
There is no repair for that within the data. It is a statement about how much information twenty correlated observations contain, and the only fixes are outside the analysis: collect more, or use what is known about the measurement process instead of asking the data to reveal it.
Reading the whole correlogram, not the first bar
The lag-one test is the summary; the correlogram is the instrument, and its shape distinguishes things the first bar cannot.
Geometric decay — each bar a fixed fraction of the one before — is the signature of an autoregressive process, and the fraction is φ. The figures in this field draw φᵏ over the bars for exactly this reason: the theoretical shape is a curve the measurement is required to follow, which is the two-routes rule applied to a diagnostic.
A single spike then nothing is a moving-average process, where each observation shares one shock with its neighbour and nothing with anything further away. A negative spike of this shape at lag one is the fingerprint of over-differencing.
Slow, nearly linear decay that is still substantial at lag fifteen or twenty suggests the series is not stationary at all, and no correction based on an effective sample size will apply.
And a spike at a seasonal lag — twelve, seven, four — with little elsewhere is a repeating pattern rather than memory, and is handled by seasonal terms rather than by any of this field’s corrections.
Four shapes, four different treatments, and all four have a lag-one bar that could be identical. That is the argument for looking at the plot rather than at the number that summarises it.
Durbin–Watson, which is the same number
Regression output frequently prints a Durbin–Watson statistic, and it is worth being able to read it because it is the same quantity in different units.
DW ≈ 2(1 − r₁)
So 2 means no autocorrelation, 0 means perfect positive autocorrelation, and 4 means perfect alternation. A DW of 1.0 corresponds to r₁ ≈ 0.5, which is the case the check above catches nine times in ten; a DW of 1.6 corresponds to r₁ ≈ 0.2, which is the case it misses four times in five.
The rule of thumb that Granger and Newbold proposed for spurious regressions reads directly off it: an R² larger than the Durbin–Watson statistic means a well-fitting model with heavily autocorrelated residuals, which is the signature of a relationship produced by the dependence rather than by anything in the data.
Its main limitation is that it looks only at lag one. A series with quarterly or weekly structure can have a Durbin–Watson near 2 and enormous correlation at the seasonal lag, which is why the correlogram — all the lags, not one of them — is the better instrument.
Where to look, which is not where residuals are usually plotted
A regression’s residuals are conventionally plotted against fitted values, and that plot cannot show autocorrelation at all: it discards the ordering, which is the only thing dependence lives in.
The plot that shows it is residuals against time, or against whatever ordered the data — sequence number, position, distance along a transect. Autocorrelated residuals appear as runs: long stretches above the line, then long stretches below.
That is a real limitation of the standard diagnostic set, and it is a specific instance of the point the residual-plot essay makes about calibration: a reader has to know what the plot looks like when nothing is wrong before a run of six points above the line means anything. The correlogram is the calibrated version of the same look, with the ±1.96/√n band supplying what the eye cannot.
The practical instruction is short. If the dataset has an order, plot the residuals in it. A column that no step of the analysis referred to has made the assumption of independent errors silently, and this is the cheapest way to find out what that assumption was worth.
What to do when it fires
The check reporting dependence is the beginning of a decision rather than the end of one, and the options are the ones this field has measured.
Correct the standard error. The effective sample size, a Newey–West estimator, or a block bootstrap. The estimate stays the same and its advertised precision changes, which recovers coverage from 47.0% to 89.1% in the strong case — most of the way back, and not all of it, because the correction is estimated from the same dependent data.
Model the process. Where the dynamics are of interest, an explicit model of them answers a question the correction merely avoids getting wrong.
Difference, if the series is non-stationary. With the cost measured, and only after establishing that the series was non-stationary — differencing a series that was not doubles its variance and installs a −0.5 correlation at lag one.
Or collect differently. If the readings are taken faster than the process changes, the effective sample size can be raised more cheaply by sampling further apart for longer than by any analysis. That is the design field’s answer to this field’s problem, and it is available only before the data exists.
Testing many lags is a multiplicity problem
A correlogram with twelve lags drawn against a 95% band invites a reading that this site has a field about: with twelve independent bars, the chance that at least one crosses the band by luck alone is about 46%, not 5%.
So a single bar over the line at lag nine, with everything else inside, is close to meaningless. It is the same arithmetic as twenty analyses of nothing, performed visually and usually without anyone noticing that twelve comparisons were made.
Two responses, both standard.
Read the shape, not the crossings. A geometric decay across the first few lags is evidence; one isolated bar past the band at a lag with no interpretation is not.
Or use a portmanteau test, which combines the first several autocorrelations into a single statistic with a single p-value — the Ljung–Box test being the usual one. It answers “is there any dependence in the first k lags” once, rather than k times, which is exactly what the corrections field recommends for any family of tests.
The reason this matters more than it might is that correlograms are usually read by eye, and the eye has no way to apply a multiplicity correction to a picture.
The check the site runs on its own recommendation
This site’s gate feeds the detector both cases it has to separate. Given independent series it must leave them alone, and it does, at 4.25%. Given series with a lag-one correlation of 0.5 it must flag them, and it does, at 90.0%.
The refusal is the half that makes the rest worth anything. A detector that fires on everything is the same as no detector, and one that never fires is what this field’s recommendations would otherwise be resting on. Both failures pass every other check the site has, because both produce a well-formed number.
The measurement that follows from running both halves is the honest limitation stated above: the same check that is 90% reliable at φ = 0.5 is 20.8% reliable at φ = 0.2, and the essays are written around that rather than in spite of it.
What the check cannot see at all
Three kinds of dependence are invisible to everything in this essay, and they are worth naming so the recommendation is not read as complete.
Dependence that is not linear. The autocorrelation measures a linear association between a series and its lagged self. A process whose volatility clusters — quiet stretches and violent stretches, the usual behaviour of financial returns — has near-zero autocorrelation in the values and enormous dependence in the squares. The correlogram of the series says independent; the correlogram of the squared series says otherwise.
Dependence at a lag longer than the window. A correlogram drawn to twelve lags says nothing about lag thirty, and a process with slowly decaying, long-range dependence has correlations too small to clear the band individually while summing to a variance inflation far larger than any short-lag estimate implies.
And dependence that is not in the ordering used. A dataset ordered by time may be dependent by location, by operator, by batch. Checking the ordering that happens to be in the file finds the dependence along that axis and nothing about any other.
The first two are diagnosable — square the series, extend the lags — and the third is not, because it requires knowing what else might have been shared. That returns to the point this essay keeps arriving at: the question of what could make these observations repeat each other is answered from how they were collected, and the check measures how much of it happened to land in the ordering recorded.
What this field established
Four essays about one assumption, which every other field on this site had been making silently.
Dependence inflates the variance of a mean by a factor with a closed form — 8.20 at fifty observations and φ = 0.8 — so the fifty are worth about six, and a 95% interval covers 47.0%.
Dependence between two series manufactures relationships: two independent random walks are called significantly related 76.7% of the time with a median R² of 0.17, and the rate rises to 88.0% with four times the data, because there is no true relationship for more data to converge on.
Differencing repairs the second — 76.7% to 4.9% — and costs a real relationship most of its visible strength, from an R² of 0.91 to 0.33, while doing measurable damage to any series that did not need it.
And the diagnostic all three rely on has a power curve, like every other test: 90.0% at a correlation of 0.5, 20.8% at 0.2, and 36.2% at 0.5 in a twenty-point series.
Taken together they change what the phrase “n = 1,000” is worth. On independent data it is a statement about precision; on ordered data it is a row count, and the quantity that governs every interval and every p-value is the effective size, which can be an order of magnitude smaller and is never printed.
That is the field’s contribution to the rest of the site. Every coverage counted in the intervals field, every error rate in the testing one, and every simulation anywhere here draws independent observations, because the generator produces them independently. Those numbers are right for the world they describe, and the assumption that puts real data into that world is the one measured here.
The last measurement is the one worth carrying, because it is the field turned on itself. The recommendation is still to run the check — it costs nothing and it catches the cases that do the most damage — with the understanding that a clean correlogram in a short series is not evidence of independence, and that the question of whether these observations could repeat each other is answered by knowing how they were collected rather than by any test applied afterwards.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- Two levels at once — both name correlation, coverage, dependence, sample size
- A coverage table with its own error — both name coverage, sample size, statistical power
- Estimating how many nulls are true — both name correlation, dependence, statistical power
- False discoveries that arrive together — both name correlation, dependence, statistical power
- One number for a table of candidates — both name autocorrelation, dependence, sample size
- The repair that keeps the question — both name autocorrelation, differencing, statistical power
Named objects
A flat tag is an object no other essay names yet.
AutocorrelationCorrelationCorrelogramCoverageDependenceDifferencingDurbin–WatsonModel diagnosticsSample sizeStatistical power