observations-that-are-not-independent

Twenty series with a lag-one correlation of 0.8

Every series has a true mean of zero and 60 observations. The marks on the right are the twenty sample means. The variance of that mean is 8.3 times what 60 independent observations would give, so the series is worth about 7 of them.

When the observations repeat each otherslider: lag-one correlation, 6 positionswide7 views

What else it draws

The same object, drawn to answer the other questions the essays put to it.

The bars are the measured correlations; the curve is φᵏ, which is what an AR(1) must have. The band is ±1.96/√n, where an independent series would stay. The first bar is 0.81 against a band of ±0.14.

Each point is 3,000 series of 50 observations. The ordinary interval covers 94.6% at φ = 0 and 28.1% at φ = 0.92. Dividing by the effective sample size n(1 − φ)/(1 + φ) instead of by n brings it back to 82.1%.

Nothing connects these two series: each is generated from its own independent draws. Regressing one on the other gives a slope with t = -10.9, R² = 0.55 and p = 0.0e+0 — a result that would be reported as a finding by any standard output.

The first pair is the false-positive rate for two independent random walks: 77% on the levels, 4.9% on the differences. The second pair is how much of a real relationship survives: R² falls from 0.91 to 0.33. The same operation does both.

Two readings at each persistence. In the darker colour, how often a regression between two independent series of 200 steps is called significant at 5%: 4.9% at φ = 0, 34.2% at 0.8, 52.4% at 0.9, 83.4% at a unit root. In the lighter, how often the standard unit-root test refuses a unit root on one of those series — the chance the analyst is told the series is stationary and may be regressed: 87.2% at φ = 0.9 and 31.9% at 0.95. At φ = 0.9 both are high at once, which is a correct diagnostic licensing a regression that is wrong half the time.

How often a regression between two independently generated series is called significant at the 5% level, for two worlds and three treatments, at 200 observations. Every pair is independent by construction, so 5% is correct everywhere and every other reading is a failure. Untreated: 82.9% and 100.0%. With a fitted line removed: 74.2% and 33.5%. Differenced: 5.0% and 5.2%. The treatment that controls the rate in both worlds is the one that discards the level and the trend, which is the quantity a study of trending series was about.

Where it is used

11 essays draw this figure, each at the numbers its own argument is about, so the same picture answers 11 different questions.

All 80 figures

FieldsThreadsSeriesConceptsAll essaysSearch