A variable that moves one thing only

Weak, and back where it started

A consistent instrumental estimate at two hundred rows and a concentration parameter of 0.32 is biased by 0.3220 ± 0.0142 against a least-squares inconsistency of 0.3594 — 89.6% of the way back to the problem it was hired to solve. Just identified, it has no mean at all, and that is measured as a rate rather than assumed.

Worth reading first: The assumption nothing tests · Correcting the persistence.

Consistency is a statement about a limit. It says that as the sample grows without bound the estimate settles on the right value, and it says nothing whatever about any sample anybody has. Two-stage least squares is consistent under the exclusion restriction; at two hundred rows, four instruments and a concentration parameter of 0.32, its mean sits 0.3220 ± 0.0142 away from the truth, where the least-squares estimate it was brought in to replace sits 0.3594 away. That is 89.6% of the way back to the problem.

The direction is not an accident and it is not specific to this world. A first stage that explains nothing of the treatment leaves the fitted treatment made of noise, and a regression of the outcome on noise correlated with the outcome’s own errors returns what an unadjusted regression returns. The weak-instrument estimator degenerates towards the estimator it was designed to correct, which is the worst available failure mode: it does not blow up, it does not oscillate, it quietly reproduces the answer the analysis was run to avoid.

What this essay measures is how fast that happens, whether the standard approximation for it is any good, and what a just-identified estimate does when there is no mean for a bias to be a statement about. The third of those turns out to be the load-bearing one, and it is measured as a rate rather than taken as a fact.

A weak instrument gives back the problem it was hired for. The counted mean bias of two-stage least squares at 4 instruments and 200 rows, over 2000 draws a setting, against the standard approximation and against the least-squares inconsistency the instrument was brought in to remove. At π = 0.02 the counted bias is 0.3220 ± 0.0142 where least squares is out by 0.3594 — 89.6% of the way back. At π = 0.3 it is 0.0118 against 0.2647. The approximation, the inconsistency over the population first-stage F, tracks the count at the weak end and sits above it in the middle: 0.968, 0.971, 0.918, 0.810, 0.740, 0.722, 0.846 as the ratio of counted to approximated bias.
Fig. 1 The counted mean bias of two-stage least squares at four instruments and two hundred rows, against the standard approximation and against the least-squares inconsistency. At the weakest setting the counted bias is 0.3220 where least squares is out by 0.3594.

Why the count is at four instruments and not at one

The sweep above runs at four instruments, which looks like an odd choice for a field whose first essay is about one. The reason is that at one there is nothing to count.

Two-stage least squares has moments up to K2K - 2, where KK is the number of instruments. With one instrument the estimator is a ratio of two correlated normals and has no mean; with two it still has none; with three it has a mean but no variance, so a sample average of estimates has no standard error and no reading from it can be compared against anything. Four is the smallest count at which the counted mean is itself an estimate with a standard error, which is what makes “the approximation is wrong here” a measurement rather than an impression.

That is a constraint the honest version of this comparison cannot get around, and it is worth stating loudly because the usual presentation does not. A plot of mean bias against instrument strength, drawn at one instrument, is a plot of a quantity that does not exist — and it will look perfectly well behaved, because a sample mean of two thousand draws from a distribution with no mean is a number, and it prints.

How fast the collapse happens, in the units the F is reported in

Between the two ends of the sweep the estimator is neither one thing nor the other, and the share of the least-squares inconsistency it is carrying is the most legible way to read where it stands.

At a concentration parameter of 0.32 the counted bias is 89.6% of the least-squares inconsistency. At 2.00 it is 61.2%, at 11.52 it is 19.1%, and at 72.00 it is 4.5% — a bias of 0.0118 against an inconsistency of 0.2647. The population first-stage F at those four settings is 1.08, 1.50, 3.88 and 19.00, so the conventional threshold of ten sits between the third and the fourth, at a point where the estimator has already shed most but not all of the problem.

That places the rule of thumb rather than endorsing it. An F of ten in this world corresponds to somewhere around a tenth of the inconsistency surviving, which is a defensible thing to accept and a poor thing to describe as strong. It is also, and this is the more important half, a threshold applied to a sample statistic. The sweep reports population F values; software reports a sample F drawn from a distribution around one. Deciding whether to report an estimate on the basis of the same data’s first-stage F is a selection rule of exactly the shape that the essay measuring what a significance filter does to the effects that survive it prices, and it is not measured here.

The standard approximation errs the opposite way from the warning

The usual approximation for this bias is the least-squares inconsistency divided by the population first-stage F:

biasλγ/(π2+G)1+μ2/K.\text{bias} \approx \frac{\lambda\gamma/(\pi^2 + G)}{1 + \mu^2/K}.

Its content is exactly the claim above — that a weak instrument returns the estimator to least squares — and it is usually quoted with a warning that it is optimistic, that the real bias is worse than the formula says, and that the formula should be read as a floor.

Counted across seven settings, it is the other way round. The ratio of counted to approximated bias is 0.968 where the instrument is worthless, so the approximation is close to right exactly where it matters most. It then falls to 0.722 in the middle of the sweep, at a concentration parameter of 25.92, and the sharpest reading is at 11.52, where the counted bias sits −3.54 standard errors below what the formula predicts. Across all seven settings the number of places where the approximation understates the counted bias is 0.

The approximation is a bound in the middle, not an estimate. The counted mean bias of two-stage least squares divided by what the standard approximation predicts — the least-squares inconsistency over the population first-stage F — at 4 instruments and 200 rows, over 2000 draws each. Four instruments rather than one because the estimator's mean exists only above two and its variance only above three, so this is the smallest count at which the counted mean has a standard error of its own. The ratio is 0.968 where the instrument is worthless and falls to 0.722 in the middle of the sweep before returning towards one: the approximation is close to right at the weak end, and overstates the bias by up to 27.8% where the instrument is moderate. It understates it at 0 of the 7 settings.
Fig. 2 Counted bias divided by the standard approximation across the sweep. The ratio runs 0.968, 0.971, 0.918, 0.810, 0.740, 0.722, 0.846 — close to right at the weak end, overstating by up to 27.8% in the middle, and never optimistic.

So the formula is a bound rather than an estimate, and the direction of its error is a property of the whole sweep rather than of one cell. That is a better result than the warning it replaces, and it is more useful: a bound that is tight where the problem is severe and loose where it is not is exactly the shape a diagnostic should have. What it is not is a description. A reader using it to correct an estimate rather than to decide whether to trust one would over-correct by a quarter through most of the range where the decision is actually difficult.

The mechanism behind the shape is visible in the sweep’s own columns. At the weak end the estimator has genuinely collapsed onto least squares and the approximation is describing a real limit; in the middle the estimator is neither collapsed nor consistent, its distribution is skewed rather than centred, and a first-order expansion around either end misses the same way a first-order expansion always misses in between. At the strong end the ratio returns towards one — 0.846 at a concentration parameter of 72 — because the bias itself is down to 0.0118 and there is very little left for either route to be wrong about.

Just identified, there is no mean to be biased

The paragraph above about moments is not a technicality about which sweep to run. It is the finding, and it wants a measurement rather than a citation.

Take one instrument at a first stage of 0.05, draw estimates in blocks, and summarise each block twice — once by its mean and once by its median. If the estimator has a finite mean, both summaries are estimates of something and both settle down as the blocks grow. If it does not, the median still is and the mean is not.

Across twenty blocks of two hundred draws each, the block means scatter with a spread of 1.7328 and the block medians with a spread of 0.0766. That is a factor of 22.61 between two summaries of the same draws. At eight hundred draws a block the factor is 33.63 — it has grown, not shrunk, which is the first sign that the two summaries are not converging to the same thing at different speeds but that only one of them is converging at all.

The mean of a ratio with no mean. Twenty independent blocks of 200 draws each, at a first stage of 0.05 and 200 rows, summarised twice. A just-identified instrumental estimate is a ratio of two correlated normals and has no mean at all, so the sample mean of a block estimates nothing: the block means scatter with a spread of 1.7328 while the block medians scatter with a spread of 0.0766, a factor of 22.61. The distinction is a rate rather than a level, and it is measured as one: quadrupling the block to 800 draws divides the medians' spread by 2.089 — a statistic with a finite variance must divide it by 2.000 — and the means' by only 1.405. That is what having no mean looks like from the inside.
Fig. 3 Twenty independent blocks of two hundred draws at a first stage of 0.05, summarised by mean and by median. The block means scatter with a spread of 1.7328 and the block medians with 0.0766.

The rate is the demonstration, not the level

A factor of twenty-two is suggestive and it is not proof of anything, because two summaries of a skewed distribution are entitled to differ. What settles it is the rate, and the rate is what a block design is for.

Quadrupling the block size divides the spread of a summary with finite variance by exactly two. Measured here, the medians’ spread divides by 2.089 against the 2.000 a finite variance demands — right, to within what twenty blocks can resolve. The means’ spread divides by 1.405.

That contrast is the whole demonstration. It is not that the mean is noisier; it is that the mean is not settling at the rate a mean settles at, because it has nothing to settle on. A sample average from a distribution with no first moment wanders indefinitely: quadrupling the sample buys some of the improvement quadrupling should buy, from the bulk of the distribution, and gives it back through the larger extreme values a larger sample is more likely to include.

Reporting it this way rather than by citing the theorem is a deliberate choice. Asymptotic statements about estimators routinely fail to describe the samples people have — a forecast band derived for known parameters and then computed with estimates in them covers 87.3% rather than 95%, and a shrinkage interval that estimates a spread and proceeds as though it knew it covers 79% — and the pattern in both is that the quantity substituted in is the one the guarantee was conditional on. A theorem about moments is a statement about an idealisation, and the practical question is whether two thousand draws at a realistic first stage behave like a quantity with a mean. They do not, and the measurement says so at a stated block size with a stated number of blocks, which is the same discipline the essay on what a single simulated figure is worth argues for: one run of anything here is one draw from a distribution of runs.

The consequence for everything downstream is that every summary of this estimator in this collection is a median. That is not a stylistic preference. A mean bias at one instrument is a number with no referent, and reporting one is the kind of quiet error the essay that computes every number twice exists to catch — except that here the two routes agree, because both of them are averaging something that does not exist.

The tail that has no mean to average away

Where the mass that breaks the mean actually sits is worth seeing, because “heavy tail” is a phrase that hides how close in the damage is.

At a first stage of 0.02, 5.55% of draws land more than ten times the true effect away from it. At 0.05 it is 4.40%. At 0.30 it is 0.00% — not small, none, over two thousand draws. The central ninety per cent of estimates at the weakest setting runs from −4.436 to 6.425 for a true effect of 1.

Where the estimate lands ten times the truth away. The share of 2000 draws whose just-identified instrumental estimate falls more than 10 times the true effect away from it, at each first stage, 200 rows. At π = 0.02 it is 5.5% of draws and the central ninety per cent of estimates runs from -4.436 to 6.425 for a true effect of 1; at π = 0.6 it is 0.0% and the same band is 0.786 to 1.188. This is the part of the distribution that has no mean to average away, and it is why every summary of this estimator here is a median.
Fig. 4 The share of draws landing more than ten times the true effect away from it, at each first stage. At π = 0.02 it is 5.55% of draws and the central ninety per cent of estimates runs from −4.436 to 6.425.

A band running from about minus four to about plus six, for a quantity whose true value is one, is not a distribution anybody would summarise by its average. Slow convergence of a summary to its limiting law is a recurring shape here — the essay that measured how slowly a sample maximum reaches its own finds the same thing about extremes — but this case is worse than slow, because the limit the average is converging to does not exist. It is also not symmetric, and the asymmetry is the reason the median behaves: the ratio’s denominator is a first-stage coefficient estimated near zero, so the draws that go wrong go wrong by passing close to a division by zero, and they scatter to both sides in a way that leaves the bulk of the distribution almost untouched. The median at that setting is 0.3385 away from the truth — badly biased, but a number about which something can be said. At a first stage of 0.60 the median bias is −0.0025, which is zero to the resolution available.

Spreading the strength brings the bias back at any single instrument’s expense

There is an obvious way to escape all of this, and measuring it is what stops the essay being a complaint. More instruments give the estimator moments. So collect more of them.

Hold the total first-stage strength fixed — a concentration parameter of 8, so the instruments together explain the same amount of the treatment however many there are — and spread it over more and more of them. The estimator’s median bias runs from −0.0163 at one instrument to 0.2788 at thirty-two, and at thirty-two it has gone 80.5% of the way back to least squares.

Where the conventional interval really does fail. Coverage of a nominal 95.0% interval when the same total first-stage strength — a concentration parameter of 8 at 200 rows — is spread across more and more instruments, over 1000 draws each. At one instrument the conventional interval covers 97.2%; at 32 it covers 51.5%, and it does so while getting SHORTER, from a median 1.454 to 0.583. Nothing in the output looks wrong: the estimate has simply gone 80.5% of the way back to least squares, so it stops moving, so its residuals stop being large, so its standard error stops being generous. The exact set covers 95.1%, 94.1%, 94.8%, 95.4%, 94.2%, 95.2% across the same row.
Fig. 5 The same total first-stage strength spread over more instruments, at a fixed concentration parameter of 8. The estimator’s median bias runs to 80.5% of the least-squares inconsistency at thirty-two instruments.

So the escape does not work, and it fails in a way that is worth naming precisely. Adding instruments buys moments and spends strength. The estimator acquires a mean and a variance, both of which can be reported, and what those reportable quantities describe is an estimate that has collapsed onto the thing it was meant to replace. The estimator becomes better behaved and worse at the same time, and every diagnostic that improves is a diagnostic about behaviour rather than about the answer.

What was ruled out, and what would have produced this without the claim

Three things could have produced the block result without the estimator having no mean, and each was checked.

A short sweep. Twenty blocks is not many, and a spread computed from twenty numbers has a wide sampling distribution of its own. That is why the claim rests on the ratio between two block sizes rather than on either level: a spread that is merely noisy is noisy at both block sizes, and the medians’ shrink of 2.089 against a demanded 2.000 shows the design has the resolution to detect the rate when the rate is there.

A coincidence of the seeds. Every block is drawn from an independent stream, and the same streams produce the medians as produce the means — so a seed that happened to be extreme is extreme in both summaries. The 22.61 factor is a comparison within blocks, not between runs.

The choice of first stage. The blocks are read at 0.05, which is weak by design. Comparing two estimators at one setting is the mistake the essay that retuned two windows before comparing them is about, and the defence here is the same: the claim is made across a sweep, and the setting a figure is drawn at is stated rather than chosen for the picture. At a strong first stage the estimator’s tail is thin enough that a sample mean over two hundred draws behaves perfectly well over any run anybody will do, and it still has no mean — which is exactly why the demonstration is at 0.05 rather than at 0.30. The claim being made is about the estimator, and the setting is chosen to make a statement about the estimator visible at a sample size somebody might actually use. A reader who wants the claim at 0.30 will not get it from two thousand draws, and saying so is more useful than a picture that seems to.

What the bias does not settle

Two things this measurement deliberately leaves open.

It says nothing about the exclusion restriction. Every world in every sweep here has the restriction holding exactly, so the estimator is consistent throughout and the entire quantity being measured is a finite-sample effect — a thing that goes away with more data, unlike the violation priced against the first stage, which does not. Reading a weak-instrument bias as evidence about validity confuses two failures that happen to share a symptom.

And it says nothing about what to do. The standard repair for many weak instruments removes each observation’s own contribution to its fitted value, and it is the reason this failure is usually described as fixable rather than fatal. That repair is not measured here — the sweep that would carry it is the one drawn above, with one more series on it — so an essay that showed the failure and not the fix is an essay that has overstated its case, and this paragraph is where that is admitted rather than hidden. What the measurement does support is narrower and firmer: at the sample sizes and first stages a reader will actually meet, a consistent estimator is returning most of the confounding it was brought in to remove, its standard approximation is a bound rather than a description, and just identified it has no mean for either of those statements to be about.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Approximation errorAsymptotic biasConcentration parameterConsistencyEndogeneityF statisticFirst stageHeavy tailInstrumental-variableJust-identifiedLeast squaresMonte CarloTwo-stage least squaresWald ratioWeak instrument