Every essay — page 3
Regression, and what the summary hides
A slope, a standard error and an R² can all be computed from data the model is grossly wrong about, and none of them says so. Leverage is a number known before the outcome is looked at, influence is a number, and whether a residual plot looks bad is a question with a calibrated answer. Two wrong rows placed together hide from every single-row diagnostic, and a loss that bounds a large residual does nothing about a far one.
A standard error for a model that is wrong
Every standard error a regression prints is a statement about a model that is true, and the repair for a model that is not is one of the most-quoted lines in applied work. It is priced here rather than recommended: the robust error reads a spread the model-based one does not, and its 95% interval covers 88.73% at twenty rows, is the worse of the two under mild heteroskedasticity until fifty, and is the smaller of the two whenever the error variance sits in the middle of the design rather than at its edges. The four corrections differ in one number — what each does with a point's leverage — and on a design with a single far-out point that number takes one of them to 0.3191 of the truth and another to 5.1127 times it.
What a wrong model estimates
A straight line fitted to a curved truth converges on the tangent at its own design's mean. Two honest studies of one world, fitting the same wrong model, report 2.600000 and 1.600000, and neither is in error.
The bread and the filling
The robust standard error is not a safety margin. At one setting of the error variance it is 1.2806 times the model-based one and at another it is 0.8246 times it, and the sign of a single dial decides which.
Robust is not free
A robust standard error's promise is asymptotic and its use is not. Its 95% interval covers 88.73% at twenty rows, and under mild heteroskedasticity it is the worse of the two intervals until a hundred.
Three corrections and a leverage
On an even design of twenty rows the four robust corrections read 0.8603, 0.9559, 1.0000 and 1.1647 of the truth and the choice barely matters. Add one point at x = 8 and they read 0.3191, 0.3419, 1.0000 and 5.1127.
Right for the wrong reason
A robust standard error costs no coverage where the risk is absent — 95.52% against 95.06% at twenty rows. It costs a 6.89% wider interval and a variance estimate 2.572 times as variable, and the pre-test that would avoid paying recovers 15.9% of what the insurance is worth.
The count that is not the rows
Three hundred rows in five clusters of sixty carry 6.9000 times the variance an independent-rows calculation reports, and the interval that counts rows covers 53.42%. The same five unequal sizes laid out two ways give design effects of 9.3158 and 5.4652.
The reference the sandwich is read against
The cluster-robust interval covers 75.05% at five clusters and 93.58% at eighty. The same estimate read against a t on G − 2 covers 87.95% at five, and the estimator is unchanged — three hundred rows grouped into five clusters cover 74.28% where the same three hundred grouped into seventy-five cover 94.63%.
A variable that moves one thing only
An instrument identifies a causal effect by assuming that a path no data can check is exactly zero, and it charges for that assumption in variance. Both prices are computed rather than described: a violation nobody can see is divided by the first stage, so the strength an instrument needs is 2.78 times the violation it is assumed not to have, and the conventional interval turns out to fail not when one instrument is weak but when many are — a failure each row's own first-stage leverage explains, and leaving that row out repairs at a price in width that depends on how strong the instruments are. What is left identified is the effect among the units the instrument actually moved, which at the standing compliance profile is 1.1000 against a population average of 0.5000.
The assumption nothing tests
An instrument buys a causal effect with an assumption no sample can check, and the price is set by the same quantity that made the method work. The first stage it needs is 2.7778 times the violation it is assumed not to have, so a direct effect of 0.05 demands a first stage of 0.1389 and least squares wins below it.
Weak, and back where it started
A consistent instrumental estimate at two hundred rows and a concentration parameter of 0.32 is biased by 0.3220 ± 0.0142 against a least-squares inconsistency of 0.3594 — 89.6% of the way back to the problem it was hired to solve. Just identified, it has no mean at all, and that is measured as a rate rather than assumed.
What the first stage does not know
A single weak instrument does not make the conventional interval undercover — it makes it cover 99.1% at a width of 7.320. Where the promise actually breaks is many instruments — coverage falls from 97.2% to 51.5% while the median width falls from 1.454 to 0.583.
Whose effect it is
With a perfectly valid instrument and no violation of anything, the estimate converges on 1.1000 where the population average effect is 0.5000. The gap is exactly θ(1 − p_c), the always-takers and never-takers cancel out of both halves of the ratio, and five per cent defiers move the answer to 1.2667.
Two instruments that disagree
The overidentification test keeps its size at 5.0% and reaches 86.4% power against a violation carried by one instrument. Against the same error carried by both in proportion to their first stages it rejects on 4.6% of draws — its own size — while the estimate is wrong by 0.3000, which is 94.2% of the confounding the instruments were brought in to remove.
Leaving each row out of its own first stage
Spread a fixed first-stage strength over thirty-two instruments and two-stage least squares covers 51.5%. Build each row's fitted treatment from a first stage that never saw that row and the same draws cover 98.7% — through an interval 5.99 times as wide, around an estimate that misses by more than the whole effect on 34.7% of draws. At eight times the strength the same repair covers 95.3% and costs a width factor of 1.66.
A standard error that knows about the instruments
Limited-information maximum likelihood came out least biased when a concentration parameter of 8 was spread over thirty-two instruments, and its conventional interval covered 79.0%. Bekker's many-instrument standard error covers 93.8% on the same draws, at 63% of the jackknife's width — and it gets there with a median standard error of 0.561 against a true spread of 0.797, because it is large on the draws that need it. At eight times the strength it covers 94.9% at 91% of the jackknife's width, and nothing measured here beats it.
What conditioning on a variable does
Putting a covariate into the regression is one arithmetic operation, and it is the right thing to do in one of the three worlds it could have come from. Here the three worlds are built to share a covariance matrix entry for entry, so the estimate they disagree about by 0.348 is a number no sample of any size can settle, and the rule that controls for everything measured is measured: it leaves a larger bias than controlling for nothing on 65.5% of four thousand structures.
One arithmetic, three decisions
A covariate beside a treatment and an outcome can be a common cause of both, a step on the path between them, or an effect of both. The regression that includes it is the same arithmetic in all three, and it is right in one — returning 0.5000, deleting 0.6300 of the effect, and turning 0.5000 into −0.0872.
The two worlds that look the same
Three causal structures were fitted to one covariance matrix and agree with it to 4.4·10⁻¹⁶. The regression returns 0.5000 under all three; the effect they hold is 0.5000, 0.8481 and 0.8481. What separates structures is a missing edge, and the signature of one is a correlation of exactly zero.
Adjusting for everything
"Control for every covariate that was measured" leaves a larger bias than controlling for nothing on 65.5% of four thousand randomly drawn structures and a smaller one on 33.8%. Its squared error is 4.110 times that of using no covariate at all, and half of it sits in its worst tenth of structures.
A collider before the treatment
A covariate measured before the treatment, on no causal path, and not a common cause of anything, still biases the estimate by exactly −0.2000 against an effect of 0.5 — while the regression that leaves it out is exact. The bias saturates at 0.3536, and the two paths that make it a collider do not appear in that bound.
The sample is a condition
Two independent standard normals, selected on their sum exceeding its median, read a correlation of exactly −1/(π − 1) = −0.4669 inside the sample. Nothing is measured badly and nothing is missing — and both halves of that split read it, in the same direction, while the population containing both reads zero.
The variable the treatment caused
Adjusting for a covariate the treatment caused stops estimating the total effect and starts estimating the direct one. When that covariate shares an unmeasured cause with the outcome it estimates neither: the total effect is 1.1300, the direct effect is 0.5000, and the regression returns 0.0500.
Adjusting for a shadow
A covariate that is 80% signal removes 68.85% of the confounding, not 80% — the share is λ(1 − ρ²)/(1 − λρ²) and it is below the reliability everywhere. The residual bias is 0.1084 against an effect of 0.5, and at 25,600 rows it is 17.6 standard errors wide.
Corrections, and what each controls
Bonferroni bounds the chance of any false positive; Benjamini–Hochberg bounds the share of the findings that are false. Both get called correcting for multiple comparisons and they are different promises, so every procedure here is made to report both rates and the power each one costs.
What the correction corrects
Twenty tests of true nulls produce at least one false positive 64% of the time, and the closed form and the count agree. Bonferroni holds it at 5% and Holm holds it at 5% while finding more. Nobody should still be using Bonferroni.
Two different promises
Bonferroni bounds the chance of any false positive. Benjamini–Hochberg bounds the share of the findings that are false. Both are called correcting for multiple comparisons, and one of them lets the familywise rate reach 20%.