Dropping the incomplete rows
Worth reading first: Three mechanisms and one dataset.
Take the missingness rule that leaves the fitted slope alone and lean on it as hard as it will go. At an index coefficient of 3.2 the rows that survive have a covariate mean of 0.543905 against a population mean of zero, a covariate variance of 0.5041 against one, and a mean outcome wrong by 0.391612. Nearly half the spread of the regressor has been deleted and the sample sits half a standard deviation to one side of the population it came from.
The fitted slope is exactly right. Not approximately: the closed bias at that setting is −1.11 × 10⁻¹⁶, which is one unit in the last place of a double, and the count over a thousand studies is 0.00086 against a standard error of 0.00389.
Most of what makes this hard to believe is a habit of reading “the sample is unrepresentative” as a single verdict. It is not one verdict. A sample can be unrepresentative of the covariate’s distribution, of the outcome’s distribution, of the joint distribution, and of the conditional law of one given the other, and those are four separate statements — three of which are true here and one of which is not.
How strong the strong setting is
“Leaning as hard as it will go” needs a number attached to it, because a probit coefficient is not a quantity anybody has an intuition for. At the field’s own setting of 1.2, the chance of a row being recorded runs 0.036080, 0.274883, 0.726376, 0.964219 and 0.998658 at covariate values of −2, −1, 0, 1 and 2. A row two standard deviations below the mean survives about one time in twenty-eight; a row two above survives essentially always.
That is not a mechanism anybody would describe as mild, and at 3.2 it is not a mechanism anybody would describe as a mechanism — it is closer to a hard cut with a soft edge. Roughly a sixth of the covariate’s range contributes almost every recorded row, and the rest of the population appears in the table as an absence. The point of pushing it there is that the claim below is an identity, and an identity that is only ever exhibited at gentle settings is indistinguishable from an approximation working well. The four rules separated by which column they read are compared at 1.2 for the same reason.
The one thing least squares conditions on
Write the outcome equation and the selection rule side by side:
with and independent standard normals. The rule reads and , and it does not read .
Least squares of on is an estimator of the conditional mean . Condition on a pair of covariate values, and among the rows carrying that pair the chance of being recorded is one fixed number — high or low, but the same for every such row, whatever its residual. So the recorded rows at that pair are a plain random sample of all rows at that pair, and the conditional mean among them is the conditional mean. The selection changes which pairs appear, and in what proportion; it does not change what the outcome does at a pair.
That is the whole mechanism, and it explains both halves at once. The marginal law of moves, because whole regions of it are deleted at different rates. The conditional law of given and does not move, because the deleting never reads the part of that is not already a function of and . A regression estimates the second and is silent about the first, so it is right about the only thing it is claiming and wrong about nothing it claimed.
The closed forms say the same in one line. The bias of the complete-case coefficient vector is , where is the covariance of the residual with the standardised selection index — and with that covariance is zero for any , any and any strength whatever. It is an identity rather than a small number, and that distinction is what the sweep below exists to make visible.
Pushed until it stops being plausible
A claim of exactness has one obvious failure mode: it holds at the setting somebody chose and decays away from it. So the dependence is swept, at a fixed 35% missing throughout, and every reading is taken twice.
| index coefficient | closed bias | counted bias | mean of | variance of | outcome’s mean wrong by |
|---|---|---|---|---|---|
| 0 | 1.11 × 10⁻¹⁶ | 0.00083 ± 0.00284 | 0 | 1.0000 | 0 |
| 0.4 | 1.11 × 10⁻¹⁶ | 0.00118 | 0.211635 | 0.9249 | 0.152377 |
| 0.8 | 0 | 0.00329 | 0.355979 | 0.7876 | 0.256305 |
| 1.2 | −1.11 × 10⁻¹⁶ | 0.00229 | 0.437767 | 0.6788 | 0.315192 |
| 1.6 | 0 | 0.00226 | 0.483227 | 0.6086 | 0.347924 |
| 2.4 | 1.11 × 10⁻¹⁶ | 0.00222 | 0.526010 | 0.5362 | 0.378728 |
| 3.2 | −1.11 × 10⁻¹⁶ | 0.00086 | 0.543905 | 0.5041 | 0.391612 |
The middle column never leaves the last bit of a double while the three to its right move by half a unit, a factor of two and four tenths respectively. The counted column drifts upward by about two thousandths in the middle of the sweep, which is a third of one standard error and is what a thousand studies can resolve; nothing there survives its own error bar.
Pushed further still, to an index coefficient of six — where a row one standard deviation below the mean is recorded about one time in a thousand — the survivors’ covariate mean reaches 0.562091, their variance falls to 0.470415, and their mean outcome is wrong by 0.404706. The closed bias at that setting is 0, with no exponent after it.
It is not about the regressor in particular
A reader who has followed the mechanism will want to check that the protection is not an artefact of the mechanism reading the very variable whose coefficient is being estimated. It is not.
Drive the missingness from instead — the second covariate, the one whose coefficient is 0.4 — and the closed bias on the slope on is again exactly zero at every strength. The counted biases across the same seven settings never exceed −0.00123 ± 0.00295, and coverage runs between 95.00% and 96.20%. The surviving rows are still not the population: the mean of among them climbs from 0 to 0.163172 as the dependence strengthens, dragged along by the correlation of 0.3 between the two covariates.
What matters is not which variable the rule reads but whether the analysis conditions on it. That qualification is sharper than it looks, and it is the one place where the protection is easy to lose by accident. Under the rule driven by , the partial slope on is exact and the marginal slope of on alone converges to 0.683878 against a true 0.72. Same rows, same mechanism, two analyses, one of them right — which is the same structural point three causal readings of one regression make about whether a column belongs in a fit. Dropping a covariate that the missingness reads converts a harmless mechanism into a biased estimate, and nothing in the output signals it.
The mirror, and it is the worse half
Change one thing — let the rule read the outcome rather than a covariate — and the same sweep produces a bias that grows monotonically: −0.04324, −0.11399, −0.16353, −0.19287, −0.22123 and −0.23323 in closed form across the same six strengths, counted at −0.04434, −0.11305, −0.16193, −0.19219, −0.22122 and −0.23276. Coverage falls with it: 93.40%, 74.10%, 52.10%, 34.80%, 21.10% and 16.50%.
Two things about that column are worth reading rather than skipping. The bias saturates: the step from 0.4 to 0.8 is worth 0.07 and the step from 2.4 to 3.2 is worth 0.012, because the whole fitted relation is multiplied by a factor that falls from 0.927933 to 0.611285 and flattens out. A mechanism that reads the outcome twice as hard does not do twice the damage, and a reader who calibrated their worry to the strength of the dependence would get the direction right and the size badly wrong.
And the damage arrives early. At an index coefficient of 0.4 — a dependence so weak that the chance of being recorded runs from about a half to about three quarters across the bulk of the outcome’s range — the bias is already −0.04324 and coverage is already down to 93.40%. There is no setting at which reading the outcome is harmless and no threshold below which it can be ignored; there is only a size, and the size is a function of a quantity nobody can measure. An effect inflated among the studies that reached significance is the same arithmetic in a setting where the filter is at least written down.
The two curves in the strength figure are the field in one picture. They are drawn at the same missing fraction, at the same strengths, from the same generator. One is a flat line on zero. The distinction between them is not visible in any dataset either could produce.
And the second reading is worse than a bias, because of what it does with more data.
At 100, 200, 400 and 800 rows the outcome-driven interval covers 70.92%, 50.42%, 21.17% and 2.42%. The bias barely moves — −0.16522, −0.16427, −0.16136, −0.16360, all within noise of the closed −0.163531 — while the standard error falls from 0.11879 to 0.04108. The ratio of the two runs −1.377, −1.973, −2.810 and −3.981, and coverage is essentially a function of that ratio alone.
So the interval is not merely optimistic. It converges, at the usual rate, onto a wrong answer, and it advertises increasing confidence in it. Under the regressor-driven rule the same four sizes give 95.33%, 96.17%, 95.25% and 93.92%: no trend, the ordinary non-monotone wobble a counted coverage has, and nothing that a larger study would fix or break.
There is no sample size at which a reader would notice. The estimate settles down, the interval narrows, every diagnostic a complete-data analysis offers is clean — because the fit really is a correct fit to the recorded rows, and the recorded rows really do follow a linear model with normal errors. The defect has no symptom. It is the shape of failure that survives longest for exactly that reason: nothing is wrong with what is there.
Coverage that sits a little above its promise
One column of the sweep is worth stopping on, because it runs the other way and a reader is entitled to be suspicious of it. Under the regressor-driven rule the counted coverage across the seven strengths reads 96.30%, 96.90%, 95.90%, 95.70%, 96.50%, 96.20% and 96.10% — consistently a little above ninety-five rather than at it, on a thousand studies whose binomial error is about 0.6 points.
That is not the mechanism. It is the interval. The number of recorded rows is itself random, the degrees of freedom of the t interval are read off it, and an interval whose width is set by a quantity that varies from study to study is slightly wider on average than the one a fixed sample size would give. The effect is small, it appears at every strength including zero, and it is present in the completely-random column too. A reading that treated the excess as evidence for the mechanism would be reading the shape of a t interval and calling it a finding.
The four thousand studies behind that figure give the same three rules 95.15%, 95.40% and 94.80% and put the fourth at 49.70%, which is the reading in its blunt form: the 95% is a claim about a procedure, and here two procedures that differ in nothing a reader can see keep and break it.
What the exactness does not buy
The slope is one quantity out of a table full of them, and the protection covers only what it covers.
On the same rows and the same draws that give an exact slope, the mean of the recorded outcomes is wrong by +0.31427 ± 0.00273, against a closed prediction of 0.315192. That is a third of a unit on a quantity whose true value is 1, and it comes from the identical arithmetic — the mean of the recorded outcomes is , which moves for precisely the reason the slope does not.
The same split runs through the whole table. The partial slope on is exact; the marginal slope of on alone is exact too under this rule, since the mechanism reads and both regressions condition on it; the mean of , the variance of , the mean of and the variance of are every one of them wrong. There is no general principle available beyond the one already stated — a quantity is protected when the analysis conditions on everything the mechanism reads, and a marginal summary conditions on nothing.
So an analysis of these rows can report a coefficient that is right to fourteen decimals and, three lines above it, a descriptive mean that is out by a third of a unit, with no warning that the two were computed under different protection. The estimand decides, and the four datasets that share a summary are the standing reminder that a table of numbers does not announce which of them the data support.
What could have made the exactness wrong
A rounding artefact. The closed biases are and , alternating in sign across the sweep, which is what a quantity that is algebraically zero looks like after two matrix solves. If it were a genuinely small nonzero number it would have a consistent sign and would scale with the strength; it does neither.
A count too coarse to see it. A thousand studies per strength resolve the bias to about 0.003. That is not enough to rule out a bias of 0.0005 and is easily enough for the comparison actually being made, since the outcome-driven curve reaches −0.23323 at the same setting. The claim of exactness rests on the identity and the count is a check on the identity, not the other way round.
Normality doing the work. It is doing some of it. The truncated moments are closed because the joint law is normal, and the sweep is therefore a check of the algebra rather than of the generality. What is not normality-specific is the mechanism: the argument that selection on the conditioning variables leaves the conditional law alone needs only that the rule not read the residual, and it is the closed forms rather than the result that would go if the law changed.
Where it stops
Two limits, both real and neither hedging.
The analysis must be a conditional-mean analysis of the variables the rule reads. The whole result is a property of least squares conditioning on the regressor. A logistic outcome model, a quantile regression, or any summary of the marginal distribution of has no such protection, and the descriptive mean above is the smallest possible example of that failing.
The gaps must be in the outcome. Everything measured here deletes . When the missingness falls on instead, the selection acts on exactly the variable least squares conditions on, and the protection does not transfer — the recorded rows then carry a truncated regressor and the conditional law being estimated is a conditional law over a restricted range. The closed machinery covers that case, since the truncation identity does not care which variable is deleted, and the analysis of it is not attempted here. That is the largest single gap this argument leaves, and it is named rather than papered over, in the same spirit as the third state a censored observation occupies being given a slot of its own rather than forced into one of the two that already existed.
What is settled is narrower than “dropping incomplete rows is safe” and sharper. Dropping them costs nothing at all for a coefficient whose regressors the mechanism reads, however unrepresentative the survivors are, and costs everything for a coefficient whose mechanism reads the outcome — with the cost growing as the study grows. Which of those two a dataset is in is not a question the dataset answers.
The reading a practitioner can take from this is narrower still, and it is about what to write down rather than what to compute. Whether the complete cases are enough is decided by two things that are both stated before the data arrive: which variables the chance of being recorded plausibly depends on, and which of those the analysis conditions on. Neither is a diagnostic and neither is checkable after the fact — counting an advertised rate rather than asserting it is available here only because the truth is known by construction. In a real study the honest artefact is the pair of lists, and an analysis that names neither has made the assumption anyway.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- What a wrong model estimates — both name closed form, estimand, least squares, sampling variation
- A penalty is a trace — both name closed form, effective sample size, least squares
- A simulation that stops when it looks settled — both name closed form, confidence interval, selection bias
- A weight fitted to balance — both name closed form, effective sample size, estimand
- One number for a table of candidates — both name closed form, effective sample size, least squares
- The count that is not the rows — both name confidence interval, effective sample size, sampling variation
Named objects
A flat tag is an object no other essay names yet.
Closed formComplete-caseConditional distributionConfidence intervalEffective sample sizeEstimandThe inverse Mills ratioLeast squaresMissing at randomMissing not at randomMissingness mechanismObservation propensitySampling variationSelection bias