The value that is not there

Dropping the incomplete rows

Push the missingness until the rows that survive have a covariate mean of 0.543905 against a population zero and a variance of 0.5041 against one, and the fitted slope is still exactly right. Where the rule reads the outcome instead, the same sweep takes coverage to 2.42% at eight hundred rows.

Worth reading first: Three mechanisms and one dataset.

Take the missingness rule that leaves the fitted slope alone and lean on it as hard as it will go. At an index coefficient of 3.2 the rows that survive have a covariate mean of 0.543905 against a population mean of zero, a covariate variance of 0.5041 against one, and a mean outcome wrong by 0.391612. Nearly half the spread of the regressor has been deleted and the sample sits half a standard deviation to one side of the population it came from.

The fitted slope is exactly right. Not approximately: the closed bias at that setting is −1.11 × 10⁻¹⁶, which is one unit in the last place of a double, and the count over a thousand studies is 0.00086 against a standard error of 0.00389.

Unrepresentative in every respect but the one that mattersThree properties of the complete cases as the chance of being observed leans harder on the regressor, in closed form, at 35.0% of outcomes missing throughout. The mean of the regressor among the rows kept climbs from 0.0000 to 0.5528 against a population mean of zero, and the mean of the outcome from 0.0000 to 0.3980 above its own. The bias in the fitted slope is exactly zero at every one of the ten settings, because selection acting on the regressor alone leaves the conditional law of the outcome given the regressor untouched and least squares conditions on exactly that. The sample is wrong about almost everything and right about the one quantity being estimated.00.2000.4000.60001234how hard the chance of being observed leans on the regressorhow far the complete cases are from the populationthe regressor's mean, among the rows keptthe outcome's mean, among the rows keptthe bias in the fitted slopeclosed form, 35.0% of outcomes missingzero means no bias
Fig. 1 Three properties of the surviving rows as the chance of being recorded leans harder on the regressor, in closed form. Two of the curves climb steadily away from the population and the third, the bias in the fitted slope, is flat on zero. The slider changes the share of outcomes that go missing.

Most of what makes this hard to believe is a habit of reading “the sample is unrepresentative” as a single verdict. It is not one verdict. A sample can be unrepresentative of the covariate’s distribution, of the outcome’s distribution, of the joint distribution, and of the conditional law of one given the other, and those are four separate statements — three of which are true here and one of which is not.

How strong the strong setting is

“Leaning as hard as it will go” needs a number attached to it, because a probit coefficient is not a quantity anybody has an intuition for. At the field’s own setting of 1.2, the chance of a row being recorded runs 0.036080, 0.274883, 0.726376, 0.964219 and 0.998658 at covariate values of −2, −1, 0, 1 and 2. A row two standard deviations below the mean survives about one time in twenty-eight; a row two above survives essentially always.

That is not a mechanism anybody would describe as mild, and at 3.2 it is not a mechanism anybody would describe as a mechanism — it is closer to a hard cut with a soft edge. Roughly a sixth of the covariate’s range contributes almost every recorded row, and the rest of the population appears in the table as an absence. The point of pushing it there is that the claim below is an identity, and an identity that is only ever exhibited at gentle settings is indistinguishable from an approximation working well. The four rules separated by which column they read are compared at 1.2 for the same reason.

The one thing least squares conditions on

Write the outcome equation and the selection rule side by side:

y=1+0.6x+0.4z+ε,recorded when c0+cxx+czz+u>0,y = 1 + 0.6\,x + 0.4\,z + \varepsilon, \qquad \text{recorded when } c_0 + c_x x + c_z z + u > 0,

with ε\varepsilon and uu independent standard normals. The rule reads xx and zz, and it does not read ε\varepsilon.

Least squares of yy on (x,z)(x, z) is an estimator of the conditional mean E[yx,z]\mathbb{E}[y \mid x, z]. Condition on a pair of covariate values, and among the rows carrying that pair the chance of being recorded is one fixed number — high or low, but the same for every such row, whatever its residual. So the recorded rows at that pair are a plain random sample of all rows at that pair, and the conditional mean among them is the conditional mean. The selection changes which pairs appear, and in what proportion; it does not change what the outcome does at a pair.

That is the whole mechanism, and it explains both halves at once. The marginal law of xx moves, because whole regions of it are deleted at different rates. The conditional law of yy given xx and zz does not move, because the deleting never reads the part of yy that is not already a function of xx and zz. A regression estimates the second and is silent about the first, so it is right about the only thing it is claiming and wrong about nothing it claimed.

The closed forms say the same in one line. The bias of the complete-case coefficient vector is δpεVsel1(px,pz)-\delta\, p_\varepsilon V_{\text{sel}}^{-1}(p_x, p_z)', where pεp_\varepsilon is the covariance of the residual with the standardised selection index — and with cy=0c_y = 0 that covariance is zero for any cxc_x, any czc_z and any strength whatever. It is an identity rather than a small number, and that distinction is what the sweep below exists to make visible.

Pushed until it stops being plausible

A claim of exactness has one obvious failure mode: it holds at the setting somebody chose and decays away from it. So the dependence is swept, at a fixed 35% missing throughout, and every reading is taken twice.

index coefficient closed bias counted bias mean of xx variance of xx outcome’s mean wrong by
0 1.11 × 10⁻¹⁶ 0.00083 ± 0.00284 0 1.0000 0
0.4 1.11 × 10⁻¹⁶ 0.00118 0.211635 0.9249 0.152377
0.8 0 0.00329 0.355979 0.7876 0.256305
1.2 −1.11 × 10⁻¹⁶ 0.00229 0.437767 0.6788 0.315192
1.6 0 0.00226 0.483227 0.6086 0.347924
2.4 1.11 × 10⁻¹⁶ 0.00222 0.526010 0.5362 0.378728
3.2 −1.11 × 10⁻¹⁶ 0.00086 0.543905 0.5041 0.391612

The middle column never leaves the last bit of a double while the three to its right move by half a unit, a factor of two and four tenths respectively. The counted column drifts upward by about two thousandths in the middle of the sweep, which is a third of one standard error and is what a thousand studies can resolve; nothing there survives its own error bar.

Pushed further still, to an index coefficient of six — where a row one standard deviation below the mean is recorded about one time in a thousand — the survivors’ covariate mean reaches 0.562091, their variance falls to 0.470415, and their mean outcome is wrong by 0.404706. The closed bias at that setting is 0, with no exponent after it.

One curve is identically zero at every strength. The bias of the complete-case slope as the missingness leans harder on what it reads, at 35.0% missing throughout, over 1000 studies at each of 7 strengths. When the rule reads the regressor the closed form is exactly zero at every strength and the counts sit within their own standard errors of it — the largest is 0.0033 against errors of about 0.0035. When the rule reads the outcome the bias grows from -0.0432 at the gentlest setting to -0.2332 at the harshest. Neither curve is about how much is missing, which is held fixed; both are about what the rule looked at.
Fig. 2 The bias in the complete-case slope against the strength of the dependence, for a rule that reads the regressor and one that reads the outcome. Lines are closed, dots are counted, and the same 35% is missing at every point on both.

It is not about the regressor in particular

A reader who has followed the mechanism will want to check that the protection is not an artefact of the mechanism reading the very variable whose coefficient is being estimated. It is not.

Drive the missingness from zz instead — the second covariate, the one whose coefficient is 0.4 — and the closed bias on the slope on xx is again exactly zero at every strength. The counted biases across the same seven settings never exceed −0.00123 ± 0.00295, and coverage runs between 95.00% and 96.20%. The surviving rows are still not the population: the mean of xx among them climbs from 0 to 0.163172 as the dependence strengthens, dragged along by the correlation of 0.3 between the two covariates.

The rows that survive are not the population. How far the complete cases depart from the population they were drawn from, in closed form, at 35.0% missing and an index coefficient of 1.2. The covariate has mean zero and standard deviation one in the population; among the rows kept when missingness reads the regressor its mean is 0.4378 and its standard deviation 0.8239, and when missingness reads the outcome the mean is 0.2672. Completely random missingness moves neither. Read beside the bias figure this is the field's whole point: the sample on the second row is thoroughly unrepresentative and the slope computed from it is exactly right.
Fig. 3 How far the surviving rows depart from the population under each rule, in closed form. Every rule but the completely random one moves both the mean and the spread of the regressor, and three of the four leave the slope untouched.

What matters is not which variable the rule reads but whether the analysis conditions on it. That qualification is sharper than it looks, and it is the one place where the protection is easy to lose by accident. Under the rule driven by zz, the partial slope on xx is exact and the marginal slope of yy on xx alone converges to 0.683878 against a true 0.72. Same rows, same mechanism, two analyses, one of them right — which is the same structural point three causal readings of one regression make about whether a column belongs in a fit. Dropping a covariate that the missingness reads converts a harmless mechanism into a biased estimate, and nothing in the output signals it.

The mirror, and it is the worse half

Change one thing — let the rule read the outcome rather than a covariate — and the same sweep produces a bias that grows monotonically: −0.04324, −0.11399, −0.16353, −0.19287, −0.22123 and −0.23323 in closed form across the same six strengths, counted at −0.04434, −0.11305, −0.16193, −0.19219, −0.22122 and −0.23276. Coverage falls with it: 93.40%, 74.10%, 52.10%, 34.80%, 21.10% and 16.50%.

Two things about that column are worth reading rather than skipping. The bias saturates: the step from 0.4 to 0.8 is worth 0.07 and the step from 2.4 to 3.2 is worth 0.012, because the whole fitted relation is multiplied by a factor that falls from 0.927933 to 0.611285 and flattens out. A mechanism that reads the outcome twice as hard does not do twice the damage, and a reader who calibrated their worry to the strength of the dependence would get the direction right and the size badly wrong.

And the damage arrives early. At an index coefficient of 0.4 — a dependence so weak that the chance of being recorded runs from about a half to about three quarters across the bulk of the outcome’s range — the bias is already −0.04324 and coverage is already down to 93.40%. There is no setting at which reading the outcome is harmless and no threshold below which it can be ignored; there is only a size, and the size is a function of a quantity nobody can measure. An effect inflated among the studies that reached significance is the same arithmetic in a setting where the filter is at least written down.

The two curves in the strength figure are the field in one picture. They are drawn at the same missing fraction, at the same strengths, from the same generator. One is a flat line on zero. The distinction between them is not visible in any dataset either could produce.

And the second reading is worse than a bias, because of what it does with more data.

More data does not repair a mechanism. What the complete-case 95% interval for the slope covers as the study grows, counted over 1200 studies at each size, at 35.0% missing. When the missingness reads the regressor the interval holds its promise at every size — 95.33%, 96.17%, 95.25%, 93.92% at 100, 200, 400, 800 rows. When it reads the outcome the coverage falls from 70.92% to 2.42% across the same sizes, because the bias is fixed at -0.1635 while the interval narrows around it: the ratio of the two runs 1.38, 1.97, 2.81, 3.98. A wrongly centred interval is the one thing more observations make worse.
Fig. 4 What the complete-case interval covers as the study grows. Under a rule that reads the regressor it holds its promise; under one that reads the outcome it falls from 70.92% to 2.42%.

At 100, 200, 400 and 800 rows the outcome-driven interval covers 70.92%, 50.42%, 21.17% and 2.42%. The bias barely moves — −0.16522, −0.16427, −0.16136, −0.16360, all within noise of the closed −0.163531 — while the standard error falls from 0.11879 to 0.04108. The ratio of the two runs −1.377, −1.973, −2.810 and −3.981, and coverage is essentially a function of that ratio alone.

So the interval is not merely optimistic. It converges, at the usual rate, onto a wrong answer, and it advertises increasing confidence in it. Under the regressor-driven rule the same four sizes give 95.33%, 96.17%, 95.25% and 93.92%: no trend, the ordinary non-monotone wobble a counted coverage has, and nothing that a larger study would fix or break.

There is no sample size at which a reader would notice. The estimate settles down, the interval narrows, every diagnostic a complete-data analysis offers is clean — because the fit really is a correct fit to the recorded rows, and the recorded rows really do follow a linear model with normal errors. The defect has no symptom. It is the shape of failure that survives longest for exactly that reason: nothing is wrong with what is there.

Coverage that sits a little above its promise

One column of the sweep is worth stopping on, because it runs the other way and a reader is entitled to be suspicious of it. Under the regressor-driven rule the counted coverage across the seven strengths reads 96.30%, 96.90%, 95.90%, 95.70%, 96.50%, 96.20% and 96.10% — consistently a little above ninety-five rather than at it, on a thousand studies whose binomial error is about 0.6 points.

That is not the mechanism. It is the interval. The number of recorded rows is itself random, the degrees of freedom of the t interval are read off it, and an interval whose width is set by a quantity that varies from study to study is slightly wider on average than the one a fixed sample size would give. The effect is small, it appears at every strength including zero, and it is present in the completely-random column too. A reading that treated the excess as evidence for the mechanism would be reading the shape of a t interval and calling it a finding.

What a complete-case interval covers, by mechanism. How often the ordinary 95% interval for the slope, computed from the complete cases alone, contains the value the data were generated at — counted over 4000 studies of 200 rows with 35.0% of the outcomes missing under each of four rules. Missingness that reads nothing covers 95.15%, that reads the regressor 95.40%, that reads the second covariate 94.80% — all three at the promise — and that reads the outcome itself covers 49.70%. The same fraction is missing in every row and the same index coefficient drives it; the only thing that changes is which variable the rule looks at.
Fig. 5 What the complete-case interval covers under each of the four rules at the field’s own settings, over four thousand studies. Three sit at the promise and the outcome-driven rule sits at 49.70%.

The four thousand studies behind that figure give the same three rules 95.15%, 95.40% and 94.80% and put the fourth at 49.70%, which is the reading in its blunt form: the 95% is a claim about a procedure, and here two procedures that differ in nothing a reader can see keep and break it.

What the exactness does not buy

The slope is one quantity out of a table full of them, and the protection covers only what it covers.

The estimand decides whether a repair is needed. The bias of four estimators of the outcome's population mean, when 35.0% of outcomes are missing at a rate that depends on the regressor, over 1500 studies of 200 rows. The mean of the complete cases is wrong by 0.3143 against a closed prediction of 0.3152, on the same rows whose fitted slope is exactly right. Weighting each complete case by the reciprocal of its chance of being observed removes most of it: 0.0183 ± 0.0051 with the true chances and 0.0229 ± 0.0044 with fitted ones — small, and several standard errors from zero, which is what consistent rather than unbiased looks like at this size. Pooling 20 imputations leaves -0.0009 ± 0.0030.
Fig. 6 Four estimators of the outcome’s population mean under the rule that reads the regressor. The complete-case mean, computed from the same rows whose slope is exactly right, is wrong by 0.31427.

On the same rows and the same draws that give an exact slope, the mean of the recorded outcomes is wrong by +0.31427 ± 0.00273, against a closed prediction of 0.315192. That is a third of a unit on a quantity whose true value is 1, and it comes from the identical arithmetic — the mean of the recorded outcomes is α+(βpx+γpz)λ\alpha + (\beta p_x + \gamma p_z)\lambda, which moves for precisely the reason the slope does not.

The same split runs through the whole table. The partial slope on xx is exact; the marginal slope of yy on xx alone is exact too under this rule, since the mechanism reads xx and both regressions condition on it; the mean of xx, the variance of xx, the mean of yy and the variance of yy are every one of them wrong. There is no general principle available beyond the one already stated — a quantity is protected when the analysis conditions on everything the mechanism reads, and a marginal summary conditions on nothing.

So an analysis of these rows can report a coefficient that is right to fourteen decimals and, three lines above it, a descriptive mean that is out by a third of a unit, with no warning that the two were computed under different protection. The estimand decides, and the four datasets that share a summary are the standing reminder that a table of numbers does not announce which of them the data support.

What could have made the exactness wrong

A rounding artefact. The closed biases are ±1.11×1016\pm 1.11 \times 10^{-16} and 00, alternating in sign across the sweep, which is what a quantity that is algebraically zero looks like after two matrix solves. If it were a genuinely small nonzero number it would have a consistent sign and would scale with the strength; it does neither.

A count too coarse to see it. A thousand studies per strength resolve the bias to about 0.003. That is not enough to rule out a bias of 0.0005 and is easily enough for the comparison actually being made, since the outcome-driven curve reaches −0.23323 at the same setting. The claim of exactness rests on the identity and the count is a check on the identity, not the other way round.

Normality doing the work. It is doing some of it. The truncated moments are closed because the joint law is normal, and the sweep is therefore a check of the algebra rather than of the generality. What is not normality-specific is the mechanism: the argument that selection on the conditioning variables leaves the conditional law alone needs only that the rule not read the residual, and it is the closed forms rather than the result that would go if the law changed.

Where it stops

Two limits, both real and neither hedging.

The analysis must be a conditional-mean analysis of the variables the rule reads. The whole result is a property of least squares conditioning on the regressor. A logistic outcome model, a quantile regression, or any summary of the marginal distribution of yy has no such protection, and the descriptive mean above is the smallest possible example of that failing.

The gaps must be in the outcome. Everything measured here deletes yy. When the missingness falls on xx instead, the selection acts on exactly the variable least squares conditions on, and the protection does not transfer — the recorded rows then carry a truncated regressor and the conditional law being estimated is a conditional law over a restricted range. The closed machinery covers that case, since the truncation identity does not care which variable is deleted, and the analysis of it is not attempted here. That is the largest single gap this argument leaves, and it is named rather than papered over, in the same spirit as the third state a censored observation occupies being given a slot of its own rather than forced into one of the two that already existed.

What is settled is narrower than “dropping incomplete rows is safe” and sharper. Dropping them costs nothing at all for a coefficient whose regressors the mechanism reads, however unrepresentative the survivors are, and costs everything for a coefficient whose mechanism reads the outcome — with the cost growing as the study grows. Which of those two a dataset is in is not a question the dataset answers.

The reading a practitioner can take from this is narrower still, and it is about what to write down rather than what to compute. Whether the complete cases are enough is decided by two things that are both stated before the data arrive: which variables the chance of being recorded plausibly depends on, and which of those the analysis conditions on. Neither is a diagnostic and neither is checkable after the fact — counting an advertised rate rather than asserting it is available here only because the truth is known by construction. In a real study the honest artefact is the pair of lists, and an analysis that names neither has made the assumption anyway.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Closed formComplete-caseConditional distributionConfidence intervalEffective sample sizeEstimandThe inverse Mills ratioLeast squaresMissing at randomMissing not at randomMissingness mechanismObservation propensitySampling variationSelection bias