the-value-that-is-not-there

Three mechanisms leave the slope alone; one does not

The bias of the complete-case slope under each of four missingness rules, counted over 4000 studies of 200 rows at 35.0% missing, with the closed form printed beside each count. Missingness that depends on nothing, on the regressor, or on the second covariate leaves the slope exactly where it was — the closed forms are zero to machine precision and the counts are -0.0005, -0.0005 and -0.0011 against standard errors of about 0.0018. Missingness that depends on the outcome moves it by -0.1635, which is 27.3% of the slope being estimated. The same share of rows is lost in every case.

The value that is not therewide8 views

What else it draws

The same object, drawn to answer the other questions the essays put to it.

Three properties of the complete cases as the chance of being observed leans harder on the regressor, in closed form, at 35.0% of outcomes missing throughout. The mean of the regressor among the rows kept climbs from 0.0000 to 0.5528 against a population mean of zero, and the mean of the outcome from 0.0000 to 0.3980 above its own. The bias in the fitted slope is exactly zero at every one of the ten settings, because selection acting on the regressor alone leaves the conditional law of the outcome given the regressor untouched and least squares conditions on exactly that. The sample is wrong about almost everything and right about the one quantity being estimated.

What a 95% interval for the slope actually covers after each way of handling 35.0% missing outcomes, counted over 4000 studies of 200 rows. Dropping the incomplete rows covers 95.93%. Filling with the observed mean covers 13.85%, because the estimate itself has moved. Filling with a fitted value covers 80.85% against a closed prediction of 79.73%: the estimate is right and the reported standard error is short by a factor of 0.6567 against a predicted 0.6500, because the residual sum of squares is divided by the whole sample's degrees of freedom. Adding residual noise recovers the spread and covers 85.78% against a predicted 84.62%, since the interval still ignores the variance of having imputed at all.

What a pooled 95% interval covers against the number of imputations, counted over 2000 studies of 200 rows at 35.0% of outcomes missing. Rubin's rules — total variance W̄ + (1 + 1/m)B, read against a t distribution on (m − 1)(1 + W̄/((1 + 1/m)B))² degrees of freedom — cover 94.10% at two imputations and reach their promise by 5, at 95.25%. Dropping the (1 + 1/m) factor takes two imputations to 93.10%; using a normal quantile instead of the degrees-of-freedom correction takes it to 92.55%; dropping both takes it to 91.45%. The median degrees of freedom at two imputations is 12.95, which is why the second correction is the larger.

Where each coefficient lands when the model that fills the missing outcomes and the model that analyses them disagree, over 1500 studies of 200 rows at 35.0% missing and 20 imputations. An imputer that omits a covariate the analysis fits attenuates that covariate's coefficient by exactly the missing fraction — -0.1405 counted against a closed -0.1400 — and pushes the coefficient it did impute on the other way by exactly the product of the omitted coefficient, the covariates' correlation and the missing fraction: 0.0402 counted against 0.0420. Both closed forms come out of the same two-by-two solve. Matching models leave both alone, and so does an imputer that knows more than the analysis.

What the slope really is, against a shift in the outcomes nobody saw — line from the closed form, dots counted over 2000 studies of 200 rows at 35.0% missing. Every point on this line produces exactly the same observed data, and the complete-case estimate is the flat line at 0.5996 regardless. The truth moves at -0.2845 per unit of shift, which is a function of the missingness model and the missing fraction and of nothing that can be estimated: across the swept range the true slope runs from 0.8845 to 0.3155, a span of 0.5691 against a value of 0.60 in the world where the shift is zero. Reporting the line is the honest form of the answer.

The bias of four estimators of the outcome's population mean, when 35.0% of outcomes are missing at a rate that depends on the regressor, over 1500 studies of 200 rows. The mean of the complete cases is wrong by 0.3143 against a closed prediction of 0.3152, on the same rows whose fitted slope is exactly right. Weighting each complete case by the reciprocal of its chance of being observed removes most of it: 0.0183 ± 0.0051 with the true chances and 0.0229 ± 0.0044 with fitted ones — small, and several standard errors from zero, which is what consistent rather than unbiased looks like at this size. Pooling 20 imputations leaves -0.0009 ± 0.0030.

The profile log-likelihood of a selection model in cy, the coefficient that lets the chance of being recorded depend on the outcome itself, for one study of 800 rows whose missingness is at random, with residuals normal; every other parameter is maximised at each fixed value. The model assumes the outcome is normal given the covariates. The curve peaks at cy = 0.35, where the fitted slope is 0.839, and the values of cy within the 95% cut run from −0.13 to 0.75; the likelihood-ratio statistic against cy = 0 is 1.47. The study was drawn with cy = 0.00. With the outcome's law left free, every value of cy fits the recorded rows equally well and this curve would be flat: its curvature is the normal assumption.

Where it is used

7 essays draw this figure, each at the numbers its own argument is about, so the same picture answers 7 different questions.

All 80 figures

FieldsThreadsSeriesConceptsAll essaysSearch