Missingness — the series
-
Three mechanisms and one dataset
Four rules for which outcomes go missing, each calibrated to lose the same 35% of the rows and each leaning on what it reads with the same coefficient. Three leave the fitted slope exactly where it was, and the one that reads the outcome moves it by 0.163531.
-
Dropping the incomplete rows
Push the missingness until the rows that survive have a covariate mean of 0.543905 against a population zero and a variance of 0.5041 against one, and the fitted slope is still exactly right. Where the rule reads the outcome instead, the same sweep takes coverage to 2.42% at eight hundred rows.
-
One imputation is not an observation
Three ways of filling a missing outcome, under a mechanism that makes dropping the rows beyond reproach. Filling with the observed mean covers 13.85%, filling with a fitted value covers 80.85%, adding noise covers 85.78%, and the thing all three were meant to improve on covers 95.93%.
-
The variance between imputations
Pooling several filled datasets covers 94.10% at two imputations and reaches its promise at five, where a single fill covered 85.78%. The correction everybody quotes is the smaller of the two doing the work — 1.00 ± 0.22 points against 1.55 ± 0.28.
-
An imputation model the analysis does not contain
A model that fills the gaps without a covariate the analysis fits attenuates that covariate's coefficient by exactly the missing share, 0.4 to 0.26, and moves the one it did carry by exactly γρf, 0.6 to 0.642. The reverse case is supposed to inflate the interval, and at four strengths of the extra knowledge it does not.
-
The mechanism the data cannot see
Two worlds produce identical covariates, identical patterns of what is recorded and identical recorded outcomes, to the last bit. Their true slopes are 0.6 and 0.315452, and the truth moves at 0.284548 per unit of an assumption nothing in the data can inform.
-
The assumption that identifies the mechanism
A selection model estimates how strongly an outcome decides whether it is recorded — the quantity two identical datasets showed no statistic can see — and it does so by assuming the outcome is normal. Where that holds and the outcome does decide, it repairs a slope complete cases put at 0.4318 to 0.5795. Where the missingness is at random and the residual is merely skewed, it reports selection that is not there, moves the slope from 0.5971 to 1.0319, and rejects missingness at random in 72.5% of studies.