Missing at random — where it appears
Named by 4 essays across one field — each of them below, with the objects they name alongside it.
Also named here as missing not at random — the same set of essays touches all of them, so they are one junction rather than several.
Three mechanisms and one dataset
Four rules for which outcomes go missing, each calibrated to lose the same 35% of the rows and each leaning on what it reads with the same coefficient. Three leave the fitted slope exactly where it was, and the one that reads the outcome moves it by 0.163531.
Dropping the incomplete rows
Push the missingness until the rows that survive have a covariate mean of 0.543905 against a population zero and a variance of 0.5041 against one, and the fitted slope is still exactly right. Where the rule reads the outcome instead, the same sweep takes coverage to 2.42% at eight hundred rows.
The mechanism the data cannot see
Two worlds produce identical covariates, identical patterns of what is recorded and identical recorded outcomes, to the last bit. Their true slopes are 0.6 and 0.315452, and the truth moves at 0.284548 per unit of an assumption nothing in the data can inform.
The assumption that identifies the mechanism
A selection model estimates how strongly an outcome decides whether it is recorded — the quantity two identical datasets showed no statistic can see — and it does so by assuming the outcome is normal. Where that holds and the outcome does decide, it repairs a slope complete cases put at 0.4318 to 0.5795. Where the missingness is at random and the residual is merely skewed, it reports selection that is not there, moves the slope from 0.5971 to 1.0319, and rejects missingness at random in 72.5% of studies.
Named alongside it
The objects these essays reach for when they reach for this one.
Closed formComplete-caseConfidence intervalMissing not at randomEstimandMissingness mechanismNon-identifiabilityObservation propensitySelection modelConditional distributionThe inverse Mills ratioLeast squares