The value that is not there

One imputation is not an observation

Three ways of filling a missing outcome, under a mechanism that makes dropping the rows beyond reproach. Filling with the observed mean covers 13.85%, filling with a fitted value covers 80.85%, adding noise covers 85.78%, and the thing all three were meant to improve on covers 95.93%.

Worth reading first: Three mechanisms and one dataset.

Thirty-five per cent of the outcomes are gone, at a rate that depends on nothing whatever. Dropping the incomplete rows is beyond reproach under that rule, and its interval covers 95.93%. Fill the gaps instead — with the mean of what was recorded, with a fitted value, or with a fitted value plus noise — and the same interval covers 13.85%, 80.85% and 85.78%.

A filled value is not an observation. What a 95% interval for the slope actually covers after each way of handling 35.0% missing outcomes, counted over 4000 studies of 200 rows. Dropping the incomplete rows covers 95.93%. Filling with the observed mean covers 13.85%, because the estimate itself has moved. Filling with a fitted value covers 80.85% against a closed prediction of 79.73%: the estimate is right and the reported standard error is short by a factor of 0.6567 against a predicted 0.6500, because the residual sum of squares is divided by the whole sample's degrees of freedom. Adding residual noise recovers the spread and covers 85.78% against a predicted 84.62%, since the interval still ignores the variance of having imputed at all.
Fig. 1 What a 95% interval for the slope actually contains after each way of handling the gaps, over four thousand studies. The only bar at its promise belongs to the method the other three exist to improve on.

Not one of the three reaches its nominal rate, and they fail for three different reasons. One moves the estimate. One leaves the estimate exactly where it was and understates its error by a factor the arithmetic can name. One repairs that and leaves a smaller version of the same fault. Separating the three is the whole of this argument, because “imputation is biased” and “imputation understates uncertainty” are different complaints and only one of them applies to any given rule.

A filled value is a number in a column

The reason all three fail, before any arithmetic, is that a completed dataset carries no record of which of its cells were measured. The analysis reads two hundred rows, computes a residual sum of squares over two hundred residuals, divides by two hundred minus three, and reports an interval built on that. Nothing in the table says that seventy of those residuals were manufactured by the same procedure that is now being checked against them.

That is a different complaint from bias, and it is the one that survives every improvement to the filling rule. A rule can be made to give the right point estimate — the second one below does, to the last digit — and to give the filled column the right variance and the right correlation — the third one does — and the interval is still computed as though the study had two hundred observations of the outcome when it had a hundred and thirty. The sample size is part of what is being imputed, and none of the three rules imputes it — which is the same currency an effective sample size is measured in when weights leave a study with fewer observations than rows, arrived at from the opposite direction.

The mechanism is chosen so that nothing else can be blamed

Every count here is taken under missingness that depends on nothing at all. That is not the interesting mechanism and it is deliberately not the interesting mechanism.

Three mechanisms leave the slope alone; one does not. The bias of the complete-case slope under each of four missingness rules, counted over 4000 studies of 200 rows at 35.0% missing, with the closed form printed beside each count. Missingness that depends on nothing, on the regressor, or on the second covariate leaves the slope exactly where it was — the closed forms are zero to machine precision and the counts are -0.0005, -0.0005 and -0.0011 against standard errors of about 0.0018. Missingness that depends on the outcome moves it by -0.1635, which is 27.3% of the slope being estimated. The same share of rows is lost in every case.
Fig. 2 The complete-case bias under four rules for what goes missing. The completely random rule, which is the one every count in this argument is taken under, sits at −0.00054 against a standard error of 0.00148.

Under a rule that reads nothing, the recorded rows are a plain random sample of the whole, the complete-case slope is exact — a counted −0.00054 against a standard error of 0.00148 — and its interval keeps its promise. So every shortfall below belongs to the filling. That matters because the usual argument for imputing is an argument about efficiency and about the rows a complete-case analysis discards the recorded halves of, and an efficiency argument is only worth having if the thing being made more efficient is still correct.

One number in every gap

The simplest fill puts the mean of the recorded outcomes into every hole. It is the rule that makes the arithmetic of the failure exact, and it is worth doing first for that reason.

Which fills move the estimate. The bias of the fitted slope after each way of dealing with 35.0% of the outcomes going missing completely at random, counted over 4000 studies of 200 rows. Dropping the incomplete rows leaves it at 0.0020 ± 0.0014. Filling every missing outcome with the observed mean pulls it to -0.2088, which is the missing fraction times the slope — -0.2100 predicted. Filling with a fitted value from the complete cases returns the complete-case estimate exactly, 0.0020, and adding residual noise leaves it at 0.0020. Two of the four repair the estimate, and the next figure asks what their intervals cover.
Fig. 3 Where the point estimate lands under each fill. Mean imputation moves the slope by −0.20883, which is the missing fraction times the slope; the other two leave it where the complete cases put it.

Every filled value is the same number, so it contributes nothing to the sum of squared deviations of the outcome column and nothing to the sum of cross-products with the regressor. Both sums keep the share of themselves that the recorded rows carry, and are then divided by the same nn as before. So

var(yfilled)=(1f)var(y),cov(x,yfilled)=(1f)cov(x,y),\operatorname{var}(y_{\text{filled}}) = (1-f)\operatorname{var}(y), \qquad \operatorname{cov}(x, y_{\text{filled}}) = (1-f)\operatorname{cov}(x, y),

and the fitted slope, being the second over the variance of a regressor that nothing was done to, comes out at (1f)(1-f) times what it should be. At f=0.35f = 0.35 that predicts a bias of exactly −0.21, and the count over four thousand studies is −0.20883 ± 0.00112.

The interval is then wrong for the ordinary reason: the estimate has moved and the interval is not wide enough to reach back. Coverage is 13.85%, and there is nothing subtle about it — the estimate is 0.20883 from the truth on a spread of 0.07115, so the interval would have to be about three times its width to reach.

What is subtle, and worth separating out, is that mean imputation’s reported standard error is very nearly honest about its own estimator. It reports 0.06660 against an actual spread of 0.07115, a ratio of 0.9361 — within seven per cent, and much closer to right than either of the other two fills manage. Mean imputation fails the first test and nearly passes the second. The two failure modes are genuinely separate, and a diagnostic built to catch one of them would clear this rule entirely: its intervals are close to the right width, drawn around the wrong centre. That is the distinction between the claim a 95% interval makes and the number it prints in its least ambiguous form.

One act, two different factors

The correlation does not fall by the same factor as the variance, and stating that carelessly is the most common way of getting mean imputation wrong.

The covariance falls by 1f1 - f and the variance of the filled column falls by 1f1 - f, so the correlation — which carries the filled column’s standard deviation in its denominator — falls by 1f\sqrt{1-f}. At the field’s own fraction that is 0.65 for the variance and the slope, and 0.806226 for the correlation. Counted, 0.6481 and 0.8076.

One act, two different factors. What survives mean imputation, as a share of the true quantity, against the share of outcomes filled — lines from the closed forms, dots counted over 1200 studies of 200 rows each. Replacing a fraction of a column by its own observed mean multiplies that column's variance by exactly one minus the fraction and multiplies its correlation with the regressor by the square root of the same number, because the correlation carries the filled column's standard deviation in its denominator and the covariance does not. At 35.0% filled the variance keeps 0.6500 of itself, the slope the same 0.6500, and the correlation 0.8062. The counts at that fraction are 0.5986 and 0.7764 at the nearest swept point.
Fig. 4 What survives mean imputation as a share of its true value, against the share of outcomes filled. Lines are closed and dots are counted, and the two curves are the same act read on two quantities.

Swept across seven fractions, both identities hold to the resolution of the count:

share filled slope, counted / closed variance correlation
0.1 0.8943 / 0.9000 0.8989 / 0.9000 0.9486 / 0.9487
0.2 0.7966 / 0.8000 0.7988 / 0.8000 0.8961 / 0.8944
0.3 0.6958 / 0.7000 0.6982 / 0.7000 0.8376 / 0.8367
0.4 0.5982 / 0.6000 0.5986 / 0.6000 0.7764 / 0.7746
0.5 0.4970 / 0.5000 0.4970 / 0.5000 0.7075 / 0.7071
0.65 0.3458 / 0.3500 0.3461 / 0.3500 0.5878 / 0.5916
0.8 0.1957 / 0.2000 0.1953 / 0.2000 0.4411 / 0.4472

The counted slope sits a few thousandths below its closed value at every fraction, consistently, and that is not noise — it is the ordinary downward bias of a ratio estimated in a finite sample, present in the complete-case column too and of no interest to the comparison. What the sweep establishes is the pair of exponents, and the two columns separate cleanly at every fraction: at 0.8 filled, the variance keeps a fifth of itself and the correlation keeps nearly half.

The practical consequence is that a filled dataset misleads two readers differently. One who looks at the regression coefficient sees it shrunk by 35%. One who looks at the correlation sees it shrunk by 19%. Both are looking at the same act, and neither number is a discount that could be undone by scaling, because the analyst does not know it was applied. This is attenuation of the kind that makes a selected group appear to improve — arithmetic doing exactly what it was told, with nothing to notice.

The fill that returns the estimate exactly

Fill each gap with the value the complete-case line predicts at that row’s covariates instead, and the point estimate stops being the problem.

It stops being the problem in the strongest possible sense. The complete-case bias is 0.00197 and the spread of the estimate is 0.09107; the regression-filled bias is 0.00197 and the spread is 0.09107. The same digits, because they are the same estimator. A filled row sits exactly on the fitted line, so its residual is zero and it contributes nothing to either normal equation: the line that would be fitted through the completed dataset is the line it was drawn from.

Those two rows of the table are the same estimate and they cover 95.93% and 80.85% — a gap of 15.08 points, every point of which is in the denominator of the reported standard error.

The comparison is worth stating with the other two columns in it. The complete-case analysis reports a standard error of 0.09284 against an actual spread of 0.09107, a ratio of 1.0194 — a shade conservative, for the ordinary reason that the number of recorded rows is itself random. The regression fill reports 0.05980 for the identical estimate. Two analyses of the same study, returning the same coefficient to fourteen decimals, differing only in whether seventy rows were written in, and reporting standard errors that differ by more than a third.

The reason is arithmetic rather than philosophy. The completed dataset has two hundred rows, so the residual sum of squares is divided by n3n - 3 rather than by the hundred and thirty that carry any residual at all, and the cross-product matrix is taken over all two hundred. The reported variance is short by (1f)2(1-f)^2 and the reported standard error by (1f)(1-f), which predicts 0.65; counted, 0.6567. Feeding that factor through a normal interval predicts a coverage of 79.7328%; counted, 80.85%.

This is the same defect as a forecast band derived for known parameters and then computed with estimates in them, and as a shrinkage weight that substitutes an estimated population spread and proceeds as though it were known. In every case the point estimate is defensible and the interval has forgotten a step. Here the forgotten step is that seventy of the two hundred rows are not data.

The correlation a fill invents

There is a reading in the filled dataset that goes the other way, and it is the one most likely to reassure somebody who should not be reassured.

Filling with fitted values raises the apparent correlation between regressor and outcome to 1.1258 times what the complete data would have given. Every filled point lies exactly on the line, so a third of the dataset now has zero scatter about the fit. The completed table looks like a cleaner measurement of a stronger relationship than the one the study actually produced.

It is worth being precise about the size of that. A hundred and thirty rows carry real scatter and seventy carry none, so the completed dataset’s residual variance is about (1f)(1-f) of the truth and its coefficient of determination is correspondingly inflated — the same direction of error as a predictor with no relationship to anything raising R² by 1/(n − 1), and by a far larger amount.

So a diagnostic run on the filled dataset does not merely fail to detect the problem; it reports better than the truth. The residual plot has a third of its points on the horizontal axis, the variance of the outcome reads 0.7917 of its true value, and every summary of fit improves. That is the inverse of the situation where a fit’s residuals are smoother than the errors that produced them — there the fit takes something out of what is left behind, here it puts something in — and both are cases of a diagnostic reading the analyst’s own arithmetic back to them.

Noise repairs the spread and leaves the smaller fault

The third rule adds a draw from the fitted residual scale to each fitted value. It is the obvious repair for the previous complaint and it does repair it: the filled column’s variance reads 1.0009 of the truth against 0.7917, the correlation reads 1.0009 against 1.1258, and the bias stays at 0.00205 ± 0.00161.

Coverage goes to 85.78%, which is better than 80.85% and is not 95%.

Two routes to the variance of one estimate. Rubin's total variance, averaged from within each study, against the variance of the pooled estimate measured across 2000 studies — the same quantity computed from two things that share no arithmetic. At 2 imputations the rules report 1.043e-2 against an actual 9.729e-3, a ratio of 1.0716; at 50 they report 8.840e-3 against 8.239e-3, a ratio of 1.0730. Both fall as the number of imputations rises, because the estimate itself gets tighter; the ratio does not, and its staying a few per cent above one is the conservatism the coverage figure shows as a rate slightly over its promise.
Fig. 5 Two routes to the variance of one estimate, computed from within each study and across studies. The gap between what a completed dataset reports and what the estimate actually does is the thing a single fill has no way to measure.

What is left is the variance of having imputed at all. The estimator is now the complete-case estimate plus a term driven by the noise added to the filled rows, and its actual variance is 1/(1f)+f1/(1-f) + f times the one it reports — 1.8885 at this fraction. Feeding that through predicts a coverage of 84.6202%; counted, 85.78%. So the added noise moves the estimator’s spread from 0.09107 to 0.10159 and moves the reported standard error from 0.05980 to 0.07432, and the second movement does not catch the first.

Two things about the third rule deserve saying plainly, because it is the one a reader is most likely to have been taught as the acceptable single fill. It is a genuine improvement: five points of coverage over the fitted value, an unbiased estimate, and a completed dataset whose second moments are right. And it is worse than the method it replaces on the only measurement that matters here, since dropping the rows covers 95.93% and this covers 85.78% while also being the wider estimator — 0.10159 of spread against 0.09107. It buys a completed table, which is a convenience, and pays for it in the thing the table is used to compute.

That residual gap is exactly the quantity a single completed dataset cannot compute, because computing it requires seeing how much the answer would have moved under a different draw, and one draw shows one answer. It is measurable only by taking several, which is the arithmetic the pooled rules are built on.

What could have made these readings wrong

The three fills could have been compared at different missing fractions. They are not: every row of the table comes from the same four thousand studies, the same rule, and the same 35%, and the complete-case row is computed on the identical draws.

The closed predictions could have been fitted after the fact. They are not fitted at all. Each is a function of ff alone, written down before the count and compared to it: 0.65 against 0.6567, 79.7328% against 80.85%, 84.6202% against 85.78%. Each count sits about a point above its prediction, in the same direction, which is the t interval’s own conservatism showing through — the predictions are computed on a normal quantile and the intervals are t intervals on the completed sample’s degrees of freedom, and the correction for not knowing the spread is worth about that much at this size.

The mean-imputation result could have been an artefact of a large fraction. It is not: the attenuation sweep runs from 0.1 to 0.8 and the identity holds at every point, and at 0.1 the slope still loses a tenth of itself.

The counts could be too noisy to order. They are not, and the margins are large compared with their errors: the four coverages carry binomial standard errors of 0.31, 0.55, 0.62 and 0.55 percentage points on four thousand studies, against gaps between them of five points at the closest and eighty-two at the widest. Nothing in the ordering is within reach of the resolution.

The ordering could reverse at another missing fraction. Two of the three rules have their failure written as a function of ff, so the direction is fixed by the algebra rather than by the setting: mean imputation’s bias is fβ-f\beta and the fitted value’s standard-error shortfall is 1f1-f, and both vanish only as ff does. What a smaller fraction changes is the size, not the sign. At a tenth missing, mean imputation still loses a tenth of the slope and the fitted value’s reported error is still short by a tenth, and the ranking of the three is the ranking of three functions that never cross on (0,1)(0,1).

A better single fill could exist. That is the reading the table is least able to refuse, and it should be stated as such. What is measured here is three rules, not the class of them. What the third row establishes is stronger than a comparison of three, though: an estimator whose point estimate is exactly right and whose completed dataset is distributionally correct still under-covers, because the defect is not in the filled values but in the sample size the analysis then believes it has.

Where the estimand decides again

Fitted weights beat the true ones. What each estimator of the outcome's population mean costs in spread and in total error, over 1500 studies of 200 rows at 35.0% missing on the regressor. The complete-case mean is the tightest of the four at 0.1058 and the worst by error at 0.3316, because it is tightly concentrated around the wrong number. Weighting by the true chance of being observed gives 0.1981; weighting by a chance fitted from the same data gives 0.1694, which is 1.37 times as efficient despite using an estimate where the other uses the truth. Twenty imputations give 0.1161, tighter than either weighting.
Fig. 6 What four estimators of a population mean cost in spread and in total error. The complete cases are the tightest of the four and the worst by error, because they are tightly concentrated around the wrong number.

One qualification runs underneath all of this and is easy to lose. Under a mechanism that reads nothing, dropping the rows is correct and everything here is a comparison of repairs to a problem that did not need repairing — the gain from filling is efficiency, and the table says the efficiency was bought with coverage. Under a mechanism that reads a covariate, the complete-case estimate is still exactly right for a slope and badly wrong for a mean, so the same four estimators separate by a third of a unit on one quantity and agree to three decimals on another.

That is why “should the gaps be filled?” has no answer without an estimand attached, and why the narrowest of four intervals being the one that misses is the reading to keep in mind when a completed dataset reports a smaller standard error than the recorded one did. Here the completed dataset reports 0.05980 against the recorded data’s 0.09284, on the same estimate, and the smaller number is the wrong one.

The reading that carries out of all four rows is a single sentence with two clauses, and the second is the one that is usually dropped. Filling the gaps can be made to leave the estimate alone; it cannot be made to leave the uncertainty alone, because a filled value carries the uncertainty of the model that produced it and a completed dataset has no column for that. Dropping the incomplete rows instead has the opposite profile — it is honest about how much it knows and throws away the recorded halves of the rows it discards — and the choice between them is not a choice between a careful method and a lazy one.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

AttenuationClosed formComplete-caseConfidence intervalDegrees of freedomImputationLeast squaresMean imputationMissing completely at randomParameter uncertaintyPlug in estimateRegression imputationSampling variationStandard error