The value that is not there

The variance between imputations

Pooling several filled datasets covers 94.10% at two imputations and reaches its promise at five, where a single fill covered 85.78%. The correction everybody quotes is the smaller of the two doing the work — 1.00 ± 0.22 points against 1.55 ± 0.28.

Worth reading first: Three mechanisms and one dataset.

Fill the gaps twice instead of once, analyse both completed datasets, and combine the two answers by Rubin’s rules: the interval covers 94.10%. At five imputations it covers 95.25% and at fifty 95.95%. The single stochastic fill it replaces — same mechanism, same 35% missing, same two hundred rows — covered 85.78%.

The extra 1/m, and the correction nobody quotes. What a pooled 95% interval covers against the number of imputations, counted over 2000 studies of 200 rows at 35.0% of outcomes missing. Rubin's rules — total variance W̄ + (1 + 1/m)B, read against a t distribution on (m − 1)(1 + W̄/((1 + 1/m)B))² degrees of freedom — cover 94.10% at two imputations and reach their promise by 5, at 95.25%. Dropping the (1 + 1/m) factor takes two imputations to 93.10%; using a normal quantile instead of the degrees-of-freedom correction takes it to 92.55%; dropping both takes it to 91.45%. The median degrees of freedom at two imputations is 12.95, which is why the second correction is the larger.
Fig. 1 What a pooled 95% interval covers against the number of imputations, with each of the two corrections removed in turn and with both removed. Counted over two thousand studies of two hundred rows.

Two imputations is not a lot, and the interesting reading is that it is nearly enough. The step from one fill to two buys eight points of coverage; the step from two to fifty buys under two. Almost everything the method delivers arrives with the second draw, because the second draw is the first moment at which the question “how much does this answer depend on which values were written in?” has an answer at all.

The half of this worth an argument is which piece of arithmetic is delivering it. The pooled variance carries two corrections, one of which is quoted in every description of the method and one of which is usually a footnote. Paired within study, the footnote is the bigger of the two.

The arithmetic, and what each term is

Take mm imputations, analyse each, and collect the mm estimates θ^j\hat\theta_j with their reported variances V^j\hat V_j. The pooled estimate is the average, and the pooled variance is

T=Wˉ+(1+1m)B,Wˉ=1mjV^j,B=1m1j(θ^jθˉ)2,T = \bar W + \left(1 + \tfrac{1}{m}\right) B, \qquad \bar W = \frac{1}{m}\sum_j \hat V_j, \qquad B = \frac{1}{m-1}\sum_j (\hat\theta_j - \bar\theta)^2,

read against a t distribution on

ν=(m1)(1+Wˉ(1+1/m)B) ⁣2\nu = (m-1)\left(1 + \frac{\bar W}{(1 + 1/m)B}\right)^{\!2}

degrees of freedom. Each of the three terms is doing a separate job.

Wˉ\bar W is what the analysis would have said had the filled values been real. It is the same quantity a single completed dataset reports, and here it is flat across every mm in the sweep — 0.005608, 0.005613, 0.005609, 0.005621, 0.005622 and 0.005620, from two imputations to fifty. That flatness is the point rather than a check: taking more draws does not change what any one completed dataset believes about itself. It changes what is known about the spread of those beliefs.

BB is that spread, and it is 0.003212 at two imputations against a Wˉ\bar W of 0.005608 — a little over half as much again on top of what the completed datasets report. It is the term a single fill has no way to compute, since a single squared difference needs two numbers.

The (1+1/m)(1 + 1/m) inflates BB because θˉ\bar\theta is itself an average of only mm draws, so the pooled estimate carries a Monte Carlo error of its own. And ν\nu is small when BB is a large share of TT, which is what makes the interval wide when the gaps are doing most of the work.

What makes an imputation proper

The whole construction depends on the filled values being draws rather than predictions, and the distinction is the same one that separates two other estimators on this site.

Each imputation here draws σ2\sigma^{*2} from an inverse chi-square on the residual degrees of freedom and then θ\theta^* from a normal around the least-squares estimate with that σ2\sigma^{*2}, before drawing the missing outcomes around θ\theta^*. Two imputations therefore differ by more than their residual noise: they differ by a whole redrawn model. An imputation that used the fitted θ^\hat\theta and σ^\hat\sigma every time would produce a BB that measured only the residual scatter and would understate the uncertainty in exactly the way a shrinkage weight that substitutes an estimated population spread and proceeds as though it were known does, or a forecast band derived for known parameters and computed with estimates in it.

That is what the extra 1/m1/m is ultimately paying for, and it is why the pooled interval can be honest at all. The completed datasets are not mm views of one answer; they are mm draws from a posterior, and averaging them is the same integration that makes a credible interval for one group of a hierarchy cover rather than a plug-in.

Two routes to one variance

Coverage is the reading anybody wants, and it is a blunt instrument: it says the interval is roughly the right width without saying whether the width was arrived at correctly. The sharper check computes the variance twice, from two things that share no arithmetic.

Two routes to the variance of one estimate. Rubin's total variance, averaged from within each study, against the variance of the pooled estimate measured across 2000 studies — the same quantity computed from two things that share no arithmetic. At 2 imputations the rules report 1.043e-2 against an actual 9.729e-3, a ratio of 1.0716; at 50 they report 8.840e-3 against 8.239e-3, a ratio of 1.0730. Both fall as the number of imputations rises, because the estimate itself gets tighter; the ratio does not, and its staying a few per cent above one is the conservatism the coverage figure shows as a rate slightly over its promise.
Fig. 2 Rubin’s total variance, averaged from within each study, against the variance of the pooled estimate measured across two thousand studies. Both fall as the number of imputations grows; the ratio between them does not.

TT is assembled inside each study from that study’s own mm analyses. The empirical variance is the spread of the pooled point estimate across the two thousand studies, which no single study can see. They read:

imputations Wˉ\bar W BB TT the estimate’s own variance ratio median ν\nu
2 0.005608 0.003212 0.010427 0.009729 1.0716 12.95
3 0.005613 0.003213 0.009896 0.009345 1.0590 17.10
5 0.005609 0.003246 0.009504 0.008899 1.0680 31.21
10 0.005621 0.003182 0.009121 0.008534 1.0688 70.54
20 0.005622 0.003150 0.008930 0.008327 1.0725 152.61
50 0.005620 0.003157 0.008840 0.008239 1.0730 390.22

Both columns fall as mm grows, because pooling more draws makes the estimate itself tighter, and the ratio does not — it sits between 1.0590 and 1.0730 throughout, which is the two-routes discipline this collection is built on applied to a quantity that is otherwise only checkable through a coverage rate.

It is worth saying what a disagreement would have meant, since a ratio near one is easy to accept without asking what it ruled out. If the (1+1/m)(1 + 1/m) term were wrong, the ratio would move with mm — the term is 1.5 at two imputations and 1.02 at fifty, so an error in it would show as a trend across the column and there is none. If Wˉ\bar W were the wrong within-study quantity, the ratio would be wrong by a constant and the coverage would be wrong in the same direction by a matching amount, which is a two-way check the coverage column alone cannot make. What the two routes agreeing establishes is that the decomposition is the right decomposition, separately from whether the interval built on it lands.

The pooled point estimate, meanwhile, is fine at every mm: its bias runs from −0.00035 at two imputations to −0.00245 at fifty, and neither exceeds 0.0025 on an estimand of 0.6. Whatever the rules are doing, they are not doing it to the centre.

Which correction is doing the work

Both corrections can be removed, separately and together, on the same draws — which is the honest way to price a correction, because the alternative is comparing two marginal rates that were computed on different studies.

What each of Rubin's two corrections is worth. How many points of coverage each correction buys, paired within study over 2000 studies of 200 rows at 35.0% missing. At two imputations, dropping the (1 + 1/m) factor from the total variance costs 1.00% ± 0.22%, and replacing the degrees-of-freedom correction by a normal quantile costs 1.55% ± 0.28%; dropping both costs 2.65%. The degrees-of-freedom correction is the larger at 2, 3 and 5 imputations, which is the reverse of which of the two is usually quoted; past that the two are inside their own standard errors of each other and the ordering is not resolved by 2000 studies. Both fall away as the number grows: at 50 imputations the pair is worth 0.20%.
Fig. 3 How many points of coverage each correction buys, paired within study. At two imputations the degrees-of-freedom correction is worth 1.55 points and the (1 + 1/m) factor 1.00.

At two imputations, dropping the (1+1/m)(1 + 1/m) from the total variance costs 1.00 ± 0.22 points of coverage. Replacing the degrees-of-freedom correction with a normal quantile costs 1.55 ± 0.28. Dropping both costs 2.65 ± 0.36.

The term that everybody quotes is the smaller one. The (1+1/m)(1 + 1/m) is the piece of Rubin’s rules that gets named in every description of the method — it is the visible departure from an ordinary variance decomposition, and it is easy to remember because it is easy to write. The correction that buys more is the one that changes which distribution the interval is read against.

The reason is in the last column of the table above. At two imputations the median degrees of freedom is 12.95. That is not a large-sample quantity: the t multiplier at thirteen degrees of freedom is about eight per cent above the normal’s, and eight per cent of width is worth more than the two or three per cent that inflating BB by a half buys when BB is a third of TT. It is the same effect that makes the correction for not knowing the spread worth twelve per cent of every interval at eight observations, arriving here from a completely different direction: the degrees of freedom are small not because the study is small but because the number of imputations is.

The degrees of freedom that has no mean

One column of the table is a median where every other is an average, and the substitution is not cosmetic.

At two imputations BB is a single squared difference between two numbers. Two draws that happen to agree closely give a BB near zero, and ν=(m1)(1+Wˉ/((1+1/m)B))2\nu = (m-1)(1 + \bar W/((1+1/m)B))^2 then blows up — so across two thousand studies the distribution of ν\nu has a tail that no finite average survives. The mean of that column would be set by whichever study drew the two most similar imputations, which is a statistic about one draw rather than about the method. The median is 12.95, and it is the number that says what the correction is worth.

That instability is itself part of what the interval is charging for. At two imputations the width is driven by a quantity estimated from one degree of freedom, so the intervals are not merely wider on average than a normal one — they are variably wider, study by study, in a way that a reader inspecting a single published interval has no way to see. The 95% belonging to the procedure rather than to the interval in front of anybody is a distinction that matters more here than in most places, because the procedure’s variability is unusually large and is invisible in any one run.

Where two imputations still come up short

Both corrections applied, two imputations cover 94.10% ± 0.53 rather than 95%. The shortfall is about nine tenths of a point and is nearly two of its own standard errors, so it is probably real and is certainly small.

What is left over is the same instability. Rubin’s ν\nu is a plug-in: it is computed from Wˉ\bar W and BB, both estimated, and an interval whose degrees of freedom are estimated from a single squared difference inherits an error that the t multiplier does not know about. The correction repairs the typical case — it is worth 1.55 points — and does not repair the tail of studies where BB came out unusually small and the interval came out unusually narrow around an estimate that had also moved. The same shape appears wherever a variance component is read off a handful of numbers, as it is when a hierarchical model reads a population spread off the distances between eight group means.

Three imputations take the shortfall to 0.15 points and five overshoot. So the practical reading is that the rules are approximately right from two and right from about five, and that the residual defect at two is a small-sample property of ν\nu rather than of the decomposition.

The claim that survives, and the one that does not

The ordering above is worth stating with its limits attached, because it is the sort of result that gets repeated past what the measurement supports.

At two imputations the gap between the two corrections is 0.55 points, on paired standard errors of 0.22 and 0.28 — a real separation. At three the two are 1.15 ± 0.24 and 1.25 ± 0.25, and at five 0.55 ± 0.17 and 0.60 ± 0.17: the degrees-of-freedom term is nominally ahead at both, by a tenth of a point and half a tenth, which is comfortably inside its own error. At ten and twenty the marginal ordering runs the other way, at 0.30 against 0.25 and 0.40 against 0.30 — differences of half a tenth of a point on errors of about 0.12, which is to say no difference at all. By fifty both are worth 0.15 points and neither is worth discussing.

So the claim that survives is narrow: both corrections are worth points, and the one usually named is not the bigger of the two at the number of imputations where either matters. The claim that does not survive is a decisive ordering across the sweep. Anywhere past five imputations, both corrections are small enough that which of them is larger is a question about Monte Carlo error rather than about Rubin’s rules, and reporting a ranking there would be reporting noise with a story attached.

Seven per cent conservative, and why that is not rounded away

The ratio column never reaches one from below. Rubin’s total variance runs about 7% above the actual sampling variance of the pooled estimate at every mm, and the coverage at large mm sits a little above its promise: 95.85% at ten, 95.70% at twenty, 95.95% at fifty, against binomial errors of about 0.45 points.

That is a real reading rather than a wobble, and the honest description of these rules in this world is mildly conservative, not exact. The excess is small enough to be worth having and large enough that a reader comparing widths across methods should know it is there. It is also the direction one would want: an interval that is seven per cent too wide in variance is about three and a half per cent too wide in half-width, and it is the failure that costs precision rather than the one that costs correctness.

What it is not is evidence of an error in the arithmetic, because the same 7% appears at two imputations and at fifty. A defect in the (1+1/m)(1 + 1/m) term would shrink with mm; a defect in ν\nu would shrink faster still, since ν\nu passes 390 by the end of the sweep. A constant offset across a twenty-five-fold range of mm is a property of the estimator rather than of the correction.

What a single fill could not compute

A filled value is not an observation. What a 95% interval for the slope actually covers after each way of handling 35.0% missing outcomes, counted over 4000 studies of 200 rows. Dropping the incomplete rows covers 95.93%. Filling with the observed mean covers 13.85%, because the estimate itself has moved. Filling with a fitted value covers 80.85% against a closed prediction of 79.73%: the estimate is right and the reported standard error is short by a factor of 0.6567 against a predicted 0.6500, because the residual sum of squares is divided by the whole sample's degrees of freedom. Adding residual noise recovers the spread and covers 85.78% against a predicted 84.62%, since the interval still ignores the variance of having imputed at all.
Fig. 4 The single fills, on the same mechanism and the same missing fraction. The best of the three covers 85.78%, and dropping the incomplete rows covers 95.93%.

Set the two tables beside each other. A single stochastic fill leaves the estimate unbiased, leaves the completed column’s variance and correlation right, and covers 85.78%; two draws of the same imputation model, pooled, cover 94.10%. Nothing about the filled values improved between those two lines. What changed is that the second procedure measured how far apart its own answers were and put that into the interval.

That is the sharpest available statement of what a filled value is and is not. The defect in a single fill is not that the numbers written in are wrong — they are drawn from the right model — it is that the analysis has no way to know they were written in. Rubin’s rules do not repair the filled values. They repair the arithmetic that reads them, by making the analysis run several times and charging for the disagreement.

The fraction of missing information is that charge, expressed as a share: (1+1/m)B/T(1 + 1/m)B / T, which settles at about 0.3555 across the sweep against a share of outcomes missing of 0.35. The two numbers agreeing is a property of this world rather than a general fact — the outcome is missing, the covariates are complete, and the imputation model is correct, so what is lost is close to what is absent — and in a world where the imputation model explained the outcome well it would be far below the missing share.

What this does not answer

Two things this argument opens and leaves.

How many imputations are worth taking. The sweep reads coverage against mm and stops there, and coverage is essentially settled by ten. The question a reader actually has is where to stop, and answering it needs the width of the pooled interval against mm and the Monte Carlo error of the pooled estimate itself, neither of which is measured here. The fraction of missing information settles at 0.3555 in this world, so a rule of thumb pinning mm to a hundred times that fraction would say thirty-five, against a sweep that stops at fifty and shows nothing moving past ten. Whether that rule is about coverage or about the reproducibility of a published number is not settled here, and the two would give different answers. A width sweep would also be the place to price what the degrees-of-freedom correction costs, since an exact partition of the degrees of freedom in a two-arm design shows how easily a count that looks like bookkeeping turns into a width.

Whether the model that fills has to be the model that analyses. Everything above imputes and analyses under the same specification, which is the case where the rules have the easiest job.

Where an uncongenial pairing loses its promise. What each pooled 95% interval covers, over 1500 studies of 200 rows at 35.0% missing and 20 imputations apiece. When the imputation model and the analysis model agree, both coefficients cover their promise — 95.80% and 95.07%. When the imputer leaves out a covariate the analysis fits, the interval for the coefficient it left out covers 72.00% and the interval for the one it kept covers 94.73% — the damage is largest where the model was short and does not stay there. When the imputer knows a covariate the analysis omits, the interval covers 95.07%, which is its promise.
Fig. 5 What the pooled interval covers when the model that fills and the model that analyses disagree. The matching pairing keeps its promise at 95.80% and the interval for a coefficient the imputer omitted covers 72.00%.

Break that and the pooled interval for a coefficient the imputation model left out covers 72.00%, with the same twenty imputations and the same arithmetic — so the correctness demonstrated above is conditional on a matching pair of models, and the condition is not a technicality.

Fitted weights beat the true ones. What each estimator of the outcome's population mean costs in spread and in total error, over 1500 studies of 200 rows at 35.0% missing on the regressor. The complete-case mean is the tightest of the four at 0.1058 and the worst by error at 0.3316, because it is tightly concentrated around the wrong number. Weighting by the true chance of being observed gives 0.1981; weighting by a chance fitted from the same data gives 0.1694, which is 1.37 times as efficient despite using an estimate where the other uses the truth. Twenty imputations give 0.1161, tighter than either weighting.
Fig. 6 What four estimators of a population mean cost in spread and in total error under a mechanism that reads a covariate. Twenty imputations give 0.11610, tighter than either form of weighting.

And nothing above compares the method to its alternatives on cost. Under a mechanism that reads a covariate, twenty imputations estimate the outcome’s mean with a spread of 0.11610 against 0.16945 for weighting by a fitted chance of being recorded — tighter, on the same draws, which is an argument for imputing that has nothing to do with coverage and is not made anywhere above.

The last thing left out is the mechanism. Everything counted here is missing completely at random, chosen so that the pooled interval’s failures belong to the pooling — and under that rule dropping the incomplete rows was already correct, so the whole apparatus is buying efficiency rather than correctness. Under a rule that reads the outcome itself, the imputation model is fitted to the recorded rows and extrapolates a relationship those rows do not carry, and no amount of pooling repairs a model estimated from the wrong population. Rubin’s rules are honest about the uncertainty of an imputation model; they say nothing about whether that model is the right one.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Between imputation varianceConfidence intervalDegrees of freedomFraction of missing informationImputationMonte CarloMultiple imputationParameter uncertaintyPlug in estimatePosterior meanRubin's rulesT intervalVariance decompositionWithin imputation variance