The variance between imputations
Worth reading first: Three mechanisms and one dataset.
Fill the gaps twice instead of once, analyse both completed datasets, and combine the two answers by Rubin’s rules: the interval covers 94.10%. At five imputations it covers 95.25% and at fifty 95.95%. The single stochastic fill it replaces — same mechanism, same 35% missing, same two hundred rows — covered 85.78%.
Two imputations is not a lot, and the interesting reading is that it is nearly enough. The step from one fill to two buys eight points of coverage; the step from two to fifty buys under two. Almost everything the method delivers arrives with the second draw, because the second draw is the first moment at which the question “how much does this answer depend on which values were written in?” has an answer at all.
The half of this worth an argument is which piece of arithmetic is delivering it. The pooled variance carries two corrections, one of which is quoted in every description of the method and one of which is usually a footnote. Paired within study, the footnote is the bigger of the two.
The arithmetic, and what each term is
Take imputations, analyse each, and collect the estimates with their reported variances . The pooled estimate is the average, and the pooled variance is
read against a t distribution on
degrees of freedom. Each of the three terms is doing a separate job.
is what the analysis would have said had the filled values been real. It is the same quantity a single completed dataset reports, and here it is flat across every in the sweep — 0.005608, 0.005613, 0.005609, 0.005621, 0.005622 and 0.005620, from two imputations to fifty. That flatness is the point rather than a check: taking more draws does not change what any one completed dataset believes about itself. It changes what is known about the spread of those beliefs.
is that spread, and it is 0.003212 at two imputations against a of 0.005608 — a little over half as much again on top of what the completed datasets report. It is the term a single fill has no way to compute, since a single squared difference needs two numbers.
The inflates because is itself an average of only draws, so the pooled estimate carries a Monte Carlo error of its own. And is small when is a large share of , which is what makes the interval wide when the gaps are doing most of the work.
What makes an imputation proper
The whole construction depends on the filled values being draws rather than predictions, and the distinction is the same one that separates two other estimators on this site.
Each imputation here draws from an inverse chi-square on the residual degrees of freedom and then from a normal around the least-squares estimate with that , before drawing the missing outcomes around . Two imputations therefore differ by more than their residual noise: they differ by a whole redrawn model. An imputation that used the fitted and every time would produce a that measured only the residual scatter and would understate the uncertainty in exactly the way a shrinkage weight that substitutes an estimated population spread and proceeds as though it were known does, or a forecast band derived for known parameters and computed with estimates in it.
That is what the extra is ultimately paying for, and it is why the pooled interval can be honest at all. The completed datasets are not views of one answer; they are draws from a posterior, and averaging them is the same integration that makes a credible interval for one group of a hierarchy cover rather than a plug-in.
Two routes to one variance
Coverage is the reading anybody wants, and it is a blunt instrument: it says the interval is roughly the right width without saying whether the width was arrived at correctly. The sharper check computes the variance twice, from two things that share no arithmetic.
is assembled inside each study from that study’s own analyses. The empirical variance is the spread of the pooled point estimate across the two thousand studies, which no single study can see. They read:
| imputations | the estimate’s own variance | ratio | median | |||
|---|---|---|---|---|---|---|
| 2 | 0.005608 | 0.003212 | 0.010427 | 0.009729 | 1.0716 | 12.95 |
| 3 | 0.005613 | 0.003213 | 0.009896 | 0.009345 | 1.0590 | 17.10 |
| 5 | 0.005609 | 0.003246 | 0.009504 | 0.008899 | 1.0680 | 31.21 |
| 10 | 0.005621 | 0.003182 | 0.009121 | 0.008534 | 1.0688 | 70.54 |
| 20 | 0.005622 | 0.003150 | 0.008930 | 0.008327 | 1.0725 | 152.61 |
| 50 | 0.005620 | 0.003157 | 0.008840 | 0.008239 | 1.0730 | 390.22 |
Both columns fall as grows, because pooling more draws makes the estimate itself tighter, and the ratio does not — it sits between 1.0590 and 1.0730 throughout, which is the two-routes discipline this collection is built on applied to a quantity that is otherwise only checkable through a coverage rate.
It is worth saying what a disagreement would have meant, since a ratio near one is easy to accept without asking what it ruled out. If the term were wrong, the ratio would move with — the term is 1.5 at two imputations and 1.02 at fifty, so an error in it would show as a trend across the column and there is none. If were the wrong within-study quantity, the ratio would be wrong by a constant and the coverage would be wrong in the same direction by a matching amount, which is a two-way check the coverage column alone cannot make. What the two routes agreeing establishes is that the decomposition is the right decomposition, separately from whether the interval built on it lands.
The pooled point estimate, meanwhile, is fine at every : its bias runs from −0.00035 at two imputations to −0.00245 at fifty, and neither exceeds 0.0025 on an estimand of 0.6. Whatever the rules are doing, they are not doing it to the centre.
Which correction is doing the work
Both corrections can be removed, separately and together, on the same draws — which is the honest way to price a correction, because the alternative is comparing two marginal rates that were computed on different studies.
At two imputations, dropping the from the total variance costs 1.00 ± 0.22 points of coverage. Replacing the degrees-of-freedom correction with a normal quantile costs 1.55 ± 0.28. Dropping both costs 2.65 ± 0.36.
The term that everybody quotes is the smaller one. The is the piece of Rubin’s rules that gets named in every description of the method — it is the visible departure from an ordinary variance decomposition, and it is easy to remember because it is easy to write. The correction that buys more is the one that changes which distribution the interval is read against.
The reason is in the last column of the table above. At two imputations the median degrees of freedom is 12.95. That is not a large-sample quantity: the t multiplier at thirteen degrees of freedom is about eight per cent above the normal’s, and eight per cent of width is worth more than the two or three per cent that inflating by a half buys when is a third of . It is the same effect that makes the correction for not knowing the spread worth twelve per cent of every interval at eight observations, arriving here from a completely different direction: the degrees of freedom are small not because the study is small but because the number of imputations is.
The degrees of freedom that has no mean
One column of the table is a median where every other is an average, and the substitution is not cosmetic.
At two imputations is a single squared difference between two numbers. Two draws that happen to agree closely give a near zero, and then blows up — so across two thousand studies the distribution of has a tail that no finite average survives. The mean of that column would be set by whichever study drew the two most similar imputations, which is a statistic about one draw rather than about the method. The median is 12.95, and it is the number that says what the correction is worth.
That instability is itself part of what the interval is charging for. At two imputations the width is driven by a quantity estimated from one degree of freedom, so the intervals are not merely wider on average than a normal one — they are variably wider, study by study, in a way that a reader inspecting a single published interval has no way to see. The 95% belonging to the procedure rather than to the interval in front of anybody is a distinction that matters more here than in most places, because the procedure’s variability is unusually large and is invisible in any one run.
Where two imputations still come up short
Both corrections applied, two imputations cover 94.10% ± 0.53 rather than 95%. The shortfall is about nine tenths of a point and is nearly two of its own standard errors, so it is probably real and is certainly small.
What is left over is the same instability. Rubin’s is a plug-in: it is computed from and , both estimated, and an interval whose degrees of freedom are estimated from a single squared difference inherits an error that the t multiplier does not know about. The correction repairs the typical case — it is worth 1.55 points — and does not repair the tail of studies where came out unusually small and the interval came out unusually narrow around an estimate that had also moved. The same shape appears wherever a variance component is read off a handful of numbers, as it is when a hierarchical model reads a population spread off the distances between eight group means.
Three imputations take the shortfall to 0.15 points and five overshoot. So the practical reading is that the rules are approximately right from two and right from about five, and that the residual defect at two is a small-sample property of rather than of the decomposition.
The claim that survives, and the one that does not
The ordering above is worth stating with its limits attached, because it is the sort of result that gets repeated past what the measurement supports.
At two imputations the gap between the two corrections is 0.55 points, on paired standard errors of 0.22 and 0.28 — a real separation. At three the two are 1.15 ± 0.24 and 1.25 ± 0.25, and at five 0.55 ± 0.17 and 0.60 ± 0.17: the degrees-of-freedom term is nominally ahead at both, by a tenth of a point and half a tenth, which is comfortably inside its own error. At ten and twenty the marginal ordering runs the other way, at 0.30 against 0.25 and 0.40 against 0.30 — differences of half a tenth of a point on errors of about 0.12, which is to say no difference at all. By fifty both are worth 0.15 points and neither is worth discussing.
So the claim that survives is narrow: both corrections are worth points, and the one usually named is not the bigger of the two at the number of imputations where either matters. The claim that does not survive is a decisive ordering across the sweep. Anywhere past five imputations, both corrections are small enough that which of them is larger is a question about Monte Carlo error rather than about Rubin’s rules, and reporting a ranking there would be reporting noise with a story attached.
Seven per cent conservative, and why that is not rounded away
The ratio column never reaches one from below. Rubin’s total variance runs about 7% above the actual sampling variance of the pooled estimate at every , and the coverage at large sits a little above its promise: 95.85% at ten, 95.70% at twenty, 95.95% at fifty, against binomial errors of about 0.45 points.
That is a real reading rather than a wobble, and the honest description of these rules in this world is mildly conservative, not exact. The excess is small enough to be worth having and large enough that a reader comparing widths across methods should know it is there. It is also the direction one would want: an interval that is seven per cent too wide in variance is about three and a half per cent too wide in half-width, and it is the failure that costs precision rather than the one that costs correctness.
What it is not is evidence of an error in the arithmetic, because the same 7% appears at two imputations and at fifty. A defect in the term would shrink with ; a defect in would shrink faster still, since passes 390 by the end of the sweep. A constant offset across a twenty-five-fold range of is a property of the estimator rather than of the correction.
What a single fill could not compute
Set the two tables beside each other. A single stochastic fill leaves the estimate unbiased, leaves the completed column’s variance and correlation right, and covers 85.78%; two draws of the same imputation model, pooled, cover 94.10%. Nothing about the filled values improved between those two lines. What changed is that the second procedure measured how far apart its own answers were and put that into the interval.
That is the sharpest available statement of what a filled value is and is not. The defect in a single fill is not that the numbers written in are wrong — they are drawn from the right model — it is that the analysis has no way to know they were written in. Rubin’s rules do not repair the filled values. They repair the arithmetic that reads them, by making the analysis run several times and charging for the disagreement.
The fraction of missing information is that charge, expressed as a share: , which settles at about 0.3555 across the sweep against a share of outcomes missing of 0.35. The two numbers agreeing is a property of this world rather than a general fact — the outcome is missing, the covariates are complete, and the imputation model is correct, so what is lost is close to what is absent — and in a world where the imputation model explained the outcome well it would be far below the missing share.
What this does not answer
Two things this argument opens and leaves.
How many imputations are worth taking. The sweep reads coverage against and stops there, and coverage is essentially settled by ten. The question a reader actually has is where to stop, and answering it needs the width of the pooled interval against and the Monte Carlo error of the pooled estimate itself, neither of which is measured here. The fraction of missing information settles at 0.3555 in this world, so a rule of thumb pinning to a hundred times that fraction would say thirty-five, against a sweep that stops at fifty and shows nothing moving past ten. Whether that rule is about coverage or about the reproducibility of a published number is not settled here, and the two would give different answers. A width sweep would also be the place to price what the degrees-of-freedom correction costs, since an exact partition of the degrees of freedom in a two-arm design shows how easily a count that looks like bookkeeping turns into a width.
Whether the model that fills has to be the model that analyses. Everything above imputes and analyses under the same specification, which is the case where the rules have the easiest job.
Break that and the pooled interval for a coefficient the imputation model left out covers 72.00%, with the same twenty imputations and the same arithmetic — so the correctness demonstrated above is conditional on a matching pair of models, and the condition is not a technicality.
And nothing above compares the method to its alternatives on cost. Under a mechanism that reads a covariate, twenty imputations estimate the outcome’s mean with a spread of 0.11610 against 0.16945 for weighting by a fitted chance of being recorded — tighter, on the same draws, which is an argument for imputing that has nothing to do with coverage and is not made anywhere above.
The last thing left out is the mechanism. Everything counted here is missing completely at random, chosen so that the pooled interval’s failures belong to the pooling — and under that rule dropping the incomplete rows was already correct, so the whole apparatus is buying efficiency rather than correctness. Under a rule that reads the outcome itself, the imputation model is fitted to the recorded rows and extrapolates a relationship those rows do not carry, and no amount of pooling repairs a model estimated from the wrong population. Rubin’s rules are honest about the uncertainty of an imputation model; they say nothing about whether that model is the right one.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A block size that changes — both name confidence interval, degrees of freedom, monte carlo
- A coverage table with its own error — both name confidence interval, monte carlo, t interval
- A lag the sample has less of — both name degrees of freedom, monte carlo, plug in estimate
- A line that beats two curves — both name degrees of freedom, monte carlo, parameter uncertainty
- A schedule that reads the mean — both name confidence interval, degrees of freedom, monte carlo
- Correcting the persistence — both name monte carlo, parameter uncertainty, plug in estimate
Named objects
A flat tag is an object no other essay names yet.
Between imputation varianceConfidence intervalDegrees of freedomFraction of missing informationImputationMonte CarloMultiple imputationParameter uncertaintyPlug in estimatePosterior meanRubin's rulesT intervalVariance decompositionWithin imputation variance