The check worth more than the check
Worth reading first: What the 95% refers to · Two routes to every number.
The discipline this field is built on is that every important number is computed twice, by arithmetic that shares nothing, and the two are required to agree. The second route’s job is to catch the first one being wrong.
It has another job, and it is available on the same draws at no extra cost. If a quantity is being estimated by simulation and there is a companion quantity whose expectation is known exactly, the companion’s own simulated error is observable — it is the gap between what the draws gave and what the exact value is — and the part of the estimate’s error that moves with it can be subtracted off.
What that is worth is entirely a matter of how closely the companion tracks, and the range is wider than it sounds.
Four thousand draws producing the precision of eight hundred and fifty-five thousand, from a quantity the simulation had already computed.
The construction, and where the exactness enters
Suppose θ = E[f] is to be estimated by averaging f over draws, and h is a second quantity computed on the same draws with E[h] known. The adjusted estimate is
, with .
The bracket is the companion’s error on this particular run: it is observable, because E[h] is known, and it has mean zero, so subtracting any multiple of it leaves an estimate still centred on θ. Choosing the multiple that best predicts f’s own error removes as much of it as the companion can see, and what is left has variance
with ρ the correlation between f and h.
The known expectation is the whole of it. A companion whose expectation had to be estimated would contribute its own error and the subtraction would buy nothing. So the requirement is exactly the one this field’s verification discipline already imposes: a second route to a number that is exact, not merely independent.
It is worth saying what the adjustment is doing in plain terms, because the algebra hides a simple idea. A simulation of four thousand draws produced a particular set of samples, and those samples are not perfectly representative — by chance they contain slightly more successes than the truth, or slightly fewer. The companion measures that: its own average differs from its known expectation by a definite, observable amount, and that difference is a direct reading of how the draws were unrepresentative.
Having read it, the estimate can be corrected for it. If this run’s samples were rich in successes by a known amount, and the quantity being estimated rises with successes at a known rate, the quantity’s estimate is too high by the product — and subtracting it is the whole method.
What the correlation measures is how much of the run’s unrepresentativeness the companion manages to see. A companion that sees all of it removes all of the error; one that sees a quarter of it removes a quarter.
The curve, and why it is the shape it is
is flat near zero and steep near one, and both ends matter.
At a correlation of 0.3 the variance falls by 9%, which is worth 10% more draws — nothing. At 0.5 it falls by 25%, worth a third more draws. At 0.9 it falls by 81%, worth five times the draws. At 0.99 it falls by 98%, worth fifty times.
So a companion is nearly worthless until it is nearly perfect, and then it is worth everything. That shape is why the method is often described as marginal and occasionally described as transformative: both descriptions are right about different companions, and the distinguishing quantity is a correlation that can be measured on the same draws.
Two companions on one simulation
The two quantities in the first figure come from the same four thousand simulated samples of forty observations each, and both have exact values to be scored against — this field’s whole point.
The coverage of the interval. Its companion is the observed count, whose expectation is exactly np = 12. They correlate at 0.2665, because whether the interval covers depends on the count and does so in a lumpy, non-monotone way: the same count can be a covering count or a missing one depending on where it sits relative to the true proportion. The companion is worth 1.08 times the draws.
The expected width of the interval. Its companion is (1 − ), whose expectation is p(1 − p)(n − 1)/n = 0.20475 exactly. They correlate at 0.9977, because the width is — a monotone function of the companion, and the only reason the correlation is not one is that a square root is not linear. The companion is worth 214 times the draws.
The width estimate’s spread falls from to — a factor of 15 in standard error, which is what a 214-fold reduction in variance looks like on the scale a reader reads. The coverage estimate’s falls by 3.3%.
The refusal, which is what stops it being a free lunch
A method that turns four thousand draws into eight hundred and fifty thousand invites the question of where the information came from, and the answer is checkable.
Take a companion computed from a random stream the interval never touched — a normal draw with mean zero, generated alongside each sample and used for nothing. Its expectation is known exactly, which is the stated requirement, so the construction applies to it unchanged. It correlates with the coverage at −0.0044 and the variance that survives is 0.99998 of what it was.
Nothing was bought, and nothing should have been: the companion’s error carries no information about the estimate’s error, so subtracting a multiple of it removes nothing. The information the method uses is the correlation, which is a property of the pair and not of the companion’s exactness alone. Exactness makes the subtraction safe; correlation makes it useful, and both are required.
What decides which case a study is in
The difference between the two is not that one companion is better chosen. It is structural, and the structure is visible before any simulation is run.
A companion the quantity is a smooth function of buys nearly everything. The width is a function of (1 − ) and of nothing else, so the only part of its error the companion cannot see is the curvature of a square root over the range wanders in. The residual is small because the range is small.
A companion the quantity depends on only through a threshold buys almost nothing. Coverage is an indicator — the interval either contains the proportion or it does not — and an indicator is a step function of the count. A linear adjustment cannot follow a step, so most of the indicator’s variability is invisible to any linear companion, however well the companion is chosen.
The arithmetic of the indicator case is worth one line. A coverage indicator at n = 40 and p = 0.3 is one for counts in a contiguous band and zero outside it, so its correlation with the count is whatever the linear fit to a step function manages — here 0.2665, and it would be smaller still if the band were centred, because a symmetric step has no linear component at all. The companion is not badly chosen; there is nothing linear in the indicator for it to find.
That distinction is worth carrying because it says which quantities the method helps with. A smooth functional of the draws — an expected width, a mean squared error, an average length — has companions worth hundreds of draws. A rate, a rejection frequency, a coverage — anything that counts events — does not, and the reason is not remediable by a better companion.
What this says about the two-route discipline
The field this essay belongs to was built on a rule about checking, and the measurement changes what the rule is worth rather than what it requires.
The second route was never only insurance. Computing every important number twice was justified as catching errors, and the catching is real — the arcsine that closed a field is an exact companion found while looking for a check. What the control-variate reading adds is that where the second route is exact and the first is simulated, the pair is also an estimator better than either.
And the discipline pays best exactly where it is cheapest to follow. A quantity with a clean closed form is the one a second route is easiest to build for, and it is also the one whose companion correlates most strongly, because both facts come from the quantity being a smooth function of something summable. The cases where verification is hard are the cases where the adjustment would have been worth least.
That is an unusual alignment and it is worth naming: the effort and the reward point the same way, which is not true of most methodological discipline.
What a defensible use looks like
Choose the companion from the quantity’s own formula, not from the simulation. The width’s companion was found by looking at what the width is a function of, and the correlation followed. Searching a list of candidate companions for the one that happens to correlate best on the run in hand is a selection with a price of its own — it is a search whose charge has not been paid — and it biases the estimated coefficient towards whatever the noise favoured.
Report both routes even when the adjustment is not used. The verification the pairing was built for is the primary purpose and the sharpening is a by-product, so a study should not stop printing the check because it started using it as an estimator. This is the same distinction a diagnostic run before a standard error draws: what a number is for and what it is worth are two separate questions.
Compute the companion whether or not it is used. A second exact route is required here for verification anyway, so the correlation is free to measure and the decision whether to use it as a control costs one pass over the draws.
Use it on the smooth quantities and not on the rates. The rule follows from the structure rather than from a threshold on ρ: a quantity that is a smooth function of a summable companion gets a large reduction, and a quantity that counts events does not. That is knowable before the simulation runs, from the quantity’s own formula, and it is the same kind of reading a coverage counted exactly rather than simulated rests on — look at what the quantity is a function of, and only then decide how to compute it.
Report the correlation with the estimate. It is the only number that says what the adjustment was worth, and a reader shown an adjusted estimate without it cannot tell a 214-fold reduction from a 1.08-fold one. It is the same discipline a width reported beside a coverage asks for, applied to a simulation rather than to an interval.
And do not read a narrower estimate as a better one without checking the centring. The adjustment is unbiased by construction, and the construction relies on E[h] being right. A companion whose expectation was derived rather than known — a closed form with an algebra mistake in it — produces an estimate that is confidently wrong, and its narrowness is exactly what makes that hard to notice. The figures here score both adjusted estimates against exactly known truths for that reason.
What is claimed here and what is not
Both exact values are finite sums. The coverage and the expected width of a textbook interval for a proportion at n = 40 are sums over the forty-one possible counts, so there is no approximation anywhere in the scoring and the adjusted estimates are measured against numbers rather than against other estimates.
The coefficient is estimated from the same draws. c = Cov(f, h)/Var(h) is computed on the run it is applied to, which introduces a small bias of order 1/draws. At four thousand draws it is invisible and at forty it would not be; a study using very few draws should estimate the coefficient on a separate pilot run or accept the bias knowingly.
The two hundred and thirteen is at one setting. It is n = 40, p = 0.3, and the correlation depends on both: at a proportion near zero or one, (1 − ) has a longer tail and the square root’s curvature over its range matters more, so the correlation falls. The claim is not that a control variate for an expected width is always worth two hundred draws; it is that a companion’s worth is and that ρ can be this close to one when the quantity is a smooth function of the companion.
The uncorrelated companion is one draw, not a sweep. The −0.0044 comes from a single run of four thousand draws, where a correlation of exactly zero would show up as something of order . The reading is consistent with zero and is quoted to show the construction is inert there rather than to estimate a correlation that is zero by construction.
And nothing here is a claim about speed. The adjustment is one pass over quantities the simulation already computed, so its cost is negligible, but “worth 214 times the draws” is a statement about variance rather than about a duration. What it means is that reaching the adjusted estimate’s precision by drawing more samples would take 214 times as many of them.
What the adjustment is not
Three things the construction resembles and is not, because each has a different requirement and mistaking one for another is how a variance reduction becomes a bias.
It is not importance sampling. That changes the distribution the draws come from and reweights them; this leaves the draws exactly as they were and subtracts an observable quantity with mean zero. Nothing about the sampling changes, which is why the adjusted estimate is unbiased with no weights to go wrong.
It is not a regression adjustment in the design sense. Regressing an outcome on a covariate to reduce variance in an experiment is the same algebra and a different licence: there the covariate’s expectation is unknown and the argument is asymptotic. Here it is known exactly and the argument is arithmetic, which is what makes five draws’ worth of companion useful.
And it is not antithetic sampling, which pairs each draw with its mirror and relies on the quantity being monotone in the driving uniform. That is a different property of the same simulation, it composes with this one, and its worth is read off a different correlation.
The common thread is that all three trade something for variance. The one measured here trades nothing except the requirement that a second exact route exist — which is a requirement this field imposes anyway, for a different reason.
The number in units a reader has
One conversion makes the two hundred and fourteen concrete. A simulation’s standard error falls like , so matching the adjusted estimate’s precision by drawing more samples would take 855,634 of them rather than four thousand.
Where a single coverage estimate is routinely a few thousand draws and a sweep is a few hundred thousand, that is the difference between a simulation that finishes in seconds and one left to run overnight. It is also the difference between a number quoted to three digits and one quoted to five: the adjusted width estimate’s standard error is , so its fourth digit is solid where the plain estimate’s third was marginal.
Whether that precision is wanted is a separate question. A figure whose caption reads to two decimal places does not need five digits of precision behind it, and the honest use of a control variate is to buy the precision a claim needs rather than as much as is available.
Where else the pair already exists
The discipline this essay reads a second use into is already in place across the site, so it is worth naming where the pairs are, because each is a control variate nobody has used as one.
A coverage computed exactly by summation sits beside a simulated one wherever the parameter is discrete — the essay that counted a coverage exactly is built on that pairing. A closed-form variance inflation sits beside a counted one wherever a dependence is priced. An enumerated reference distribution sits beside a sampled one wherever a randomisation test is small enough to enumerate.
In each case the exact half is currently used to check the simulated half and then discarded. The measurement above says that where the two track closely, the exact half is also worth somewhere between a little and a great deal of extra precision, and that which of those it is can be read off a correlation computed on the draws already taken.
There is one place it would not help and the reason is instructive. Where the exact route is available for the whole quantity — a coverage for a proportion, say — the simulation is not needed at all, so the pairing is a check and nothing more. The method has room only where the exact companion is a related quantity rather than the quantity itself, which is the ordinary case and is the one worth looking for.
Still open: the companion for a rate
The method’s failure case is the one this site needs most. Almost every number on it is a rate — a coverage, a rejection frequency, an error rate — and the measurement above says a linear control variate is nearly useless for one, because an indicator is a step function of anything smooth.
Two repairs are available in principle and neither is measured here. The companion could be made non-linear — regressing the indicator on a function of the count rather than on the count — which recovers whatever a flexible fit can see and brings back a selection question about how flexible. Or the indicator could be replaced by its conditional expectation given something partially observable, which is a different technique with a different requirement and which, for a proportion, would amount to summing exactly over the thing being simulated and would make the simulation unnecessary.
That second observation is the interesting one and it points at a boundary rather than a technique: where the conditional expectation can be computed, the simulation is not needed, and where it cannot, the companion cannot see the step. Whether there is a useful middle — a quantity partially summable and partially simulated — is the question this field has not answered.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The same draws for both methods — both name binomial proportion, closed form, correlation, coverage, monte carlo, variance reduction
- An interval that covers and says nothing — both name binomial proportion, closed form, coverage, exact enumeration, interval width
- The arm whose variance is its answer — both name binomial proportion, closed form, efficiency, monte carlo, variance reduction
- A coverage table with its own error — both name binomial proportion, closed form, coverage, monte carlo
- A rate times a size — both name closed form, estimation error, monte carlo, variance reduction
- A simulation that stops when it looks settled — both name binomial proportion, closed form, coverage, monte carlo
Named objects
A flat tag is an object no other essay names yet.
Binomial proportionClosed formCorrelationCoverageEfficiencyEstimation errorExact enumerationInterval widthMonte CarloNumerical methodsRound tripSampling variationVariance reductionVerification