Weights from an approximate law
Worth reading first: What the 95% refers to · Two routes to every number.
A companion with two coordinates ended with a clean division. A band weight computed exactly from the companion’s law keeps everything post-stratification offers; a band weight estimated from samples forfeits all of it, because the samples would have been worth more simply counted. Every companion in that line of essays had an exact law, because the simulations were built so that one existed — the logarithms of a lognormal sample are a normal sample, and their t statistic is exactly Student’s t.
Most simulations are not built that way. What they have is a companion whose law is known approximately — a normal limit, an Edgeworth correction, a saddlepoint — accurate, but not exact. A weight computed from such an approximation sits between the two cases the earlier essay separated. It carries an error, so it is not exact; the error is not sampling error, so it does not behave like an estimated weight either. The question this essay answers is what that error costs, and where.
The simulation is the one the field has used throughout: samples of twenty lognormal observations, the t test’s actual size at a nominal lower-tail 2.5%, which comes out at about 6.2%. The companion is the logarithms’ t statistic, cut into sixteen bands from below −4 to above zero. Its law is exactly Student’s t on nineteen degrees of freedom, so the exact weights are available as a reference — and so are three approximations a practitioner might use instead.
A bias the run multiplies
The arithmetic is short. Post-stratification estimates the rate as : each band’s observed rejection rate, weighted by the band’s probability. With the exact weights the estimate is unbiased and its variance is , where is the variance left inside the bands. With approximate weights the estimate converges instead to , and the gap
is a bias. Its mean squared error is , and its worth in multiples of the draws — the plain rate’s variance over its mean squared error — is
The bias is multiplied by the run. A larger simulation shrinks the variance that stratification attacks and leaves the bias exactly where it was, so every approximate weight has a run size beyond which stratifying on it is worth less than counting: .
With exact weights the bands are worth 2.709 times the draws at every run size. The three approximations are the standard normal, the normal rescaled to the t’s variance, and Fisher’s one-term expansion of the t distribution, . Reading the t statistic as standard normal makes the bands worth 0.762 at a thousand draws and 0.242 at four thousand; it breaks even with counting at 669 draws, below the size of almost any simulation. The variance-matched normal is worth 1.849 at four thousand and breaks even at 14,698. Fisher’s expansion is worth 2.708 at four thousand — indistinguishable from exact — and 2.523 at a million, and breaks even at 23.7 million draws.
So the three approximations are not three degrees of the same thing. One is useless at any realistic size, one is useful for a quick run and harmful for a careful one, and one is as good as the exact law for every simulation anybody is likely to run.
How wrong the weights are
The obvious guess is that the ranking follows how accurate each approximation is. Look at the weights themselves.
Every approximation is badly wrong in the outermost band, below −4, because a normal tail is thinner than a t tail on nineteen degrees of freedom. The standard normal puts −91.7% of the true weight there and is still −10.9% off in the eighth band. The variance-matched normal is −79.8% in the outermost band. Fisher’s expansion is −60.5% there, −7.5% in the fourth band, and within 2.1% from the fifth band on. All three are exact in the last band, from zero up, which holds half the probability and in which the t test almost never rejects — a symmetric approximation to a symmetric law cannot be wrong about the half above zero.
By the largest error in any band, Fisher’s expansion is the best of the three, and not by an order of magnitude: sixty per cent against eighty. By what it costs the run it is better by three orders. The weight errors do not explain the bias on their own.
Where the errors go
The bias is , and because both sets of weights sum to one it can be rewritten with any constant subtracted from the rates. It is a covariance — between the weight errors and the band rates — rather than a size. An approximation that is wrong in bands where the rate is zero costs nothing; one whose errors in high-rate bands cancel each other costs little; one whose errors all lean the same way where the rate is high costs the most.
Fisher’s expansion puts too little weight in the outer bands, where the rate is near one, and pushes the estimate down by 5.86 in units of 10⁻⁴. It puts slightly too much in the bands between −2.9 and −1.4, where the rate falls from 99% to 35%, and pushes it up by 5.46. The two nearly cancel, and the net bias is −0.39 — six hundredths of a per cent of the rate. The variance-matched normal pushes 34.27 up and 18.45 down, for a net 15.81, about 40 times Fisher’s. Its single-band errors are of the same order; what differs is that they do not cancel.
Both approximations lose weight in the far tail and must put it back somewhere, because both sum to one. What separates them is where it comes back, and the per-band numbers say exactly where.
Fisher’s expansion is short in the five outermost bands, below about −2.9, where essentially every sample falls in the rejection region, and over in the next five, from −2.9 to −1.4, where between 99% and 35% of samples do. From −1.4 up its errors are under a tenth of a per cent. So the weight it removes at a rate of one comes back at rates mostly above two thirds, and the two pushes nearly cancel.
The variance-matched normal is short in the six outermost bands, by up to 79.8%, and over in the next six, by up to 8.5% — and those six run from −2.6 to −0.9, where the rejection rate falls from 97% to 3%. Its largest surpluses sit in the bands where the rate is 88%, 66% and 35%. And it pays for part of that surplus by being short again in the three bands between −0.9 and zero, by up to 4.0%, where the rate is under one per cent. Weight moved from a band with a rate near zero to a band with a rate near a half is bias in its purest form, and it is what most of the rescaled normal’s 15.81 is made of.
The standard normal does the same thing more crudely: it is short in every band below −1.7, by amounts that reach −14.99 in units of 10⁻⁴ in a single band where the rate is 97%, and its surplus lands in bands where the rate is at most about a third. It moves probability out of the rejection region altogether, which is why it biases the rate by −11.9% of itself.
The measurement does not say why Fisher’s balance comes out so close. It says that the balance, not the size, is what an approximation has to be checked for before its weights are used — and that the check is one sum over the bands, with rates a short pilot can supply. A pilot cannot supply the weights, as the earlier essay showed; it can supply the rates the weight errors are multiplied by, and a few thousand samples are enough to see which way an approximation leans.
Computed and counted
The worth formula is first-order: it treats the band rates as known and ignores the empty bands a finite run leaves. Two routes to every number is the habit this field runs on, so the formula is checked against the simulation itself.
A pool of a million samples splits into 250 runs of four thousand, each post-stratified with each set of weights and scored by its mean squared error about the pool’s own rate. Exact weights are worth 2.709 computed and 2.574 counted. Fisher’s expansion: 2.708 and 2.566. The variance-matched normal: 1.849 and 1.834. The standard normal: 0.242 and 0.235. The counted values sit a little below the formula in every case, as they did for exact weights in the essay that cut the companion into many bands, because a run of four thousand leaves some outer bands empty and fills them from their neighbours. The ordering, and the size of each loss, are the formula’s.
The same arithmetic, applied to estimated weights
The bias formula also explains the earlier essay’s result about estimated weights from the other side. A weight estimated from an independent pilot of samples has an error too, and it can be written the same way: in every band, multiplied by the band’s rate. The difference is that the pilot’s errors are random, of both signs, and of a size that shrinks as . Averaged over pilots the bias is zero; for any one pilot it is not, and its square, averaged, is the variance of the band rates across bands divided by — exactly the second term the earlier essay found in its estimated-weight formula.
So the two cases are one case at two settings of the same dial. An approximation’s weight error is fixed and structured; a pilot’s is random and shrinks with the pilot. Both cost the run times their squared bias. The pilot’s version can be made small by spending samples, and the earlier essay showed the samples are better spent counted. The approximation’s version is free once the approximation exists, and whether it is small depends on how its errors line up — which is the question the figures above answer for three approximations and three patterns.
There is a third member of the family that the field has already met. A companion with an exactly known expectation used as a control variate needs that expectation exactly, and an approximate expectation biases the controlled estimate by the approximation’s error times the control’s coefficient. The arithmetic is the same: a fixed error, multiplied through to the estimate, competing with a variance the run is shrinking. Every variance-reduction device that leans on an exactly known quantity has a run size at which an approximate version of that quantity stops paying, and it is always the variance removed divided by the square of the bias carried.
A tolerance that tightens with the run
The three approximations are particular. The general question the earlier essay left — whether a weight with a saddlepoint’s accuracy, an error of a fraction of a per cent in each band, costs less than it saves — needs a weight error of controlled size and shape. So perturb the exact weights by a relative error δ in a stated pattern, renormalise so they still sum to one, and find the δ at which the stratified rate stops beating counting.
Three patterns bracket what an approximation might do. A constant relative error in every band below zero is the shape a saddlepoint’s error has, since a saddlepoint’s relative error is nearly flat across a tail; after renormalising it biases the rate by 0.50 times its own size, because the tail bands hold half the probability. Errors of random sign are what an approximation with no structure would do. Errors lined up with the band rates — up where the rate is above average, down where it is below — are the worst a pattern of a given size can do.
At four thousand draws a run can afford a constant tail error of 9.8%, random errors of 7.7% and aligned errors of 2.9%. At a million draws the same three tolerances are 0.61%, 0.48% and 0.18%. Every one falls as the square root of the run, because the bias is fixed and the variance it competes with falls as .
That answers the question directly. A weight with a saddlepoint’s flat relative error of half a per cent biases this rate by about a quarter of a per cent of itself, which a run can afford up to about a million and a half draws. For the simulations this field runs — thousands to tens of thousands of draws — a saddlepoint-accurate weight keeps essentially all of what exact weights keep. For a simulation of a hundred million draws it does not, and a hundred million is not an absurd size for a rate at one in ten thousand.
What this does to the method
The essays that put a rate beside a rate and then cut the companion into bands established that post-stratification on a companion with an exact law is close to free: it costs one integral per band and buys a multiple of the draws. The essay on two coordinates showed that estimated weights are worse than useless. Between the two there is now a measured middle.
An approximate weight is a bias, and the run is its multiplier. There is no approximation so good that a large enough simulation cannot expose it, and none so bad that a small enough simulation will notice. The useful statement is a run size, and for any approximation it is one division: the variance stratification removes over the square of the bias.
The bias is a covariance, not an error size. An approximation is cheap to use for weights if its errors cancel over the bands where the target rejects, and expensive if they lean one way there. Fisher’s one-term expansion is worse in its outermost band than a careful analyst would accept for a probability, and it is the best of the three by a factor of forty in bias because its errors change sign where the rate is high.
The tolerance depends on the run’s purpose as well as its size. A simulation whose rate is being compared with another simulation’s — two tests’ sizes, two designs’ power — can absorb a bias common to both and is limited only by the difference. A simulation whose rate is the answer cannot. The run sizes above are for the second kind.
And the check costs almost nothing. The bias is one sum over sixteen bands, the weight errors can be estimated by the gap between the approximation and its own next order, and the band rates come from the run that is about to be stratified. A simulation can therefore report, beside its stratified rate, the run size at which its own weights would stop paying — a number that says how far the result can be trusted before anyone asks.
Still open: an approximation that knows its own error
The next measurement is the repair the formula points at. If the bias is , and the run estimates each , then an approximation that comes with an estimate of its own error — an Edgeworth series with its next term, a saddlepoint with its second-order correction — gives a correction to the bias that the run can compute. The corrected estimator removes the bias at the cost of the error in the next term, which is smaller by a power of the companion’s sample size. Whether that correction keeps the stratified rate worth its exact-weight value out to runs of a hundred million draws, and what it costs when the companion is the score built on two coordinates rather than a statistic with a named law, has not been measured. It would say whether the method reaches the simulations that need it most, whose companions have only a tilt’s accuracy to offer.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- An ordering that depends on the rule — both name bias-variance, closed form, mean squared error, monte carlo
- The error no window repairs — both name bias-variance, closed form, mean squared error, monte carlo
- The gap a sample shows — both name bias-variance, closed form, mean squared error, monte carlo
- The length nobody has — both name bias-variance, closed form, mean squared error, monte carlo
- What choosing the length costs — both name bias-variance, closed form, mean squared error, monte carlo
- A covariate with no levels — both name closed form, monte carlo, stratification
Named objects
A flat tag is an object no other essay names yet.
Bias-varianceBinningClosed formEdgeworth expansionMean squared errorMonte CarloSaddlepoint approximationStratificationStudent's tVariance reduction