A companion cut into many
Worth reading first: What the 95% refers to · Two routes to every number.
A rate needs a rate beside it found that the right companion for a simulated rate is a step, not a line. A companion used as a straight line can remove at most a fifth of a 5% rate’s variance however closely it tracks the statistic, and the same companion’s own rate at the target’s threshold removes 38.2% at a correlation of 0.9. It also found that a single step leaves something behind: the best any function of the companion can do at that correlation is 48.1%, and that best function is not a step but a smooth curve — the probability that the target event happens, given the value of the companion.
The way to follow a curve without fitting one is to cut the companion in several places, estimate the target’s rate separately in each band, and add the bands back together weighted by their exact probabilities. That is post-stratification, and on a companion whose distribution is known it needs nothing estimated except the rates inside the bands.
What a band estimator is, exactly
Cut the companion’s line at points . Each band has a probability known exactly, because the companion’s distribution is. In a run of draws, the estimator takes the share of draws in band that had the target event, multiplies by , and sums over the bands. The plain rate weights each band by the share of draws that happened to land in it; the band estimator replaces that share by the true one, and the difference is exactly the part of the plain rate’s error that came from the draws being unrepresentative of the companion.
Its variance per draw, to first order in the number of draws, is
against for the plain rate, and for bivariate normal and every is a difference of two orthant probabilities divided by . So the share of the variance the bands remove is exact, for any number of bands and any placement, and it is
Two cases anchor it. With a single cut at the target’s own threshold, the band estimator is the control variate on the companion’s rate — the same estimator to first order — and removes the same 38.2%. With infinitely many bands, becomes the conditional probability itself and the share becomes the best any function of the companion can do. Between the two, the bands follow the curve as finely as their number and placement allow.
Where the cuts go
The cheapest choice of bands is equal probabilities of the companion: the quartiles, the deciles, the sixty-fourths. It is what a histogram does, and it wastes nearly all its cuts.
For a 5% rate and a companion correlated 0.9, the conditional probability is essentially zero for every companion value below 0.70 and essentially one above 2.95; it climbs from 1% to 99% across that stretch, and nowhere else does it change. A cut placed where the curve is flat divides a band in which the target’s rate is the same on both sides, and dividing it buys nothing. Eight equal-probability strata put six of their seven cuts below 0.70, in the flat, and leave one to describe the whole rise.
Spending the cuts across the rise — equally spaced from to — follows the curve instead. At eight strata the band cuts remove 46.8% of the variance and the equal-probability cuts 29.9%. The band cuts reach 95% of the best any function can do at eight strata; equal-probability strata need thirty-two.
The rule is the one where a cut belongs kept arriving at in this field from other directions: a cut’s worth is decided by what changes across it. A median split of a covariate is informative about an outcome because the outcome’s expectation differs on the two sides; a cut of a companion is informative about a rate for the same reason, and a cut in the flat of the conditional probability is a median split of a quantity nothing depends on.
How the ceiling moves with the companion
The best a companion can do is set by its correlation with the statistic, and the bands approach it at a pace that depends on the same number.
The gap the bands exist to close is the one between the upper two curves, and it is widest in the middle of the range. At weak correlations there is little for any companion to remove, and at correlations near one the companion’s own rate already sits close to the best. A single step leaves 40.1% of the achievable reduction behind at a correlation of 0.6, 20.6% at 0.9 and 14.9% at 0.95.
At a correlation of 0.7 the best function removes 19.8% of a 5% rate’s variance, and eight band cuts remove 18.1% of it; the companion’s own rate, one cut, removes 13.0%. At 0.97 the best is 70.5% and eight band cuts reach 69.6%, against 62.2% for the single cut. In every case the steep part of the gain is from one cut to eight, and by sixteen band strata the curve is within a fraction of a percentage point of its ceiling.
The equal-probability strata catch up only slowly, and more slowly the stronger the companion, because a stronger companion’s conditional probability rises over a narrower stretch and an equal-probability grid lands fewer cuts on it. At 0.97, sixteen equal-probability strata remove 61.4% — less than the single cut at the target’s own threshold — and they need sixty-four to come within a percentage point of the best. A grid that ignores where the rate lives is beaten by a single well-placed step until it has several times more cuts than the rise can use.
On a real simulation, where the bands have to be filled
The exact arithmetic assumes every band’s rate is estimated with the precision a band of its probability deserves. A simulation has a fixed number of draws, and each band’s rate is estimated from the draws that fall in it. With few bands that is no constraint; with many, some bands receive a handful of draws and some receive none, and the estimator has to do something about a band it has no draws for.
The companion used for a rate beside a rate is the one to test it on: the t test’s lower-tail size on samples of twenty lognormal observations, 6.1% at a nominal 2.5%, with the t statistic of the logarithms as the companion, exactly Student’s t on nineteen degrees of freedom. Its bands were placed two ways — across the lower stretch from −4 to 0, where the data’s test rejects, and at equal probabilities of the t distribution — and the band estimator was run on a hundred simulations of four thousand samples at each number of bands.
The gain rises and then falls. Across the lower stretch, two bands were worth 1.77 times the draws, eight bands 2.69, sixteen 2.95 and thirty-two 2.99. Beyond that the bands start costing draws: sixty-four were worth 2.80 and two hundred and fifty-six 2.14, at which point sixty-nine of the bands in a typical run held no draws at all. Equal-probability bands were worth 1.09 at two and 1.73 at eight, and reached their own best, 3.03, also at thirty-two.
Thirty-two bands are worth more than the single matched cut, 2.32, and more than twice what the companion used as a straight line was worth, 1.43. The step from one well-placed cut to a few dozen is where most of the remaining gain lies. In the units a table of rates is read in, that is the difference between a cell whose own simulation error flags it as wrong on a good share of honest runs and one run three times as long, for the price of sorting each draw into a band.
The two placements finish level on this simulation, where they did not in the exact arithmetic, and the reason is the stretch chosen. The lower band from −4 to 0 covers half the companion’s distribution, so its cuts are not much more concentrated than equal-probability ones, and the lognormal test’s conditional probability rises over a wider range than a bivariate normal’s at 0.9 would. A band fitted more tightly to where this test rejects would do better at a few strata; the measurement uses a plain stretch, fixed in advance, so that nothing about the bands was tuned on the runs that scored them.
Post-stratified is not stratified
The band estimator decides the bands after the draws are made. A stratified simulation would decide them before: draw a fixed number of samples inside each band of the companion, in proportion to its probability or more heavily where the target’s rate is uncertain, and never leave a band empty or under-filled. That is the simulation’s version of allocating units where the variance is, and when it can be done it is better than anything post-stratification can reach, because it removes the randomness in how many draws each band receives as well as the imbalance in where they fell.
It usually cannot be done here, and the reason is specific. To draw a lognormal sample whose logarithms’ t statistic falls in a given band, the simulation would have to generate normal samples conditional on their t statistic, which is possible for a normal sample — the statistic and the direction of the sample are independent — but awkward, and impossible for most companions of interest. Post-stratification asks for nothing of the kind. It takes the draws as they come, which is why it is the one that fits an existing simulation, and it pays for that convenience with the variance of the band counts, which is the second-order term that grows as the bands multiply.
The same distinction runs through experimental design. A trial that stratifies its randomisation fixes the balance in advance; a trial that adjusts for a covariate afterwards corrects the balance it got. A threshold in the tail measured how much of a threshold’s imbalance a balanced covariate removes, and found the same φ©²/p(1 − p) that caps a straight companion here; the band estimator is the after-the-fact correction, applied to a simulation’s draws rather than to a trial’s units.
Why too many bands give the gain back
Two things go wrong when the bands outnumber what the draws can fill, and only one of them is in the first-order formula.
The first is noise. A band’s rate estimated from five draws is a poor estimate, and the band estimator adds times that poor estimate into the total. To first order this does not matter — the formula above assumes each band’s rate is as good as its share of the draws allows — but at second order a band’s contribution to the variance grows like once the expected count in it is small, and with hundreds of bands those terms add up to more than the bands remove.
The second is an empty band. It must be given some rate, and any rule for doing so is a model. The rule used here borrows the pooled rate of the two neighbouring bands, which is sensible where the conditional probability is smooth and wrong at the edges of the rise, where the neighbours on either side differ. At two hundred and fifty-six bands that rule is used sixty-nine times a run and the estimate drifts down, to 5.94% against 6.13% for the plain rate — a bias that was not there at thirty-two bands, bought by giving the estimator more flexibility than its draws could inform.
That is the flexibility question the check worth more than the check warned of when it named a non-linear companion as the first repair for a rate. Post-stratification avoids most of it, because nothing is fitted: the weights are exact and only the band rates are estimated. But the number of bands is a choice, and past the point where the bands are thinner than the draws can fill, it is a choice between noise and a model for empty bands.
How many bands a run can afford
The measurement gives a rule of thumb with a reason behind it. At four thousand draws the gain peaked at thirty-two bands, about a hundred and twenty-five draws a band on average — but the average hides where the draws go. Across the lower stretch, the outer bands near −4 receive very few draws, because the companion is rarely that low, and the bands where the rejections happen receive more. What limits the number of bands is the thinnest band that still matters, not the average.
So the practical placement is the one the exact arithmetic already recommended, with a floor: cut where the conditional probability is changing, and stop cutting before any band that carries a meaningful share of the target’s events expects fewer than a few dozen draws. On a simulation of a hundred thousand draws the same rule would allow several hundred bands and a gain close to the ceiling; at four thousand it stops near thirty.
What this has to do with a simulation that could be avoided
The field’s question was whether there is a useful middle between simulating a rate, where a smooth companion cannot see its step, and computing the rate exactly, where the simulation is unnecessary. Post-stratification on an exact companion is a precise version of that middle. It computes exactly the part of the rate that the companion determines — the bands’ probabilities — and simulates only the part it does not — the rate inside each band. The finer the bands, the more of the rate is computed and the less is simulated, until at the limit the only simulated quantity is how often the event happens at a given value of the companion.
The ceiling on what the middle can buy is therefore a statement about how much of the target event the companion explains, and for the lognormal test it is an honest three draws for one. Where the companion explains nearly all of it — a correlation near one — the middle approaches the exact calculation; where it explains little, the middle collapses back to plain simulation. Neither end requires a new idea, and the whole of the craft is in the placement of the cuts. It is also the opposite trade to draws aimed at the tail, which changes where the draws come from and weights them back; the bands leave the draws alone and correct what they happened to be. It is the same division of labour a coverage counted exactly makes when it sums over every possible count, applied to a companion whose “counts” are bands of a continuous variable.
What is exact here and what is measured
Post-stratifying a 5% rate on a companion correlated 0.9 removes 46.8% of its variance with eight strata placed across the stretch where the conditional probability changes, and 29.9% with eight at equal probabilities; the ceiling, reached by infinitely many strata, is 48.1%. These are exact for bivariate normals, from orthant probabilities.
One cut at the target’s threshold is the companion’s own rate, 38.2%, and the band estimator and that control variate agree to first order.
On the lognormal t test’s size, thirty-two bands of the logarithms’ t statistic were worth 2.99 times the draws, and two hundred and fifty-six were worth 2.14, with sixty-nine empty bands a run, measured over a hundred simulations of four thousand samples at each number of bands.
Not claimed: that thirty-two bands is right for other simulations. The peak depends on the number of draws, the rate’s rarity and where the companion puts its mass, and the rule stated above — no band that matters with fewer than a few dozen expected draws — is a reading of this one measurement rather than a derived bound. Not claimed either that the neighbour rule for empty bands is a good one; it is the simplest, and a better one would move the right-hand end of the curve without changing where it peaks.
Still open: a companion with more than one coordinate
Every companion here is one number per draw. A simulation often has several exact companions available at once — the logarithms’ t statistic, their sample variance’s chi-square, the sign of their median — and bands in several dimensions multiply quickly: eight cuts in each of three companions is five hundred and twelve cells, which four thousand draws cannot fill.
The multivariate version of the question is whether the companions can be combined into a single score first — the conditional probability of the target given all of them, estimated from a pilot — and the bands cut on that score, which keeps the number of cells small while using every companion. The score’s own distribution would then have to be known exactly for the weights to be exact, and it usually is not: a combination of exactly distributed statistics is not in general exactly distributed. Whether the weights can be computed for the combinations that matter, and what is lost by estimating them when they cannot, has not been measured.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The sample is a condition — both name closed form, conditional distribution, monte carlo, threshold
- A basis is a subspace — both name monte carlo, threshold, variance reduction
- A covariate with no levels — both name closed form, monte carlo, stratification
- A rate times a size — both name closed form, monte carlo, variance reduction
- An ordering that depends on the rule — both name bias-variance, closed form, monte carlo
- Balanced on the wrong function — both name stratification, threshold, variance reduction
Named objects
A flat tag is an object no other essay names yet.
Bias-varianceBinningClosed formConditional distributionMonte CarloStratificationThresholdVariance reduction