A rate needs a rate beside it
Worth reading first: What the 95% refers to · Two routes to every number.
The check worth more than the check found that a quantity computed exactly beside a simulation can sharpen it: subtract the part of the simulated estimate’s error that moves with the companion’s observable error, and what survives is a share of the variance. On one simulation a companion for an interval’s expected width was worth 214 times the draws. A companion for the same interval’s coverage was worth 1.08 times them, and the essay traced the failure to the shape of the thing being estimated: a coverage is an average of indicators, an indicator is a step, and a straight line fitted to a step sees little of it.
It ended on a problem, because almost every number a simulation study reports is a rate — a coverage, a size, a power, a false-discovery proportion — and the method it had just measured is nearly useless for rates. The repair it named first was to make the companion non-linear. The obvious non-linear function of a companion is its own step: the companion’s rate at the same threshold. How much that recovers turns out to have an exact answer, and the answer is a cut correlation — the quantity the arcsine that closes it evaluated.
The ceiling on a straight line
Take the cleanest case. A statistic is standard normal, the rate being simulated is , and a companion is standard normal too, correlated with , with its expectation known to be zero. A linear control variate removes the squared correlation between the indicator and , and that correlation has a closed form:
The factor that matters is the second one, and it does not depend on the companion at all. Even a companion perfectly correlated with the statistic — — removes only
of the rate’s variance. At the median that is 2/π, 63.7%. For a 10% rate it is 34.2%; for a 5% rate, 22.4%; for a 1% rate, 7.2%; for one in a thousand, 1.1%.
The ceiling is a familiar function. It is the share of a normal quantity’s information that survives cutting it at — the curve an outcome cut in two priced a responder analysis with, and a baseline cut in two found again for an adjustment, and a threshold in the tail found as the share of a threshold’s imbalance a balanced covariate removes. Here it arrives from the other side of the same identity. There, a measured quantity was replaced by its indicator and kept of what it knew. Here, an indicator is being predicted from the measured quantity, and the best straight line through it explains the same share. The squared correlation between a normal and its own cut is one number, whichever way round it is read.
That is why the coverage companion in the essay on control variates was worth so little. A coverage near 95% is a miss rate near 5%, and no straight-line companion of any kind can remove more than a fifth of a 5% rate’s variance. The count it used correlated 0.2665 with the covering indicator — 7.1% removed, against a ceiling of about 22% — and no better choice of straight-line companion was available to find.
A companion’s own rate
The straight line fails because it spends its fit on the whole range of the companion, most of which says nothing about the event. A draw where the companion sits three standard deviations above the threshold and a draw where it sits one above predict the same thing — the event happened — and the straight line insists on predicting different things for them. The companion’s own indicator at the same threshold, , has the step in the same place, and its expectation is known exactly whenever ’s distribution is.
Its correlation with the target indicator is the correlation between two cuts of two correlated normals, which is an orthant probability:
At the median it is Sheppard’s , the arcsine that closed the cut-point geometry. Elsewhere it is the same integral evaluated at a different corner, and it tends to one as does — a rate companion perfectly tracking the statistic removes everything, where a straight one removed at most a fifth.
For a 5% rate at a correlation of 0.9, the straight companion removes 18.1% and the companion’s rate 38.2%. At 0.99, the straight line is stuck at 21.9% against its ceiling of 22.4%, and the rate companion removes 77.0%, which is worth 4.35 times the draws.
Where the straight line still wins
The comparison does not go one way everywhere, and the exception is informative.
At the median the straight line’s ceiling is 63.7%, high enough that it stays competitive: at a correlation of 0.9 it removes 51.6% where the companion’s rate removes 50.8%, and the rate companion overtakes it only above a correlation of 0.91. At a 5% rate the crossover is at 0.64. The further into a tail the rate sits, the lower the straight line’s ceiling and the earlier the rate companion wins; at the median, where a step and a line look most alike, a line is nearly as good as a step until the companion is very close to the statistic.
Both fall short of the third curve. The best any function of the companion can do is its conditional expectation, , and for bivariate normals its explained share is again an orthant probability — the correlation of two cuts at the squared correlation, because two copies of that share only correlate . At 0.9 that ceiling is 48.1% for a 5% rate; at 0.99, 82.8%. The companion’s own rate reaches four fifths of the best at 0.9 and more than nine tenths at 0.99, without anything having been fitted.
A companion that exists: the logarithms of a lognormal sample
The bivariate-normal case is the arithmetic. A simulation the site actually needs is where it earns its keep, and one of the commonest is a test’s size on data the test was not built for.
The one-sample t test assumes normal data. On a sample of twenty lognormal observations — the logarithms normal with a standard deviation of 0.5 — testing the true mean, the lower tail rejects 6.14% of the time at a nominal 2.5%, because the sample mean and standard deviation of skewed data move together and a low mean comes with a small standard deviation. It is the defect where the two tails disagree measured on an exponential source, one tail too heavy while the total looks nearly right, and it is the number a limit read from one side depends on. That rate has no closed form, which is why it is simulated.
The same draws carry a companion whose rate is exact. The logarithms of each sample are a normal sample, so the t statistic of the logarithms, tested against their own true mean, is exactly Student’s t on nineteen degrees of freedom: its lower-tail rate at any cut is known to every digit. The two statistics are computed from the same twenty numbers and move together, most tightly where it matters.
The picture shows why the straight companion has little to work with. The cloud is nearly a straight band, and most of it sits far from either critical line; the rejections the rate counts are a thin sliver at the lower left, and a straight line fitted to the indicator over the whole band spends its slope on the middle, where nothing happens. The companion’s own rejection region, by contrast, is a sliver in the same corner.
What each companion was worth
Over two hundred simulations of four thousand samples each, each companion was used in turn and the spread of the estimates measured directly.
Used as a straight line, the logarithms’ t statistic was worth 1.43 times the draws; its measured correlation predicts 1.31, and the two agree within the error of two hundred runs. Used as its own rate at the test’s level, 2.5%, it was worth 1.60, predicted 1.55. Used as its own rate at 6.5% — a level chosen from a short pilot run to sit near the rate being estimated — it was worth 2.32, predicted 2.23.
The last step is the one worth drawing out. The companion’s rate at 2.5% counts a sliver of the lower corner and the target’s rate at 6.14% counts a wider one, so the two steps are in different places and correlate only 0.60. Moving the companion’s cut until its rate matches the target’s puts the steps together, and the correlation rises to 0.74. Nothing about the companion’s exactness is spent in the move: the cut is fixed before the run, so the companion’s expectation at that cut, 6.5%, is still known to every digit.
All three adjusted estimates are centred on the plain one to within a twentieth of a percentage point, inside the error of two hundred runs. The companions change the spread and not the centre, which is the check that nothing was bought with bias.
What the gain is worth in a table
A factor of 2.32 sounds modest beside the 214 an expected width earned, and for a single number it is. Its value shows in the place rates are actually reported, which is a table of them. A coverage table with its own error found that twenty correct cells at a thousand replications each flag at least one of themselves as wrong on most honest runs, and that ten times the replications does not cure it, because the standard error falls only as the square root. A companion worth 2.32 draws is the same as running every cell 2.32 times as long, at the cost of computing one more t statistic per sample — which the simulation was nearly computing anyway.
It is also the same construction as running two methods on the same draws, seen from a different angle. There, two procedures shared their samples so that the difference between their rates was estimated more precisely than either rate. Here one of the two procedures has a rate known exactly, so the shared draws sharpen the other one’s rate itself rather than a difference. Whenever a simulation compares a procedure with a textbook procedure whose behaviour is exact on the same data, the textbook procedure is a companion waiting to be used.
Rarer rates, and the other tool for them
The further into a tail the rate sits, the worse the straight companion does — its ceiling falls to 7.2% at 1% and 1.1% at one in a thousand — and the more a matched rate companion has to offer, because a step in the right place still sees the event however rare it is. But a rare event also means few draws land in the region where either step is, and a companion cannot correct an estimate built from a handful of events: it can only remove the part of the error that moves with its own.
For genuinely rare rates the tool is different. Draws aimed at the tail change where the samples come from and weight them back, and they turn a probability of three in ten million from a job of hundreds of millions of draws into one of hundreds. A companion leaves the draws where they were. The two compose — an aimed simulation can still carry a companion computed on its own draws — and they answer different shortages: aiming fixes having too few events, a companion fixes having too few draws for the events there are. A size near 5% is the second kind of problem, which is why a companion is the right first repair for the rates simulation studies count most.
Why matching the level is the right move
A pilot run choosing the companion’s cut is a small selection, and the essay on control variates warned against choosing a companion by how well it correlates on the run in hand, because the coefficient is then fitted to the noise that chose it. The warning does not bite here, for two reasons. The pilot is a separate run, so the main run’s coefficient is estimated on draws the choice never saw. And the choice is of a level rather than of a companion: the target’s rate is the only thing the pilot is used to read, to the nearest half a percentage point, and the companion’s exactness does not depend on the level being right.
What the level buys follows from the orthant arithmetic. Two cuts correlate most when they cut the same share off correlated variables — a cut at 2.5% and a cut at 6% of the same normal cannot correlate perfectly even at , because a draw can fall between them. So a companion’s rate should be cut where the target’s rate is, and when the target’s rate is what is unknown, a rough estimate of it is enough to place the cut.
What changes in practice
For a rate, use the companion as a rate. A simulation estimating a size, a coverage or a power, beside a statistic whose distribution is known, should use the known statistic’s own exceedance indicator as its control, not the known statistic itself. The ceiling on the straight line is not a matter of choosing the companion well; it is , and at the tails where tests live it is small.
Cut the companion where the target is. A companion’s rate at the nominal level is the obvious choice and is not the best one when the target’s rate is far from nominal — which is exactly when the simulation is interesting. A short pilot fixes the level.
Look for a transformation with an exact law. The logarithms of a lognormal sample are one case of a general pattern: data generated by transforming normal variables carry, inside every simulated sample, the normal variables themselves, and any statistic computed on those has an exact distribution. A simulation of a procedure’s behaviour on non-normal data can almost always compute the same procedure on the normal variables it was built from, for free, and use that as its companion.
And report what the companion was worth. The correlation of the two indicators, measured on the run, is the one number that says whether the adjustment was worth doing, and at 0.60 against 0.74 the difference between a companion worth 1.6 draws and one worth 2.3 is visible in it before any spread is measured.
What is exact here and what is measured
A companion used as a straight line removes at most φ©²/p(1 − p) of a simulated rate’s variance — 63.7% at the median, 22.4% for a 5% rate, 7.2% for a 1% rate — exactly, for a normal statistic, whatever the companion.
The companion’s own rate at the same cut removes the squared cut correlation, an orthant probability: 38.2% for a 5% rate at a correlation of 0.9, 77.0% at 0.99. The best function of the companion removes the cut covariance at the squared correlation, 48.1% and 82.8%.
On the lognormal t test’s size at σ = 0.5 and samples of twenty, the companion used linearly was worth 1.43 times the draws, as a rate at 2.5%, 1.60, and as a rate at a matched 6.5%, 2.32, measured over two hundred runs of four thousand samples and agreeing with what each run’s own correlation predicted.
Not claimed: that the lognormal’s gains carry to other skewed families, where the transformation back to normality may not exist or may not be known. Not claimed either that the bivariate-normal ceilings bound the lognormal case: the two t statistics are not jointly normal, their relation is tighter in the lower tail than in the middle, and the measured worth of each companion is its own.
Still open: more than one cut of the companion
The companion’s own rate is one step placed where the target’s step is. The best function of the companion is the whole curve , a smooth rise from nothing to one, and the gap between the two — 38.2% against 48.1% at a correlation of 0.9 — is what a single step leaves on the table. The natural way to follow the curve without fitting it is to cut the companion in several places and estimate the target’s rate separately within each band, weighting the bands by their exact probabilities.
That is post-stratification on the companion, and it raises two questions a single cut did not. Where should the cuts go — at equal probabilities of the companion, as a histogram would place them, or where the conditional probability is changing? And how many can a simulation of a given size afford before the bands are too thin to estimate a rate in? Both have answers in the same orthant arithmetic, and neither has been worked out here.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A copula that halves a marginal — both name closed form, correlation, median split, skewness
- The fourth moment that was missing — both name closed form, correlation, monte carlo, threshold
- The sample is a condition — both name closed form, correlation, monte carlo, threshold
- The zero that survives both — both name closed form, median split, skewness, threshold
- A basis is a subspace — both name monte carlo, threshold, variance reduction
- A boundary for giving up — both name closed form, error rate, monte carlo
Named objects
A flat tag is an object no other essay names yet.
Closed formCorrelationError rateMedian splitMonte CarloSkewnessThresholdVariance reduction