A cut point, at a correlation

A rate needs a rate beside it

A simulated rejection rate is an average of indicators, and a companion with an exactly known expectation can remove at most φ(c)²/p(1 − p) of its variance when it is used as a straight line — 22.4% for a 5% rate, 7.2% for a 1% rate — however closely it tracks the statistic. The ceiling is the share of a normal's information a cut keeps. Used as its own rate at the same cut, the same companion removes 38.2% at a correlation of 0.9 and 77.0% at 0.99. On the t test's size for lognormal data it was worth 1.43 times the draws used linearly and 2.32 times used as a rate at a matched level.

Worth reading first: What the 95% refers to · Two routes to every number.

The check worth more than the check found that a quantity computed exactly beside a simulation can sharpen it: subtract the part of the simulated estimate’s error that moves with the companion’s observable error, and what survives is a share 1−ρ21 - \rho^2 of the variance. On one simulation a companion for an interval’s expected width was worth 214 times the draws. A companion for the same interval’s coverage was worth 1.08 times them, and the essay traced the failure to the shape of the thing being estimated: a coverage is an average of indicators, an indicator is a step, and a straight line fitted to a step sees little of it.

It ended on a problem, because almost every number a simulation study reports is a rate — a coverage, a size, a power, a false-discovery proportion — and the method it had just measured is nearly useless for rates. The repair it named first was to make the companion non-linear. The obvious non-linear function of a companion is its own step: the companion’s rate at the same threshold. How much that recovers turns out to have an exact answer, and the answer is a cut correlation — the quantity the arcsine that closes it evaluated.

What a companion with a known expectation can remove from the variance of a simulated 5% rateFor a rate of 0.05, a companion used linearly removes at most 22.4% of the variance however well it correlates. At a correlation of 0.9 it removes 18.1%, the companion's own rate at the same cut 38.2%, and the best function of the companion 48.1%; at 0.99, 21.9%, 77.0% and 82.8%. The companion's rate overtakes the linear companion at a correlation of 0.64.00.2500.5000.750100.2000.4000.6000.8001correlation between the statistic and its companionshare of the simulated rate's variance removedthe linear ceiling, φ(c)²/p(1 − p) = 22.4%the best function of the companionthe companion's own ratethe companion, linearlybivariate normal, orthant probabilities, exacta rate is best predicted by a rate
Fig. 1 The share of a simulated 5% rate’s variance a companion can remove, against the correlation between the statistic and its companion: the companion used as a straight line, the companion’s own rate at the same cut, and the best function of the companion there is. The dashed line is the straight line’s ceiling. The slider sets how rare the event is.

The ceiling on a straight line

Take the cleanest case. A statistic TT is standard normal, the rate being simulated is p=P(T>c)p = P(T > c), and a companion SS is standard normal too, correlated ρ\rho with TT, with its expectation known to be zero. A linear control variate removes the squared correlation between the indicator 1{T>c}1\{T > c\} and SS, and that correlation has a closed form:

Corr(1{T>c},S)=ρ φ(c)p(1−p).\mathrm{Corr}(1\{T > c\}, S) = \frac{\rho\,\varphi(c)}{\sqrt{p(1-p)}}.

The factor that matters is the second one, and it does not depend on the companion at all. Even a companion perfectly correlated with the statistic — S=TS = T — removes only

φ(c)2p(1−p)\frac{\varphi(c)^2}{p(1-p)}

of the rate’s variance. At the median that is 2/π, 63.7%. For a 10% rate it is 34.2%; for a 5% rate, 22.4%; for a 1% rate, 7.2%; for one in a thousand, 1.1%.

The most a companion used linearly can remove from a simulated rate, against how rare the event is. The ceiling is φ(c)²/(p(1 − p)) at the cut c whose tail is p: 63.7% for a median rate, 34.2% at 10%, 22.4% at 5%, 7.2% at 1% and 1.1% at 0.1%. It is reached only by a companion perfectly correlated with the statistic.
Fig. 2 The most a companion used as a straight line can remove from a simulated rate, against how rare the event is. Even knowing the statistic itself exactly, used linearly, buys a fifth of the variance at a 5% rate.

The ceiling is a familiar function. It is the share of a normal quantity’s information that survives cutting it at cc — the curve an outcome cut in two priced a responder analysis with, and a baseline cut in two found again for an adjustment, and a threshold in the tail found as the share of a threshold’s imbalance a balanced covariate removes. Here it arrives from the other side of the same identity. There, a measured quantity was replaced by its indicator and kept φ(c)2/p(1−p)\varphi(c)^2/p(1-p) of what it knew. Here, an indicator is being predicted from the measured quantity, and the best straight line through it explains the same share. The squared correlation between a normal and its own cut is one number, whichever way round it is read.

That is why the coverage companion in the essay on control variates was worth so little. A coverage near 95% is a miss rate near 5%, and no straight-line companion of any kind can remove more than a fifth of a 5% rate’s variance. The count it used correlated 0.2665 with the covering indicator — 7.1% removed, against a ceiling of about 22% — and no better choice of straight-line companion was available to find.

A companion’s own rate

The straight line fails because it spends its fit on the whole range of the companion, most of which says nothing about the event. A draw where the companion sits three standard deviations above the threshold and a draw where it sits one above predict the same thing — the event happened — and the straight line insists on predicting different things for them. The companion’s own indicator at the same threshold, 1{S>c}1\{S > c\}, has the step in the same place, and its expectation is known exactly whenever SS’s distribution is.

Its correlation with the target indicator is the correlation between two cuts of two correlated normals, which is an orthant probability:

Corr(1{T>c},1{S>c})=P(T>c,S>c)−p2p(1−p).\mathrm{Corr}(1\{T > c\}, 1\{S > c\}) = \frac{P(T > c, S > c) - p^2}{p(1-p)}.

At the median it is Sheppard’s (2/π)arcsin⁡ρ(2/\pi)\arcsin\rho, the arcsine that closed the cut-point geometry. Elsewhere it is the same integral evaluated at a different corner, and it tends to one as ρ\rho does — a rate companion perfectly tracking the statistic removes everything, where a straight one removed at most a fifth.

For a 5% rate at a correlation of 0.9, the straight companion removes 18.1% and the companion’s rate 38.2%. At 0.99, the straight line is stuck at 21.9% against its ceiling of 22.4%, and the rate companion removes 77.0%, which is worth 4.35 times the draws.

Where the straight line still wins

The comparison does not go one way everywhere, and the exception is informative.

What a companion with a known expectation can remove from the variance of a simulated median rate. For a rate of 0.5, a companion used linearly removes at most 63.7% of the variance however well it correlates. At a correlation of 0.9 it removes 51.6%, the companion's own rate at the same cut 50.8%, and the best function of the companion 60.1%; at 0.99, 62.4%, 82.8% and 87.3%. The companion's rate overtakes the linear companion at a correlation of 0.91.
Fig. 3 The same three companions for a median rate. The straight line’s ceiling is 2/π, high enough that it keeps pace with the companion’s own rate until the correlation passes 0.9.

At the median the straight line’s ceiling is 63.7%, high enough that it stays competitive: at a correlation of 0.9 it removes 51.6% where the companion’s rate removes 50.8%, and the rate companion overtakes it only above a correlation of 0.91. At a 5% rate the crossover is at 0.64. The further into a tail the rate sits, the lower the straight line’s ceiling and the earlier the rate companion wins; at the median, where a step and a line look most alike, a line is nearly as good as a step until the companion is very close to the statistic.

Both fall short of the third curve. The best any function of the companion can do is its conditional expectation, P(T>c∣S)P(T > c \mid S), and for bivariate normals its explained share is again an orthant probability — the correlation of two cuts at the squared correlation, because two copies of TT that share only SS correlate ρ2\rho^2. At 0.9 that ceiling is 48.1% for a 5% rate; at 0.99, 82.8%. The companion’s own rate reaches four fifths of the best at 0.9 and more than nine tenths at 0.99, without anything having been fitted.

A companion that exists: the logarithms of a lognormal sample

The bivariate-normal case is the arithmetic. A simulation the site actually needs is where it earns its keep, and one of the commonest is a test’s size on data the test was not built for.

The one-sample t test assumes normal data. On a sample of twenty lognormal observations — the logarithms normal with a standard deviation of 0.5 — testing the true mean, the lower tail rejects 6.14% of the time at a nominal 2.5%, because the sample mean and standard deviation of skewed data move together and a low mean comes with a small standard deviation. It is the defect where the two tails disagree measured on an exponential source, one tail too heavy while the total looks nearly right, and it is the number a limit read from one side depends on. That rate has no closed form, which is why it is simulated.

The same draws carry a companion whose rate is exact. The logarithms of each sample are a normal sample, so the t statistic of the logarithms, tested against their own true mean, is exactly Student’s t on nineteen degrees of freedom: its lower-tail rate at any cut is known to every digit. The two statistics are computed from the same twenty numbers and move together, most tightly where it matters.

Two t statistics computed on the same 1,500 samples of 20 lognormal observations. The t statistic of the data, which rejects its true mean more often than 2.5% in the lower tail, against the t statistic of the logarithms, which is exactly Student's t. Of 1,500 samples, 27 are rejected by both, 57 by the data's test alone and 1 by the logarithms' alone.
Fig. 4 The t statistic of fifteen hundred lognormal samples against the t statistic of their logarithms. The vertical line is the data’s lower critical value and the horizontal line the logarithms’. The data’s test rejects well beyond its nominal share and most of its rejections fall where the logarithms’ statistic is also low.

The picture shows why the straight companion has little to work with. The cloud is nearly a straight band, and most of it sits far from either critical line; the rejections the rate counts are a thin sliver at the lower left, and a straight line fitted to the indicator over the whole band spends its slope on the middle, where nothing happens. The companion’s own rejection region, by contrast, is a sliver in the same corner.

What each companion was worth

Over two hundred simulations of four thousand samples each, each companion was used in turn and the spread of the estimates measured directly.

What each companion was worth on the t test's size for lognormal data, σ = 0.5, samples of twenty. The rejection rate being simulated is 6.14%. Over 200 simulations of 4,000 samples, the logarithms' t statistic used linearly was worth 1.43 times the draws (its correlation predicts 1.31); the logarithms' own rejection at 2.5%, 1.60 (predicted 1.55); at 6.5%, a level matched to the target's from a pilot run, 2.32 (predicted 2.23).
Fig. 5 What each companion was worth on the lognormal t test’s lower-tail size, in multiples of the draws: the logarithms’ t statistic used as a straight line, the logarithms’ own rejection at the test’s level of 2.5%, and their rejection at a level matched to the target’s rate.

Used as a straight line, the logarithms’ t statistic was worth 1.43 times the draws; its measured correlation predicts 1.31, and the two agree within the error of two hundred runs. Used as its own rate at the test’s level, 2.5%, it was worth 1.60, predicted 1.55. Used as its own rate at 6.5% — a level chosen from a short pilot run to sit near the rate being estimated — it was worth 2.32, predicted 2.23.

The last step is the one worth drawing out. The companion’s rate at 2.5% counts a sliver of the lower corner and the target’s rate at 6.14% counts a wider one, so the two steps are in different places and correlate only 0.60. Moving the companion’s cut until its rate matches the target’s puts the steps together, and the correlation rises to 0.74. Nothing about the companion’s exactness is spent in the move: the cut is fixed before the run, so the companion’s expectation at that cut, 6.5%, is still known to every digit.

All three adjusted estimates are centred on the plain one to within a twentieth of a percentage point, inside the error of two hundred runs. The companions change the spread and not the centre, which is the check that nothing was bought with bias.

What the gain is worth in a table

A factor of 2.32 sounds modest beside the 214 an expected width earned, and for a single number it is. Its value shows in the place rates are actually reported, which is a table of them. A coverage table with its own error found that twenty correct cells at a thousand replications each flag at least one of themselves as wrong on most honest runs, and that ten times the replications does not cure it, because the standard error falls only as the square root. A companion worth 2.32 draws is the same as running every cell 2.32 times as long, at the cost of computing one more t statistic per sample — which the simulation was nearly computing anyway.

It is also the same construction as running two methods on the same draws, seen from a different angle. There, two procedures shared their samples so that the difference between their rates was estimated more precisely than either rate. Here one of the two procedures has a rate known exactly, so the shared draws sharpen the other one’s rate itself rather than a difference. Whenever a simulation compares a procedure with a textbook procedure whose behaviour is exact on the same data, the textbook procedure is a companion waiting to be used.

Rarer rates, and the other tool for them

The further into a tail the rate sits, the worse the straight companion does — its ceiling falls to 7.2% at 1% and 1.1% at one in a thousand — and the more a matched rate companion has to offer, because a step in the right place still sees the event however rare it is. But a rare event also means few draws land in the region where either step is, and a companion cannot correct an estimate built from a handful of events: it can only remove the part of the error that moves with its own.

For genuinely rare rates the tool is different. Draws aimed at the tail change where the samples come from and weight them back, and they turn a probability of three in ten million from a job of hundreds of millions of draws into one of hundreds. A companion leaves the draws where they were. The two compose — an aimed simulation can still carry a companion computed on its own draws — and they answer different shortages: aiming fixes having too few events, a companion fixes having too few draws for the events there are. A size near 5% is the second kind of problem, which is why a companion is the right first repair for the rates simulation studies count most.

Why matching the level is the right move

A pilot run choosing the companion’s cut is a small selection, and the essay on control variates warned against choosing a companion by how well it correlates on the run in hand, because the coefficient is then fitted to the noise that chose it. The warning does not bite here, for two reasons. The pilot is a separate run, so the main run’s coefficient is estimated on draws the choice never saw. And the choice is of a level rather than of a companion: the target’s rate is the only thing the pilot is used to read, to the nearest half a percentage point, and the companion’s exactness does not depend on the level being right.

What the level buys follows from the orthant arithmetic. Two cuts correlate most when they cut the same share off correlated variables — a cut at 2.5% and a cut at 6% of the same normal cannot correlate perfectly even at ρ=1\rho = 1, because a draw can fall between them. So a companion’s rate should be cut where the target’s rate is, and when the target’s rate is what is unknown, a rough estimate of it is enough to place the cut.

What changes in practice

For a rate, use the companion as a rate. A simulation estimating a size, a coverage or a power, beside a statistic whose distribution is known, should use the known statistic’s own exceedance indicator as its control, not the known statistic itself. The ceiling on the straight line is not a matter of choosing the companion well; it is φ(c)2/p(1−p)\varphi(c)^2/p(1-p), and at the tails where tests live it is small.

Cut the companion where the target is. A companion’s rate at the nominal level is the obvious choice and is not the best one when the target’s rate is far from nominal — which is exactly when the simulation is interesting. A short pilot fixes the level.

Look for a transformation with an exact law. The logarithms of a lognormal sample are one case of a general pattern: data generated by transforming normal variables carry, inside every simulated sample, the normal variables themselves, and any statistic computed on those has an exact distribution. A simulation of a procedure’s behaviour on non-normal data can almost always compute the same procedure on the normal variables it was built from, for free, and use that as its companion.

And report what the companion was worth. The correlation of the two indicators, measured on the run, is the one number that says whether the adjustment was worth doing, and at 0.60 against 0.74 the difference between a companion worth 1.6 draws and one worth 2.3 is visible in it before any spread is measured.

What is exact here and what is measured

A companion used as a straight line removes at most φ©²/p(1 − p) of a simulated rate’s variance — 63.7% at the median, 22.4% for a 5% rate, 7.2% for a 1% rate — exactly, for a normal statistic, whatever the companion.

The companion’s own rate at the same cut removes the squared cut correlation, an orthant probability: 38.2% for a 5% rate at a correlation of 0.9, 77.0% at 0.99. The best function of the companion removes the cut covariance at the squared correlation, 48.1% and 82.8%.

On the lognormal t test’s size at σ = 0.5 and samples of twenty, the companion used linearly was worth 1.43 times the draws, as a rate at 2.5%, 1.60, and as a rate at a matched 6.5%, 2.32, measured over two hundred runs of four thousand samples and agreeing with what each run’s own correlation predicted.

Not claimed: that the lognormal’s gains carry to other skewed families, where the transformation back to normality may not exist or may not be known. Not claimed either that the bivariate-normal ceilings bound the lognormal case: the two t statistics are not jointly normal, their relation is tighter in the lower tail than in the middle, and the measured worth of each companion is its own.

Still open: more than one cut of the companion

The companion’s own rate is one step placed where the target’s step is. The best function of the companion is the whole curve P(T>c∣S)P(T > c \mid S), a smooth rise from nothing to one, and the gap between the two — 38.2% against 48.1% at a correlation of 0.9 — is what a single step leaves on the table. The natural way to follow the curve without fitting it is to cut the companion in several places and estimate the target’s rate separately within each band, weighting the bands by their exact probabilities.

That is post-stratification on the companion, and it raises two questions a single cut did not. Where should the cuts go — at equal probabilities of the companion, as a histogram would place them, or where the conditional probability is changing? And how many can a simulation of a given size afford before the bands are too thin to estimate a rate in? Both have answers in the same orthant arithmetic, and neither has been worked out here.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Closed formCorrelationError rateMedian splitMonte CarloSkewnessThresholdVariance reduction