A companion with two coordinates
Worth reading first: What the 95% refers to · Two routes to every number.
A companion cut into many post-stratified a simulated rate on one companion statistic whose distribution is known exactly. The rate was the actual size of a t test on samples of twenty lognormal observations, 6.2% at a nominal lower-tail 2.5%; the companion was the t statistic of the same samples’ logarithms, which is exactly Student’s t on nineteen degrees of freedom because the logarithms are a normal sample; and a few dozen bands of it were worth about three times the draws. The essay ended on the harder case every simulation actually has. A companion is rarely one number. The logarithms offer their mean and their spread separately, each with its own exact law, and the t statistic of the logarithms is only one function of the two.
Cutting two companions into bands at once multiplies the cells, and the natural repair — combine them into a single score first and cut the score — produces a companion whose distribution is not a named law. The question left open was whether its band probabilities can be computed at all, and what is lost by estimating them when they cannot. On this simulation they can be computed, the combination is worth more than half as much again as the t statistic alone, and estimating the weights instead throws all of it away.
Two statistics the logarithms already have
Write a sample’s logarithms in units of their own standard deviation, , so that are a standard normal sample. They have two summaries that matter here, and both have exact laws that do not depend on anything unknown:
and they are independent — the defining property of a normal sample, and the one Student’s derivation rests on. The logarithms’ t statistic is , a function of the pair. So the pair is worth at least what the t statistic is worth, and it can be worth more if the target — whether the t test on the lognormal data rejects — depends on the mean and spread of the logarithms in some way the ratio does not capture.
It does, and the reason is visible in a picture of where the rejections are.
The rejections occupy a slanted wedge: low logarithmic mean, and — for any given mean — a smaller spread makes rejection more likely. The t statistic of the logarithms is constant along rays through the origin, and the wedge’s edge is not a ray. Along the ray on which it is , the share of samples rejected runs from 43.3% where the spread is large to 100.0% where it is small. A companion that knows only which ray a sample is on averages over that whole range, and every band of it is a mixture of samples that were certain to reject and samples that were close to even.
The mechanism is the one where the two tails disagree found for t intervals on skewed data, seen from underneath. A lognormal sample whose data t statistic falls far below its critical value is one that has missed the large observations, and on the logarithmic scale that is a sample whose logarithms are low and bunched: missing the large values pulls the mean down and removes the spread they would have supplied. The data’s t statistic divides its mean deficit by its own small spread, so the spread of the logarithms carries information about the rejection that their ratio to the mean discards.
A grid in both coordinates
The obvious use of two companions is a grid. Cut the mean into bands across its lower stretch, cut the spread at equal probabilities of its chi-square, and post-stratify on the cells, each of whose probability is a product of a normal probability and a chi-square one — exact, because the two are independent.
To first order the grid does what the picture promises. Eight by eight cells are worth 3.49 times the draws, sixteen by sixteen 4.31, and thirty-two by thirty-two 4.68, against the 2.77 that bands of the t statistic reach however finely they are cut. But it needs its cells: a two-by-two grid is worth 1.33 and four-by-four 1.86, less than the same number of bands of the t statistic, because a coarse grid in two coordinates wastes most of its cells in the half of the plane where nothing rejects.
That is the arithmetic that sank equal-probability strata in the one-coordinate case, squared. A grid has to be fine in both directions to follow a boundary that is diagonal, so its cells grow as the square of its resolution, and a simulation has to fill them. On runs of four thousand samples a sixteen-by-sixteen grid counts at 3.07 times the draws, against its 4.31 to first order; the thirty-two-by-thirty-two grid, 1,024 cells, counts at 1.97, with 94 of its cells left empty in a typical run.
Worth here is counted as mean squared error about the rate over the whole million samples, not as variance, and the difference matters for the grid. An empty cell has to be given some rate, and the rule used here is the gentlest available — the pooled rate of the cell’s own row of the mean — but the empty cells are disproportionately the rare cells at the edge of the wedge, where nearly every sample rejects, and any borrowed rate is too low there. On runs of a thousand samples the thousand-cell grid is biased low by 0.31 percentage points against a rate of 6.2% and is worth 1.30 — barely better than not stratifying at all. The hero figure’s slider shows it.
One number fitted to both
The alternative the earlier essay named is to collapse the pair into a single score before cutting — an estimate of the probability that a sample rejects, given its two coordinates — and cut the score into bands. Bands of one number cost cells linearly, so sixteen of them are sixteen cells however many coordinates went into the score.
The score used here is the simplest that can follow a slanted edge: a logistic regression of the rejection on , and their product, fitted on a pilot of a thousand samples drawn separately from the run and never counted in the rate. Its bands are cut at equal spacings of the fitted log-odds from to . Sixteen of them are worth 4.64 times the draws to first order and 4.50 counted on runs of four thousand; thirty-two bands, 4.68 and 4.25; sixty-four, 4.69 and 3.95. The first-order worth barely moves after sixteen, which says the score’s bands have already followed the conditional probability as far as the score can, and the counted worth falls slowly as the bands grow thinner than a run can fill evenly. Sixteen bands of the score are worth more than a thousand cells of a grid and do it with no empty cells at all.
What four and a half times the draws is for
A simulated rate is almost never wanted alone. It is a cell in a table — the size of a test at several sample sizes and skewnesses, the coverage of several intervals — and a coverage table with its own error showed how easily such a table misleads when each cell carries its sampling error unannounced. On runs of four thousand samples the plain estimate of this rate has a standard error of 0.382 percentage points, so a cell reading 6.2% could honestly read anything from about 5.4% to 7.0%. Sixteen score bands shrink that to 0.178 points, which is what the plain rate would give on 18,500 samples.
The check worth more than the check found one exact companion worth 214 times the draws on a smooth quantity and another worth 1.08 on a rate, and the difference was the difference between a quantity that moves smoothly with its companion and one that jumps; on this rate a straight line was worth about 1.4. The arithmetic of these essays has been a steady reclassification of what a companion should be: first a line, then a step, then several steps, and now several steps of a function of several companions — each change following the target’s conditional probability more closely, because that probability, and not the target’s correlation with anything, is what a companion can remove.
It also composes with the other economies a simulation study makes. Post-stratification changes only how each run’s draws are weighted, so a comparison of two tests on the same draws can stratify both on the same score and keep the correlation that common draws buy; and the score can be fitted once, on one pilot, for every cell of a table whose cells share a sample size, since the logarithms’ law does not depend on which test is being scored.
And the reason the second coordinate helps is the one the correction a t test would use found from the other end. There, the samples that made a t test’s long tail fail were the ones whose own skewness gave the least warning, because they had missed the large values. Here, the same samples are the ones whose logarithms are bunched as well as low; a statistic that reads the bunching finds them, and a statistic that divides it out does not. What makes those samples hard to correct for makes them easy to stratify on, provided the companion looks at their spread directly.
A band of a score has an exact probability
The part the earlier essay could not settle is the weights. Post-stratification replaces the share of the run’s samples that fell in a band by the band’s true probability, and the whole gain comes from that replacement. For bands of the t statistic the probability is a difference of two Student distribution functions. For bands of a fitted score it is the probability that a nonlinear combination of a normal and a chi-square falls between two values, and no table holds it.
It does not need one. For any fixed value of the spread the score is a straight line in — the fitted log-odds is , which in has slope — so the set of that puts a sample in a band is an interval, and its probability is a difference of two normal distribution functions. Integrating that over the chi-square law of gives the band’s probability, and the integral is one-dimensional and smooth. It is computed here over four thousand equal-probability points of the chi-square; the sixteen band probabilities sum to one to within , and each agrees with the share of the million samples that fell in it to within 2.6 of that share’s standard errors.
So the combination is not exactly distributed in the sense of having a name, but it is exactly computable, because the pair it is built from has a known joint law and the score is simple enough in one coordinate to be integrated over the other. That is the general condition, and it is worth stating plainly: a combination of exact companions has exact band probabilities whenever the bands can be written as regions whose probability under the companions’ joint law can be integrated. A score that is monotone in one coordinate given the rest always qualifies. The combination’s distribution does not need to be known; the probability of each band does.
How large the pilot has to be
The score is fitted, and the fit could be the weak point: a pilot too small to locate the wedge would cut bands in the wrong places.
It is not a weak point. A score fitted on a hundred samples — four of which were rejections — gives sixteen bands worth 3.88 times the draws, already well above anything the t statistic can do. A pilot of 250 gives 4.48, a thousand 4.64, and four thousand 4.85, after which a larger pilot changes nothing. The pilot need only locate the edge of the wedge roughly; the bands then measure the rate on each side of it from the run’s own draws. A badly fitted score costs a little of the gain and never introduces a bias, because the band probabilities are exact for whatever score was fitted — a poor score is a poor choice of strata, not a wrong weighting of them.
That separation is the design’s central property, and it is the same one the draws aimed at the tail relied on for importance sampling: the pilot chooses where to look, and correctness comes from the exact probabilities of where it looked, so an imperfect pilot degrades efficiency rather than truth.
Weights that have to be estimated
Suppose the band probabilities could not be computed. The obvious substitute is to estimate them — from the run’s own samples, or from a pilot.
From the run’s own samples the substitution is exact and useless. If each band’s weight is the share of the run that fell in it, the post-stratified rate is — the plain rate, to the last digit, whatever the bands are. The gain from post-stratifying is precisely the difference between a band’s true probability and its share in the run; estimate the first by the second and there is no difference left.
From an independent pilot of samples the estimate is not the plain rate, but it carries the pilot’s error in every weight.
To first order the estimator’s variance is the within-band variance the exact weights leave, divided by the run’s size, plus the variance of the band rates across bands divided by the pilot’s size. The second term is the part the exact weights remove, and it comes back scaled by the ratio of run to pilot. At a pilot the size of the run the two effects cancel exactly and the bands are worth 1.00 times the draws — the plain rate’s precision, bought with a second simulation. A pilot of a thousand, a quarter of the run, makes the bands worth 0.30: worse than not stratifying at all. It takes a pilot of 256,000 samples, sixty-four times the run, to recover 4.39 of the 4.64 that exact weights give for nothing.
And at every pilot size the same samples would have been worth more counted. A pilot of 16,000 added to a run of four thousand is worth five times the run; used for weights it is worth 2.43. A weight estimated from samples is paid for twice — once in the samples, and again in the error it carries into every band — and the samples would have been better spent on the rate. Post-stratification is worth doing only with weights that are computed, and the whole content of these companions is the class of statistics whose weights can be.
What the second coordinate bought, and what it needed
The pair is worth 4.64 times the draws against 2.77 for the one function of it that was being used, on a simulation where the logarithms’ t statistic had already seemed a natural and nearly complete companion. The gain came from a coordinate that looks irrelevant — the logarithms’ spread — and that carries information because a skewed sample’s failure has a shape on the logarithmic scale, low and bunched, that a ratio of mean to spread folds together.
A grid is the wrong way to use it. A thousand exact cells are worth 4.68 to first order and 1.97 on a run of four thousand, with a bias that reaches 0.31 points on a run of a thousand; the loss is not a matter of precision alone, since cells that are rare are cells at the edge of the rejection region, and every rule for an empty one borrows a rate from where the region is not.
A score is the right way, and its fit is cheap and its weights exact. A pilot of a hundred samples already beats the t statistic; a thousand gives nearly all there is; and the band probabilities are one-dimensional integrals that sum to one within and agree with a million counted samples.
The worths are computed on one pool of a million samples. First-order worth is over the within-band variance , with each band’s rate taken from the pool; counted worth is the ratio of mean squared errors over the pool’s runs, about the pool’s own plain rate, which charges the stratified estimates for any sampling error in that reference and so, if anything, understates them. Not measured: a score whose bands are not intervals in some coordinate, where the integral would be two-dimensional; companions without a known joint law, which is most real ones; and whether the earlier count for the t statistic’s bands, 2.99 on a hundred runs, is the same quantity as the 2.62 counted and 2.76 first-order here — both are within the spread a hundred runs allow.
Still open: a companion whose law is only approximately known
Every companion in these essays has an exact law, because the simulation was built so that one existed: the logarithms of a lognormal are normal. Most simulations have companions whose law is known only approximately — a statistic with an Edgeworth or saddlepoint approximation, a bootstrap quantity with a known limit — and the result above says that an estimated weight forfeits everything while an exact one keeps everything.
Between the two is a weight computed from an approximation that is accurate but not exact. Its error is not sampling error and does not shrink with any pilot; it is a fixed bias in each band’s probability, and it biases the rate by the sum of those errors times the band rates. Whether a saddlepoint-accurate weight — error a fraction of a per cent in each band — costs less than the variance it removes, and at what level of approximation error the trade turns, is computable on this same simulation by perturbing the exact weights by the approximation’s known error, and has not been done. It decides whether the method reaches the simulations where it would be most useful, which are the ones whose companions have no exact law.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The arm whose variance is its answer — both name closed form, monte carlo, pilot study, variance reduction
- A covariate with no levels — both name closed form, monte carlo, stratification
- A rate times a size — both name closed form, monte carlo, variance reduction
- A threshold in the tail — both name closed form, stratification, variance reduction
- An ordering that depends on the rule — both name bias-variance, closed form, monte carlo
- The effect a stopped trial reports — both name closed form, conditional distribution, monte carlo
Named objects
A flat tag is an object no other essay names yet.
Bias-varianceBinningClosed formConditional distributionLogistic regressionMonte CarloPilot studyStratificationVariance reduction