Concept

Variance ratio — where it appears

How much larger one arm's variance is than the other's, written λ. It is the only unknown in the weights that make a two-arm fixed-width interval exact, it is a ratio of within-arm contrasts so a blinded rule may compute it, and misstating it by a factor of two costs half a point of coverage.

Named by 13 essays across 6 fields — each of them below, with the objects they name alongside it.

The stopping rule costs more than the weighting does. Coverage over 2000 runs of a trial whose variance ratio drifts by a factor of twenty, at three ways of deciding when to stop. Twelve blocks fixed in advance is the top line and reproduces what a trial of fixed length delivers. Stopping when the reported interval is short enough is the bottom line, and it costs between 3.0% and 5.5% of coverage — including for the rule that is told every block's true ratio, which is what says the shortfall belongs to the stopping and not to the weights. Stopping on a width predicted from the within-arm sums of squares is the middle line, and it is back at the fixed-length values. The standard error on each point is 0.49%.

A width the trial has to stop for

The weighting that covers at 94.9% on twelve blocks covers at 91.5% when the trial stops as soon as its interval is short enough — and so does the rule that is told every block's true variance ratio. The shortfall is the stopping, not the weights.

stop · Stopping
What a sample shows, and what the algebra does. The difference between a rectangular block's implied long-run variance and a trapezoidal one's, as a share of the truth. Above the axis the rectangle is less biased and below it the trapezoid is. The heavy line is exact — computed from the law's own autocovariances — and it crosses at 19.2. The others are what samples of 120, 240, 480, 960 rows report, and every one of them exaggerates whichever window is ahead: at ℓ = 20, where the exact difference is 0.28 points, a sample of 120 rows shows 4.31 points — 15 times larger. That is the number the earlier reading of this comparison was missing: three tenths of a point is what the algebra says and not what a hundred and twenty rows report.

The gap a sample shows

The exact difference between two block windows at a block length of twenty is three tenths of a point. What a hundred and twenty rows report is four and a third, because the autocovariances the window is applied to are attenuated too.

crossing · Bootstrap
Exact in the corner, where nothing was. Coverage of a nominal 95% interval on five designs, at a required half-width of 0.3. The first four are the two-arm field's own and the fifth is its corner — two variances, block sizes that swing by eight, and an allocation that alternates between five to one and one to five — where neither of that field's two conditions holds. The effective-size weights over-cover there at 98.40%; the weights h_b(λ) = (1/m_A + λ/m_B)⁻¹ cover at 94.84%, and at 94.84% when λ is estimated from the within-arm contrasts rather than known. Nothing here is supposed to move.

Weights that need only a ratio

A fixed-width interval about a difference is exact under either of two conditions and under neither in the corner. It is exact there too, and the only thing it needs is how much larger one arm's variance is than the other's.

corner · Nuisance
How wrong the ratio is allowed to be. λ enters only through the weights, so misstating it leaves the estimate unbiased and moves two things — the interval's calibration and its efficiency — both of which are closed forms of the design. Coverage stays at its level over a factor of two in either direction (94.93% at half the truth, 94.27% at twice it) and starts to go at a factor of five. An estimate on hundreds of within-arm degrees of freedom is never wrong by anything like that, which is what makes the feasible rule usable rather than merely definable.

Blinded, and still exact

The one number the exact interval needs is a ratio of within-arm spreads, which is a contrast and contains no mean — so a rule forbidden to look at the effect may compute it, on more degrees of freedom than the interval itself has.

corner · Width
Two promises, and no rule here keeps both. A fixed-width procedure promises two things: that the interval covers at its nominal rate, and that it is no wider than the width asked for. Over 1500 runs of the modelled weighting, a rule that stops when the interval it will report is short enough keeps the width — only 2.0% of runs come out wider than 0.34 — and covers at 91.13% against a nominal 95%. A rule that stops on a width predicted from the within-arm sums of squares covers at 94.80% and comes out wider than promised on 42.3% of runs. The two promises are in conflict because keeping the second one exactly requires conditioning on the very quantity that has to be independent of the stopping time for the first.

Stopping on the arms

The width a trial will report is predictable from quantities the interval is not about. A rule that stops on the prediction covers at 94.5% where one that stops on the interval covers at 91.5, and it costs two blocks and half of the width promise.

stop · Width
What each error is a claim about, and what the claim comes out as. Each variance estimate's average over 20000 draws, divided by the variance the slope actually has across those same draws, at 80 rows with the error variance leaning towards the edges of the design (γ = 0.8). One is a standard error that is right. The model-based estimate reads 0.6081 of the spread, so its standard error is 77.98% of the one it should report; the four robust corrections read 0.9576, 0.9821, 0.9961, 1.0362. Two further routes agree with the count and share none of its arithmetic: n times the counted variance is 4.8905 against a population sandwich of 4.9200, and the counted ratio of the two standard errors is 1.2799 against a closed form of 1.2806.

The bread and the filling

The robust standard error is not a safety margin. At one setting of the error variance it is 1.2806 times the model-based one and at another it is 0.8246 times it, and the sign of a single dial decides which.

sandwich · Misspecification
Three answers to how much sample is left. What a set of inverse-probability weights leaves of the treated arm, by three routes, at six settings of the assignment rule. The integral 1/(π∫φ/e) reads the whole covariate space and falls from 0.9392 to 2.655e-3. Kish's effective size counted in samples of 600 falls only to 0.2861, because almost all of the integral's fall is in a region a sample of six hundred never draws from. And the fraction the variance of the weighted mean actually delivers is lower again — 0.1155 — because the variance is the average of one over the effective size and the effective size averaged is not the same number. At the widest overlap all three agree to 0.05%.

How many observations a weight leaves

Kish's effective sample size is exact — for an outcome whose mean does not move with the covariates the weights are built from, the studentised variance reads 1.0680 where the formula says one. For the population's own outcome the same reading is 6.769, rising to 52.497.

weights · Weighting
The weights may not read the block they weight. A weighted least squares decomposition needs weights that are constants, or at least independent of the differences they multiply. One λ̂ pooled across the trial is estimated on hundreds of degrees of freedom and is effectively a constant; a λ̂ estimated inside each block is estimated on that block's own two or three, and is correlated with the difference it weights. Coverage falls from 94.68% to 82.76% — and the interval gets wider while doing it, 0.5163 against 0.3024, which is the signature of weights that are noise.

The condition that cannot be dropped

The weights may not read the block they weight. Estimate the variance ratio inside each block rather than across the trial and the coverage falls to 83% — on an interval that is at the same time seventy per cent wider.

corner · Allocation
A wrong weight costs width; a random weight costs level. Five weightings on a trial whose variance ratio drifts by a factor of 20.1 between the first block and the last, over 4000 runs. The rule that knows every λ_b covers at 95.1% and sets the width. One ratio for the whole trial is wrong for every block and costs nothing in level — 94.8% — while being 20% wider; equal weights are calibrated by an identity and 22% wider. The ratio estimated inside each block is the only rule aimed at the quantity that actually varies, and it is the only one that misses the level, at 92.0%: a weight computed from a handful of degrees of freedom is mostly noise, and noise in a weight is not a wrong weight. Modelling the drift across blocks recovers the oracle's width at 94.8%.

A ratio that changes between blocks

A wrong weight costs width and a random weight costs level. The rule aimed at the quantity that actually varies is the only one that misses its own coverage, and the rule that models it across blocks recovers the whole of what knowing it is worth.

blocks · Nuisance
The fixed-width trial's coverage when the outcomes are not normal, for both stopping rules. normal: stopping on the arms 94.05% after 18.1 blocks, on the report 89.95%; log-normal, skewness 0.95: stopping on the arms 94.70% after 18.5 blocks, on the report 90.80%; log-normal, skewness 2.26: stopping on the arms 94.15% after 19.3 blocks, on the report 90.25%; log-normal, skewness 4.75: stopping on the arms 94.45% after 18.7 blocks, on the report 90.50%; t, five degrees of freedom: stopping on the arms 94.35% after 18.3 blocks, on the report 90.30%; skewness 4.75, arm A only: stopping on the arms 93.80% after 26.0 blocks, on the report 89.90%; skewness 4.75, arm B only: stopping on the arms 93.60% after 14.2 blocks, on the report 89.95%; equal variances, normal: stopping on the arms 94.75% after 11.4 blocks, on the report 90.90%; equal variances, skewness 4.75: stopping on the arms 94.05% after 11.1 blocks, on the report 92.00%.

A width rule on skewed outcomes

The blinded fixed-width rule rests on a within-arm spread being independent of the arm means, which only normal samples guarantee. On outcomes with a skewness of 4.75 the independence fails and the overall coverage barely notices — 93.60% to 94.70% across every shape counted, against 94.05% on normal outcomes. What skew moves is the runs that stop by twelve blocks, which cover about 90% with the skew in one arm, and the trial's length: a variance ratio corrected on normal theory lengthens it from 18.1 blocks to 26.0 with the skew in the first arm and shortens it to 14.2 with the skew in the second.

stop · Width
Where the bias lands. The drift in the log variance ratio, fitted across 12 blocks over 4000 trials. E[log λ̂_b] is log λ_b plus ψ(k_B/2) − log(k_B/2) − ψ(k_A/2) + log(k_A/2), which depends on nothing but the degrees of freedom — so the tempting sentence is that it goes into the intercept and leaves the slope alone. It does not, because the blocks alternate between allocations and the alternation is correlated with the covariate being fitted: the lopsided blocks carry 0.5383 of bias and the even ones carry none. Uncorrected the slope reads 1.5597 against a truth of 1.5, which is 8.0 standard errors. Subtracting the two digammas block by block leaves 1.4976.

The bias that lands in the slope

The bias in a log variance estimate depends on nothing but its degrees of freedom, so it goes into the intercept — unless the degrees of freedom alternate with the design, which is exactly what a block-randomised trial makes them do.

blocks · Width
Estimating a weight you already know is worth doing. The variance of an inverse-probability estimate weighted by a propensity fitted from the sample, over the variance of the same estimate weighted by the true propensity, paired on the same 500 samples of 600 units at each of five settings. Every reading is below one: the stabilised estimator keeps 27.8% of its true-weight variance where the assignment is nearly a coin toss and 72.0% where it is nearly decidable, and the unstabilised one 30.0% and 49.3%. Neither estimator is materially biased, so this is a variance rather than a trade. The true weights are right about the population and know nothing about the draw; the fitted weights are the value that sets this draw's own imbalance to zero, and that imbalance was what the variance was made of.

The estimated weight is the better one

The propensity is known exactly here, so it can be weighted by — and estimating it from the same data and weighting by that gives a variance ratio of 0.4769 on paired draws. The reason is a projection: the draw's own imbalance explains 56.33% of the true-weight variance and 0.05% of the estimated-weight one.

weights · Weighting
Weights that balance a sample by construction. What three sets of weights leave of the standardised difference between the arms on each covariate, as a root mean square over 1200 samples of 600 units. The true propensity leaves 0.1317 and 0.1186 — a sampling error, since it is right about the population and knows nothing of the draw. A likelihood fit leaves 0.0770 and 0.0657, having absorbed part of the draw's imbalance as a side effect of fitting the treatment. Weights fitted so that each arm's weighted means are the sample's leave 1.4e-14 and 1.2e-14, which is the arithmetic's floor rather than a small number: the largest gap between a weighted arm mean and the sample mean in any draw is 9.8e-14.

A weight fitted to balance

Weights fitted so that each arm's weighted covariate means equal the sample's leave a difference of 1.4×10⁻¹⁴ between the arms and give the estimate a third of the variance of weights fitted by likelihood — 0.011883, within a relative 5.8% of the bound no estimator can beat. In the world where the assignment carries a square nobody named, the same exact balance leaves the square further apart than no weighting at all, and where the outcome carries it too the estimate is wrong by 0.6973 with an interval that covers 1.5%.

weights · Weighting

Named alongside it

The objects these essays reach for when they reach for this one.

CoverageEstimated varianceFixed-width intervalBlindingInterval widthInverse variance weightingWeighted least squaresBlockingClosed formDegrees of freedomAllocation ratioEfficiency

All concepts