Concept

Efficiency — where it appears

How much of the information an ideal procedure would extract a given one actually does, expressed as a ratio and read as a multiplier on the experiment. Read as a multiplier it says how many more runs the worse procedure needs, which is the form an experimenter can budget against.

Named by 20 essays across 12 fields — each of them below, with the objects they name alongside it.

The price of each thing the rule is not told. What each rule gives up against the best model available, at a persistence of 0.85 on a fifteen-candidate table, over 400 draws. Reading down: least squares with the ordinary penalty; the whitening at the true ρ; the same at a ρ̂ estimated per candidate; that rule with the term the Gaussian likelihood carries and it omits; a Bartlett-tapered Ω̂ estimated once from the fullest candidate at L = 8; the same estimated per candidate; and the truncated Ω̂, which exists on only 45.0% of draws and is averaged over those. Knowing ρ recovers 89.9% of what counting rows gives up, estimating it 83.4%, and estimating a whole covariance 75.6%.

A covariance with no parameter in it

The whitening that repairs a criterion is told the dependence is a first-order autoregression and left to find one number. A real dependence is not one number, and the obvious estimate of it is not a covariance matrix.

banded · Dependence
Four rules, three shapes, and no ordering that survives. The variance of the unadjusted treatment estimate under each rule, as a fraction of the variance a coin gives, over 450 trials of 200 units each. Against a covariate that enters linearly the rule that reads the number nearly halves it. Against a threshold at 1 it removes about a fifth. Against a quadratic every rule here is at or worse than a coin — they are all optimising a criterion that is one over the variance of an estimate in a model this outcome does not obey, and a constraint that helps nothing still costs something. Nothing in a trial says which column it is in.

Balanced on the wrong function

A rule that reads a covariate's numbers halves the variance of the treatment estimate, if the covariate enters the outcome as a straight line. If it enters as a threshold the rule is worth a fifth of that, and if it enters as a curve every rule here is worse than a coin.

shape · Blocking
One curve, three designs, and the disagreement is the weights. The compartmental response exp(−θ₁t) − exp(−θ₂t) at the guess θ = (0.2, 1.2), with three designs underneath it. Each mark is a setting and its height is the share of the experiment spent there. The D-optimal design splits the runs equally between two settings — that is the determinant's answer and it is equal for every model of this kind. The design for the first parameter alone puts 1.9% of the experiment at its early setting and the rest at its late one, because the early runs are there to identify the nuisance and nothing more. The design that protects a 4-fold rectangle needs 3 settings and spends 9.1% at the earliest of them.

The guess with two numbers in it

Every optimal design for a non-linear model is optimal at a guess. Where the model has one parameter that moves the settings, that guess is a number and everything about it comes out in closed form; where it has two, three constants become functions and one of them becomes zero.

blind · Local design
Exact in the corner, where nothing was. Coverage of a nominal 95% interval on five designs, at a required half-width of 0.3. The first four are the two-arm field's own and the fifth is its corner — two variances, block sizes that swing by eight, and an allocation that alternates between five to one and one to five — where neither of that field's two conditions holds. The effective-size weights over-cover there at 98.40%; the weights h_b(λ) = (1/m_A + λ/m_B)⁻¹ cover at 94.84%, and at 94.84% when λ is estimated from the within-arm contrasts rather than known. Nothing here is supposed to move.

Weights that need only a ratio

A fixed-width interval about a difference is exact under either of two conditions and under neither in the corner. It is exact there too, and the only thing it needs is how much larger one arm's variance is than the other's.

corner · Nuisance
A threshold in the tail is a threshold nothing balances. The share of a coin's imbalance in an indicator 1{x > c} that survives a rule which balances the covariate itself. The smooth curve is 1 − ρ² with ρ = φ(c)/√(p(1−p)), a closed form with no trial in it; the points are counted over 500 trials of 200 units at each threshold. At the median the two agree that about a third survives — the removed share is exactly 2/π — and by two standard deviations 86.9% survives. The closed form is exact in the limit and optimistic by a few points at this many units, because the rule balances the sample's mean rather than the population's.

A threshold in the tail

How much of a threshold's imbalance a balanced covariate removes is a correlation, and the correlation is a closed form. At the median it is exactly 2/π — the same 2/π a median split throws away — and two standard deviations out it is an eighth.

shape · Blocking
How wrong the ratio is allowed to be. λ enters only through the weights, so misstating it leaves the estimate unbiased and moves two things — the interval's calibration and its efficiency — both of which are closed forms of the design. Coverage stays at its level over a factor of two in either direction (94.93% at half the truth, 94.27% at twice it) and starts to go at a factor of five. An estimate on hundreds of within-arm degrees of freedom is never wrong by anything like that, which is what makes the feasible rule usable rather than merely definable.

Blinded, and still exact

The one number the exact interval needs is a ratio of within-arm spreads, which is a contrast and contains no mean — so a rule forbidden to look at the effect may compute it, on more degrees of freedom than the interval itself has.

corner · Width
Four designs, and what each of them guarantees. The worst Ds-efficiency each design achieves anywhere in a rectangle of parameter values 4 times wide in each coordinate. The design built at the guess guarantees 7.3%: it is perfect where it was built and nearly useless at one corner. The D-optimal design at the same guess guarantees 25.6% — it answers the wrong question everywhere and is therefore not concentrated on being right anywhere. The third is the one this field exists to measure: maximin over the first parameter's whole range, with the second held at its guess. It guarantees 7.0%, which is no better than the design that protects nothing. Protecting both is worth 45.9%, and it needs 3 settings to do it.

The worst case in two directions

A design that protects a range of one parameter is robust. Protect the range of one parameter while holding the other at a guess and the design is still robust, still has a guarantee, and guarantees no more than a design that protects nothing at all.

blind · Criterion
The guarantee, as the basis is allowed more functions. The lower line is the best worst case over the six named shapes for a basis of each size, found by scoring every subset of the dictionary — an exact answer, since the problem is finite. One function guarantees 2.3%, which is nearly nothing; three guarantee 59.0% and the basis that does it is the covariate, its square and its cube, with no indicator in it. The upper line is the same problem with the basis drawn rather than fixed, which is worth 2.09 times as much at two functions and 1.32 at three. The two lines converge because a basis large enough to protect everything has nothing left to randomise over.

Which shapes are worth protecting

Choosing a basis by its worst case is a finite problem with an exact answer. The answer has no tie in it, which a maximin optimum is supposed to have — and the tie comes back, along with twice the guarantee, when the basis is drawn rather than chosen.

basis · Blocking
The weights may not read the block they weight. A weighted least squares decomposition needs weights that are constants, or at least independent of the differences they multiply. One λ̂ pooled across the trial is estimated on hundreds of degrees of freedom and is effectively a constant; a λ̂ estimated inside each block is estimated on that block's own two or three, and is correlated with the difference it weights. Coverage falls from 94.68% to 82.76% — and the interval gets wider while doing it, 0.5163 against 0.3024, which is the signature of weights that are noise.

The condition that cannot be dropped

The weights may not read the block they weight. Estimate the variance ratio inside each block rather than across the trial and the coverage falls to 83% — on an interval that is at the same time seventy per cent wider.

corner · Allocation
The row count entered twice, and a penalty is one place. Regret against the best available model as the errors are made persistent. Counting rows more than quadruples; both penalty repairs — the trace, and the scalar effective sample size — are worse than it at every persistence measured; and whitening the sample and keeping the ordinary penalty falls, recovering 86.9% of what counting rows gives up at ρ = 0.85. Doing it at an estimated ρ recovers 80.6%, so having to estimate the dependence from the rows being selected on costs 7.3% of what knowing it is worth. Mallows' forms are drawn beside the logarithmic ones and behave the same, which is what rules the linearisation out.

The repair that was exact and made it worse

A penalty computed from the trace is exactly the optimism it estimates, and selecting with it gives up a fifth more than not correcting anything. The row count entered the criterion twice, and a penalty is the second place.

effective · Forecast
What a rule reads, against what the outcome uses. The variance of the treatment estimate relative to a coin's, for four things a rule might balance against three shapes the outcome might have, over 350 trials of 200 units. The diagonal is the easy part — a rule that reads the function the outcome uses removes about half the variance. What the table is for is the off-diagonal: reading the covariate alone is worth nothing against a quadratic (0.755), and reading all three is worth nearly as much against every shape as the matching rule is against its own (0.532, 0.493, 0.514).

Three functions of one number

A rule that balances the covariate is exposed to every shape the outcome might have. A rule that balances three functions of it costs two points of variance against the shape the first was built for and takes the worst case from a coin's to about half of it.

shape · Criterion
The overshoot is the last block size and nothing else. A run stops at a multiple of its own block sizes and cannot land between them, so it ends past its own target by about half a block. Fixed sizes overshoot by 1.5, 2.2, 3.1, 4.7, 8.5 observations as the size goes 2, 3, 5, 8, 16. Every schedule here ends in blocks of two and every one of them lands where blocks of two land — 1.62, 1.32, 1.37 against 1.48 — while having spent most of the run inside blocks four and eight times larger. That is the one thing on this page a schedule genuinely takes from both ends.

What a schedule actually buys

Big blocks early and small blocks late is the right instinct and it does not take both ends of the trade, because there are not two ends to take. What it does take is the overshoot — about four per cent of the observations — and a steadier stopping point.

pace · Nuisance
A wrong weight costs width; a random weight costs level. Five weightings on a trial whose variance ratio drifts by a factor of 20.1 between the first block and the last, over 4000 runs. The rule that knows every λ_b covers at 95.1% and sets the width. One ratio for the whole trial is wrong for every block and costs nothing in level — 94.8% — while being 20% wider; equal weights are calibrated by an identity and 22% wider. The ratio estimated inside each block is the only rule aimed at the quantity that actually varies, and it is the only one that misses the level, at 92.0%: a weight computed from a handful of degrees of freedom is mostly noise, and noise in a weight is not a wrong weight. Modelling the drift across blocks recovers the oracle's width at 94.8%.

A ratio that changes between blocks

A wrong weight costs width and a random weight costs level. The rule aimed at the quantity that actually varies is the only one that misses its own coverage, and the rule that models it across blocks recovers the whole of what knowing it is worth.

blocks · Nuisance
Fitting them together is worth something under one law. Four fits of the same regression under four dependences: least squares, the two-step plug-in every whitened rule in this collection runs, the coefficients and the band maximised together, and a whitening at the law's own covariance that nobody has. Under the moving average — the one law the band family contains — the joint fit beats the two-step by 0.0077 at 3.3 paired standard errors. Under the autoregression, long memory and the break it is a tie: 0.4, 1.0, 0.3 standard errors. That is the same ordering the likelihood gap gave, arrived at through the coefficients rather than through the objective.

What fitting them together buys

Maximising over the coefficients and the covariance together beats the two-step under one of four dependences and ties under the other three. It is the one the band family contains, and the likelihood said so before any coefficient was compared.

family · Dependence
What the exactness costs, and the dial it is bought with. The median half-width of the interval each rule reports, at a requirement of 0.4 and a first look after 5 observations. The flat line is the interval a practitioner writes at the purely sequential rule's stopping time, which covers 91.72% rather than 95%. The curve is the blinded rule, which covers its nominal level at every block size: it reads b − 1 degrees of freedom where the other reads n − 1, and pays for the exactness in width. The best block size is 3, at 0.4712. Larger blocks give the stopping rule a better estimate and the interval a worse one, and the two costs go opposite ways, which is what puts the minimum in the middle.

What the blindfold costs

The exactly-covering rule pays for it in the width of the interval, and the block size is a dial between two costs that run in opposite directions. And on an interval whose width was fixed in advance, the same repair buys nothing at all.

blind · Nuisance
What the guess is worth, when it is worth anything. The variance cost of an even split relative to the variance-minimising one for a risk difference, against the first arm's proportion, with the second at 0.3. The cost is a pure number: it does not depend on the trial's size. It is exactly zero at 0.3 and at 0.70, where the two arms have the same p(1 − p); it is 0.19% at a half and 4.36% at a tenth. Across the whole range from a tenth to nine tenths it never exceeds 4.36%, which is what the variance-minimising rule is worth here — and what it is worth is the reason it is safe to use with a guess.

The arm whose variance is its answer

With a binary outcome the allocation rule is a function of the proportions the trial exists to estimate. It costs at most 4.36% of variance to ignore it anywhere between a tenth and nine tenths, because √(p(1−p)) stays within a factor of two of its peak across 98% of the unit interval.

allocation · Allocation
One row at x = 9 on a wrong line, fitted three ways. Least squares gives slope −0.389, Huber 0.171 with the extra row at weight 0.115, and least trimmed squares 0.420, fitted to the 12 rows it keeps. The twenty clean rows alone give 0.495. Open circles are the rows the trimmed fit leaves out.

A robust loss and a far x

One far row drags least squares to a slope of −0.389. Huber's loss, the standard robust line, reaches only 0.171, and carried further out the same row gets its full weight back. Least trimmed squares reads 0.420 at every distance, and at the normal model keeps 7.13% of least squares' efficiency to do it.

regression · Leverage
Three contrasts on one dataset, three different splits. The variance-minimising allocation for each of three ways of reporting the same two-arm comparison, against the first arm's proportion, with the second at 0.1. A risk difference wants the arm with the larger p(1 − p) to get more units; a log odds ratio wants it to get fewer, and the two curves are exact reflections of each other in the half line. A log risk ratio wants something else again. At a first-arm proportion of 0.6 they ask for 62.0%, 21.4% and 38.0% of the units. A trial reporting more than one of them cannot be optimal for either.

Two contrasts, one split

A risk difference wants 62.0% of the units in the first arm, a log risk ratio wants 21.4% and a log odds ratio wants 38.0% — on one dataset, with one pair of proportions. The difference's rule and the odds ratio's are exact reflections of each other, so no split can be near-optimal for both.

allocation · Allocation
One row at x = 9 on a wrong line, fitted three ways. Least squares gives slope −0.389, Huber 0.171 with the extra row at weight 0.115, and least trimmed squares 0.420, fitted to the 12 rows it keeps. The MM-estimator carried on from the trimmed fit gives 0.479, with the extra row at weight 0.000 and the scale fixed at 0.319. The twenty clean rows alone give 0.495. Open circles are the rows the trimmed fit leaves out.

The start an efficient robust line inherits

The MM-estimator carries a trimmed fit on through a redescending loss, and it does what it promises on one far row: slope 0.479 at every distance, the row at weight exactly zero, and 87.2% of least squares' efficiency at twenty rows. What it cannot do is choose. At eight far rows of twenty the exact trimmed fit picks the wrong half on 111 datasets; the efficient step repairs none of them, spoils none of the other 89, and ends nearer the wrong line than the start did.

regression · Leverage
Two companions on one simulation, two hundredfold apart. How many times as many draws each companion is worth, on the same 4,000 simulated samples of 40 observations. The coverage of the interval is estimated with the observed count as its companion, whose expectation is 12 exactly; they correlate at 0.2665 and the companion is worth 1.08 times the draws. The expected width is estimated with p̂(1 − p̂) as its companion, whose expectation is 0.20475 exactly; they correlate at 0.9977 because the width is a monotone function of it, and the companion is worth 214 times the draws — 856 thousand simulated samples' worth of precision from four thousand.

The check worth more than the check

The same exactly known companion that verifies a simulation can sharpen it. On one set of four thousand draws, one companion is worth 1.08 times the draws and another is worth 214 times them, and the factor is 1 − ρ² with nothing else in it.

method · Routes

Named alongside it

The objects these essays reach for when they reach for this one.

Nuisance parameterClosed formCoverageVariance reductionAllocation ruleDegrees of freedomFixed-width intervalBlindingInterval widthModel misspecificationTreatment effectAllocation ratio

All concepts