Every essay — page 10
More arms than two
Minimisation balances a trial by making the arms' counts even inside every factor level. With two arms there is one way to measure how uneven two counts are. With three there are several, they are all called minimisation, and they send different patients to different arms — while the ratio a trial was designed to deliver quietly disappears unless the score was told about it.
The analysis after three arms
An unadjusted analysis after a two-arm balancing rule rejects 0.6% of true nulls where it claims 5%. With three arms and a deterministic rule it rejects none at all — and the repair is the same repair, which is a sentence and a column in the model.
Guessing one arm in three
A balancing rule is guessable because it is balancing. With three arms the next assignment is worked out less often than with two — and by more, relative to what a guesser gets for nothing, and the damage they can do is almost unchanged.
The best of a set, and what the search costs
Comparing two forecasters is a test. Comparing eight is a multiplicity problem on top of a dependence problem, and the two do not separate: eight windows of one series carry the multiplicity of about two independent comparisons and eight separate problems carry eight, so a correction that charges for the number of models is wrong in both directions. The reference distribution has to be over the whole set — and once it is, the models nobody would have run turn out to cost more than the ones that were close.
What the other forecast adds
Two forecasters, one series, and two different questions about them. Which is more accurate has an answer that changes with the persistence of the series; whether either is redundant has an answer that never changes at all.
Eight forecasters and one benchmark
A set of forecasters is a multiplicity problem on top of a dependence problem, and the two do not separate. Eight windows of one series carry the multiplicity of two and a half independent comparisons; eight separate problems carry eight.
The models that were never in the running
A reference distribution for a set has to assume something about every candidate in it. Assuming that all of them are as good as the benchmark is what makes the reality check honest, and it is what sixteen hopeless candidates use to destroy it.
A distribution drawn from the null
Between nested models the ordinary comparison statistic has a null distribution centred at minus one and a 95% point of a quarter. A correction to its mean repairs the centre and leaves the shape; simulating the null repairs both.
What the design is asked to guarantee
Two halves of one question that were never separated: which parameters a design is for, and how much precision is enough. Protecting one parameter of a non-linear model over a range of its own values is a worst case of a ratio of determinants, and it is not a special case of either problem it is made of. Letting the experiment stop when it is precise enough is the adaptation with a theorem against it — the rule stops when its own noise estimate is low, so the interval it produces is short.
An efficiency that is a ratio
A design chosen for a model is not a design chosen for the parameter somebody wanted. Asking for one of two parameters moves the runs, unbalances the weights, and costs the other question exactly 15.07% — at every setting, because it is algebra.
Protecting one parameter over a range
A design for a non-linear model is optimal at a guess. A design for one of its parameters over a range of guesses is a worst case of a ratio of two determinants, and it is not a special case of either problem it is made of.
Stopping when it is precise enough
An experiment that runs until its estimate is precise enough is the natural design and the one with a theorem against it. Its two-stage cousin keeps its promise exactly, for every unknown spread, and pays twice the observations for it.
The interval after a stop it chose
A rule that stops when the estimated precision is good enough stops on the samples whose estimate was small. Its interval covers 90% and claims 95%, and a fresh sample of the same random size covers 95.4%.
Balancing what has no levels
Every balancing rule in the two fields before this one reads a level. Age, blood pressure and a baseline score have none, and the first thing that happens to them is that somebody invents some — a choice with a cost available in closed form before any data exists: a median split can see exactly 2/π of a normal covariate, so a rule that balances its two halves perfectly still leaves three fifths of a coin's imbalance. A rule that reads the number instead does not beat that by a factor; it beats it by a rate.
A covariate with no levels
Every balancing rule on this site reads a level. Age and blood pressure have none, so somebody cuts them into categories — and a median split can see exactly 2/π of a normal covariate, whatever the rule does with the halves.
The rule that reads the number
Stop categorising and let the rule read the covariate itself. What it should minimise is not an invented distance but the variance of the effect being estimated — and what comes back is not a better constant but a different rate.
What the balanced trial is worth
A rule that reads the covariate removes three quarters of the imbalance. An analysis that does not know it happened prices the imbalance anyway, rejects one true null in two hundred instead of one in twenty, and finds a real effect less often than a coin-tossed trial does.
Balancing more than one number
The criterion generalises to several covariates without a word changing, which makes the question what it is worth rather than whether it can be done. Each one added takes a share of the assignment's freedom, and the imbalance left in every one of them rises.
Searching among fitted models
The set field compares forecasts nobody estimated. A specification search compares a benchmark with a table of variants that all contain it, and two things change at once: every variant is behind before the search begins, by an amount with a closed form in the shape of the table, and the reference distribution for the winner can no longer be resampled from the data — it has to be generated from a model. The repair the nested case asks for, applied row by row, takes a table from its nominal level to a quarter.
A table of nested models
A benchmark and eight variants of it, each adding one thing. Every variant is behind before the search begins, by an amount that can be written down before the data exists — and the two most natural ways of reading the table are wrong in opposite directions.
A null with a model in it
The distribution to read the winner of a table against cannot be resampled from the data, because the data does not contain the null. It has to be generated from a model — which is the assumption the resampling was chosen to avoid.
The weight that is a vector
Two forecasts have a best combination and one number describes it. Eight have a best combination too, and the vector describing it puts nothing at all on the forecast with the smallest mean squared error.
When every null is true
A reality check assumes that every candidate in the set is exactly as good as the benchmark, which is a configuration nobody's data is ever in. Test a combination against its own parts and that configuration is not assumed — it is what the arithmetic makes true.
What the procedure may not read
Two restrictions that were cheap until now. A design for a non-linear model may not read the parameter values, and where the model has two of them the guess is a point in a rectangle: the weights that were exactly 1/√2 become a function, and a design robust in one coordinate turns out to guarantee no more than one built at a single point. A stopping rule may not read the mean it will report — and there is an exact way to obey that, at a price in the width of the interval rather than in the number of observations.
The guess with two numbers in it
Every optimal design for a non-linear model is optimal at a guess. Where the model has one parameter that moves the settings, that guess is a number and everything about it comes out in closed form; where it has two, three constants become functions and one of them becomes zero.
The worst case in two directions
A design that protects a range of one parameter is robust. Protect the range of one parameter while holding the other at a guess and the design is still robust, still has a guarantee, and guarantees no more than a design that protects nothing at all.
The rule that cannot see the mean
A sequential rule stops when its own estimate of the spread is small, which is more often on the samples whose spread came out low — so the interval afterwards is short. There is a way to keep updating the estimate and stop being able to see the mean at all.
What the blindfold costs
The exactly-covering rule pays for it in the width of the interval, and the block size is a dial between two costs that run in opposite directions. And on an interval whose width was fixed in advance, the same repair buys nothing at all.
The shape the covariate enters by
Every balancing rule in the two fields before this one optimises one over the variance of the treatment estimate in a model where the covariate enters linearly — which is not an assumption about the analysis, it is what the criterion is. Balancing the mean of a covariate removes exactly 2/π of the imbalance in a median split of it, a quarter of a threshold in the tail, and nothing at all from a quadratic. The ranking of the rules reverses between shapes, and against one of them every rule here is worse than a coin.
Balanced on the wrong function
A rule that reads a covariate's numbers halves the variance of the treatment estimate, if the covariate enters the outcome as a straight line. If it enters as a threshold the rule is worth a fifth of that, and if it enters as a curve every rule here is worse than a coin.
A threshold in the tail
How much of a threshold's imbalance a balanced covariate removes is a correlation, and the correlation is a closed form. At the median it is exactly 2/π — the same 2/π a median split throws away — and two standard deviations out it is an eighth.