Every essay — page 11
The shape the covariate enters by
Every balancing rule in the two fields before this one optimises one over the variance of the treatment estimate in a model where the covariate enters linearly — which is not an assumption about the analysis, it is what the criterion is. Balancing the mean of a covariate removes exactly 2/π of the imbalance in a median split of it, a quarter of a threshold in the tail, and nothing at all from a quadratic. The ranking of the rules reverses between shapes, and against one of them every rule here is worse than a coin.
Three functions of one number
A rule that balances the covariate is exposed to every shape the outcome might have. A rule that balances three functions of it costs two points of variance against the shape the first was built for and takes the worst case from a coin's to about half of it.
The analysis and the shape
An unadjusted analysis after a rule that read the covariate is too cautious — by a third against a linear outcome, by nothing at all against a quadratic. And an adjustment for the wrong function recovers almost none of the precision the right one would.
A search with no fixed point
The search field compares a benchmark with variants that all contain it. Take the fixed point away — let the benchmark be one candidate among sixteen, selected by the same data as its rivals — and the closed form for what every candidate is behind by survives intact, because it was never about nesting but about counting parameters. Three readings of one true null give 2.5%, 8.0% and 76.8%; Bonferroni is three times too strong on one table and not strong enough on another; and searching the benchmark makes the table harder to reject with rather than easier.
When the benchmark is a candidate
A specification search with a benchmark nailed down is the case with a closed form. Take the nail out — let the model that would have been reported be one of sixteen, chosen by the same data as its rivals — and the same true null is read three ways, at 2.0%, 7.8% and 76.2%.
The displacement is a parameter count
A nested variant is behind its benchmark out of sample before anything is searched for. The closed form for how far turns out to have nothing about nesting in it — only two integers and a window length — and it prices a table where no candidate contains any other.
Residuals that keep their own variance
A reference distribution for a search has to be generated from a fitted model, and the generator draws residuals. Four ways of drawing them keep four different things — and the one this site has reached for three times repairs nothing at all here.
The corner the test is calibrated at
"No candidate is better than the benchmark" is not a null but a face of a region, and a reality check is calibrated at one corner of it. Fill the table with candidates that are hopeless rather than equal and the test finds a genuine improvement 0.0% of the time.
Choosing what the rule reads
A balancing rule cannot tell one basis from another with the same span, so choosing what to hand it is choosing a subspace, and everything about it is exact geometry: what a rule removes of a shape is that shape's squared multiple correlation on the span. The maximin basis over a list of shapes is then a finite problem with an exact answer — and over a whole subspace the guarantee is exactly zero rather than small, which is why the list cannot be avoided. Randomising which functions the rule reads is worth twice the best fixed choice.
A basis is a subspace
A balancing rule cannot tell one basis from another with the same span, so choosing what to hand it is choosing a subspace — and then what it removes of any outcome shape is a projection, computable exactly, with no trial anywhere in it.
Which shapes are worth protecting
Choosing a basis by its worst case is a finite problem with an exact answer. The answer has no tie in it, which a maximin optimum is supposed to have — and the tie comes back, along with twice the guarantee, when the basis is drawn rather than chosen.
Where the guarantee is exactly zero
An experimenter who declines to name the shapes, and asks instead to be protected against anything in a class, is asking for a number that is not small but zero. Bounding the class is unavoidable, and the two ways of doing it choose different bases.
When the constraints run out
Every function added to a basis is a constraint the assignment has to satisfy with the same units. At sixteen units and a stated tolerance the admissible assignments run 3,874, then 1,006, then 314, then none — and the count is exact, because the assignment space is finite.
The block size as a schedule
The blinded rule's block size is a dial between an interval's degrees of freedom, a stopping rule's, and the overshoot. Letting it change during the run keeps the coverage exact — the argument never needed the sizes to be equal — and then runs into an identity: the two kinds of degree of freedom add to N − 1 on every run, so a schedule cannot make more of both. What it can do is spend each where it is worth most, which is worth a few per cent and is not the free lunch the dial looked like.
A block size that changes
The blinded rule's exactness never needed the blocks to be the same size. Letting the size be chosen from the contrasts as the run goes on leaves the coverage exactly where it was — and runs straight into an identity that says what a schedule can and cannot buy.
Two degrees of freedom, one total
The block size is a dial, and the two things a fixed-width procedure claims move in opposite directions along it. Divide the width by the square root of the sample size and one of them turns out to depend on the number of blocks and on nothing else.
What a schedule actually buys
Big blocks early and small blocks late is the right instinct and it does not take both ends of the trade, because there are not two ends to take. What it does take is the overshoot — about four per cent of the observations — and a steadier stopping point.
A schedule that reads the mean
The block sizes may be anything at all provided they are functions of the contrasts. Two natural schedules break that, in opposite directions — and the most natural mistake of the three is not a schedule at all but a stopping rule, at 86.87% coverage and fewer observations.
Scoring a search without spending data
A hold-out is expensive: every row spent scoring is a row not spent fitting. There are two standard ways of not paying — an information criterion, which predicts the hold-out from the in-sample numbers, and a resampling, which builds a reference distribution from one sample. Neither avoids what it replaces. The criterion's penalty *is* the displacement, computed rather than counted, and it beats a rolling hold-out at every split there is; both it and the resampling are undone by the same defect, which is rows that repeat each other, because both are counting independent things and there are fewer of those than there are rows.
A criterion is a prediction of the hold-out
A rolling hold-out spends half the sample measuring what a criterion computes from all of it. Against an oracle that is arithmetic rather than an estimate, the criterion gives up 0.01701 and the hold-out 0.03200 — and the number the hold-out reports for its own winner is optimistic by more than either.
Where the two searches cross
The obvious dial between a criterion and a hold-out is how much of the sample to hold out, and moving it never changes the answer. The dial that does is one nobody chooses — how much each row repeats the one before it — and the two rules change places at about 0.81.
Two defects and one resampling
Four resamplings, each the repair for one defect and wrong about the other. Put both defects in the same world and the statistic's 5% point is 3.8028, where the best of the four reaches 2.8326 — until a multiplier that stays on its own row and shares a sign with its neighbours reaches 2.9988.
How long a block a multiplier shares
Sharing a sign over more rows keeps more of the dependence and leaves fewer independent signs to build a distribution from. The bias falls from 1.6885 to 0.8479 and the spread rises from 1.3073 to 2.1716, and the rejection rate walks straight through its nominal level on the way from 11.3% to 1.3%.
When the set is too large to walk
Two constructions that were each measured by enumerating everything, at the size where enumerating everything stops being possible. A balancing dictionary over two covariates is an outer product, every inner product in it is still closed form, and a rule holding every main effect of both removes exactly nothing of any interaction. A trial of two hundred units has assignments that cannot be walked — and the acceptance rate turns out not to depend on the size of the trial at all, so the exhaustion a sixteen-unit count runs into is a fact about sixteen units.
A dictionary that is a product
Two covariates make what a balancing rule may read an outer product — eight main effects and sixteen interactions — and every inner product in it is still closed form. What a rule holding all eight main effects removes of a pure interaction is not small. It is zero.
Where the enumeration stops
A maximin over an eight-function dictionary is a walk over seventy subsets. Over twenty-four it is 735,471 at eight functions, and the exchange algorithm that replaces the walk scores 421. What licenses the second curve is four sizes where both exist and agree, which is a weaker warrant than it looks.
A count that has to be estimated
At sixteen units the admissible assignments can be counted by walking all 12,870 of them. At four hundred there are about 2^393.70, and the share admitted is 0.31885 against a closed form of 0.31818 that has no trial size in it at all. The exhaustion a small trial runs into is a fact about small trials.
What a reference distribution costs to sample
A randomisation test on a trial too large to enumerate has to sample its reference distribution, at 1/p attempts per draw and a p-value resolved to 1/(B + 1). Six constraints cost 9,878 attempts per thousand draws, and a thousand draws resolve p to 9.99·10⁻⁴ and not one digit finer.
A promise about two arms
The fixed-width interval whose coverage is exact was built for one mean, and almost nothing anybody runs an experiment for is one mean. The construction survives with the harmonic effective size in place of the block size, and four things do not: a unit of that size costs four observations rather than one, the second variance puts Neyman's allocation in reach of a blinded rule, the conservation identity comes up one degree of freedom per block short, and there is an exact estimator under either of two conditions and none under both.
A width promised for a difference
The exact fixed-width interval was built for one mean. Two arms make the target 42.7 units of effective size and each unit costs four observations, so the same promise about a difference costs 169.4 rather than 42.7 — and the theorem survives untouched with the harmonic size in place of the block size.
Which weights are the inverse variances
There is an exact estimator when the two arms share a variance and another when every block has the same two counts, and between them they cover every trial anybody designs on purpose. In the corner where neither holds, both cover 98.45% instead of 95%, and the only estimator at its level is the one with no theorem behind it.