The collection

Every essay — page 11

Essays 241 to 264 of 436, in the same order.

The shape the covariate enters by

Every balancing rule in the two fields before this one optimises one over the variance of the treatment estimate in a model where the covariate enters linearly — which is not an assumption about the analysis, it is what the criterion is. Balancing the mean of a covariate removes exactly 2/π of the imbalance in a median split of it, a quarter of a threshold in the tail, and nothing at all from a quadratic. The ranking of the rules reverses between shapes, and against one of them every rule here is worse than a coin.

A search with no fixed point

The search field compares a benchmark with variants that all contain it. Take the fixed point away — let the benchmark be one candidate among sixteen, selected by the same data as its rivals — and the closed form for what every candidate is behind by survives intact, because it was never about nesting but about counting parameters. Three readings of one true null give 2.5%, 8.0% and 76.8%; Bonferroni is three times too strong on one table and not strong enough on another; and searching the benchmark makes the table harder to reject with rather than easier.

Choosing what the rule reads

A balancing rule cannot tell one basis from another with the same span, so choosing what to hand it is choosing a subspace, and everything about it is exact geometry: what a rule removes of a shape is that shape's squared multiple correlation on the span. The maximin basis over a list of shapes is then a finite problem with an exact answer — and over a whole subspace the guarantee is exactly zero rather than small, which is why the list cannot be avoided. Randomising which functions the rule reads is worth twice the best fixed choice.

The block size as a schedule

The blinded rule's block size is a dial between an interval's degrees of freedom, a stopping rule's, and the overshoot. Letting it change during the run keeps the coverage exact — the argument never needed the sizes to be equal — and then runs into an identity: the two kinds of degree of freedom add to N − 1 on every run, so a schedule cannot make more of both. What it can do is spend each where it is worth most, which is worth a few per cent and is not the free lunch the dial looked like.

Scoring a search without spending data

A hold-out is expensive: every row spent scoring is a row not spent fitting. There are two standard ways of not paying — an information criterion, which predicts the hold-out from the in-sample numbers, and a resampling, which builds a reference distribution from one sample. Neither avoids what it replaces. The criterion's penalty *is* the displacement, computed rather than counted, and it beats a rolling hold-out at every split there is; both it and the resampling are undone by the same defect, which is rows that repeat each other, because both are counting independent things and there are fewer of those than there are rows.

When the set is too large to walk

Two constructions that were each measured by enumerating everything, at the size where enumerating everything stops being possible. A balancing dictionary over two covariates is an outer product, every inner product in it is still closed form, and a rule holding every main effect of both removes exactly nothing of any interaction. A trial of two hundred units has assignments that cannot be walked — and the acceptance rate turns out not to depend on the size of the trial at all, so the exhaustion a sixteen-unit count runs into is a fact about sixteen units.

A promise about two arms

The fixed-width interval whose coverage is exact was built for one mean, and almost nothing anybody runs an experiment for is one mean. The construction survives with the harmonic effective size in place of the block size, and four things do not: a unit of that size costs four observations rather than one, the second variance puts Neyman's allocation in reach of a blinded rule, the conservation identity comes up one degree of freedom per block short, and there is an exact estimator under either of two conditions and none under both.

FieldsThreadsSeriesConceptsFigure librarySearch