Every essay — page 12
A promise about two arms
The fixed-width interval whose coverage is exact was built for one mean, and almost nothing anybody runs an experiment for is one mean. The construction survives with the harmonic effective size in place of the block size, and four things do not: a unit of that size costs four observations rather than one, the second variance puts Neyman's allocation in reach of a blinded rule, the conservation identity comes up one degree of freedom per block short, and there is an exact estimator under either of two conditions and none under both.
The degrees of freedom in the sums
One arm partitions N − 1 exactly. Two arms give the rule N − 2b and the interval b − 1, which is short by one per block — and the missing ones are in the block sums, which are correlated with the differences at −0.79 and are usable anyway.
What a two-arm rule may not pool
A spread computed "within the block" without the arm label carries a share of the effect, so the trial runs 173 observations at a null and 282 at an effect of 1.5. The stopping rule is reading the thing it exists to measure, and the phrase that produced it is one word long.
Counting what is independent
Akaike's penalty is 2q because the optimism of a fit is a trace and the trace collapses to the parameter count when the rows are independent. Repair the trace and the criterion gets worse, on either scale, because the row count entered twice and a penalty is the second place: whiten the fit and the ordinary penalty is correct again, which recovers 88% of what counting rows gives up. Estimating the dependence from the rows being selected on costs 8% of that. And a multiplier resampling cannot keep more dependence than the residuals have, which is a ceiling rather than a tuning problem.
A penalty is a trace
Akaike's 2q is not a count of coefficients. It is the answer a trace collapses to when the rows are independent — and once they are not, the trace is still the right object and is no longer the count.
One number for a table of candidates
An effective sample size is a real quantity, it is exactly right about one thing, and that thing is a mean. Substituted into Akaike's criterion it changes nothing at all, because the penalty it is meant to fix has no sample size in it.
The repair that was exact and made it worse
A penalty computed from the trace is exactly the optimism it estimates, and selecting with it gives up a fifth more than not correcting anything. The row count entered the criterion twice, and a penalty is the second place.
The residuals are not the errors
A fit removes the part of the errors lying in its own column space, and a persistent design's column space is itself slow — so what is left behind is smoother than what went in, at every lag, by an amount that grows with the lag.
What a multiplier cannot keep
Two reasons were named for the quarter a blocked resampling falls short, and taking either away makes the gap larger. What is left is a bound — a multiplier can only take dependence out, and the residuals' own is already below the errors'.
When the two are not independent
The dictionary that is an outer product needs the covariates independent, and the count that is a rate times a binomial coefficient needs the draws independent. Neither holds in a real trial. A linearisation turns the missing fourth-order expectation into Mehler's formula applied to products, which makes the whole geometry closed at any correlation — and the result it was built to test does not survive: what a rule removes of a pure interaction is exactly nothing at independence and 4ρ²/(1+ρ²)² everywhere else. A walk on the admissible set is exactly uniform, needs a burn-in, and is dearer than hunting until almost nothing is admissible.
The fourth moment that was missing
Mehler's formula makes the main effects exact at any correlation and stops there, because the interactions need an expectation of four Hermite functions rather than two. A linearisation turns the four into two, and the whole geometry becomes closed again.
A zero that was an assumption
A rule handed every main effect of both covariates removes exactly none of a pure interaction. That is true at machine precision, it is a fact about independence, and it dies as the square of the correlation.
Walking the admissible set
A rerandomisation test hunts for admissible assignments and throws away the rest. A walk visits them instead — and it is exactly uniform only because it stands still when a proposal fails, which is the step that looks like waste.
Draws that repeat each other
A hunt costs 1/p evaluations per independent draw. A walk costs one per step and yields an effective draw every τ steps. Both are counted in the same unit, and the walk is dearer at every tolerance a trial is designed at.
The weights the corner needs
A fixed-width interval about a difference is exact when the arms share a variance or the allocation ratio is constant, and exact under neither when both fail. It is exact there too, with h_b(λ) = (1/m_A + λ/m_B)⁻¹ — the inverse variances, written as a function of the variance ratio alone, which is a contrast and so is readable by a blinded rule. The estimated precision weights that had no theorem behind them turn out to be that rule at an estimated ratio. And an interval's own scale estimate is right in exactly two cases: inverse-variance weights, and equal ones.
Weights that need only a ratio
A fixed-width interval about a difference is exact under either of two conditions and under neither in the corner. It is exact there too, and the only thing it needs is how much larger one arm's variance is than the other's.
Blinded, and still exact
The one number the exact interval needs is a ratio of within-arm spreads, which is a contrast and contains no mean — so a rule forbidden to look at the effect may compute it, on more degrees of freedom than the interval itself has.
The condition that cannot be dropped
The weights may not read the block they weight. Estimate the variance ratio inside each block rather than across the trial and the coverage falls to 83% — on an interval that is at the same time seventy per cent wider.
Estimating the dependence, not naming it
Whitening a sample repairs a criterion, and the whitening that repairs it is told the dependence is a first-order autoregression and left to find one number. A real dependence has no parameter. The obvious estimate — the sample autocovariances, cut off at some lag — is not a covariance matrix on half the draws there are, so the rule built on it does not exist; the tapered estimate that is always a covariance matrix costs a further seven points of what the repair is worth. And the two constructions named as escaping the resampling's ceiling turn out to be one construction and one identity: a moving block attenuates exactly as a blocked multiplier does, because the attenuation is the join.
A covariance with no parameter in it
The whitening that repairs a criterion is told the dependence is a first-order autoregression and left to find one number. A real dependence is not one number, and the obvious estimate of it is not a covariance matrix.
The window that has to be chosen, and the term that was dropped
An estimated covariance has a bandwidth in it, and both ends of the dial are wrong for different reasons. The rule a practitioner would reach for is two thirds worse than the best window there is.
The triangle that was not the multiplier's
A resampling that leaves each residual on its own row can keep only what the residuals have, times a triangle. A construction that moves every one of them has the same triangle — and the one in this collection's own table has a different taper entirely.
Errors generated from a fitted model
The one construction that is not bounded by the residuals, because a model extrapolates past the lags it was told about and a truncated sample sequence cannot. It is nearly exact where the only defect is dependence, and it pays for it where there are two.
A cut point, at a correlation
The geometry of a balancing dictionary over two dependent covariates is closed for polynomials and was taken to be asymptotic for cut points, because a threshold's Hermite coefficients never terminate. Conditioning on the second variable closes it exactly: every mixed inner product is ρ^j times a one-variable answer and the only two-dimensional object left is an orthant probability, which at the median is (2/π) arcsin ρ. The truncation that was feared falls geometrically in the correlation rather than algebraically in the order — and the interaction guarantee a correlation destroys for powers survives it exactly for median splits, because a two-valued function squares to a constant.
A cut is not a polynomial, and it does not have to be
A threshold's expansion never terminates, which is why a balancing dictionary's geometry was closed for powers and taken to draws for cut points. Conditioning on the second variable closes it for both.
The arcsine that closes it, and the error that was overstated
Two median splits of a correlated pair agree with probability ½ + arcsin(ρ)/π, exactly. And the truncation the field was avoiding falls geometrically in the correlation, not algebraically in the order.
The zero that survives a cut
A rule holding both main effects removes half of a pure interaction between correlated powers and exactly none between correlated median splits. The guarantee that a correlation destroyed was never about interactions.
What a block may vary
Two questions from two corners of the collection with the same answer. A walk over the admissible assignments that exchanges more than one unit per arm mixes faster and is refused more often, and the trade is exactly computable: the gain is a factor of six where a hunt is fifty times cheaper anyway, and nothing at all where the comparison is actually decided. And a variance ratio that drifts between blocks cannot be estimated inside the block it weights — but it can be modelled across them, once the bias in a log variance estimate is subtracted, because that bias depends on the degrees of freedom and the degrees of freedom alternate with the allocation.
A proposal that moves more than two units
The walk's autocorrelation is a fact about its step size and not about its acceptance rate. Exchanging three units from each arm mixes nearly twice as fast as exchanging one, and is refused a third more often.
Stationary is not convergent
A walk that exchanges every unit in each arm preserves the uniform distribution exactly and never gets near it. Every doubly stochastic matrix has the same stationary distribution; only some of them have a limit.
Where the gain is, and where the decision is
A bigger proposal is worth a factor of six at a loose tolerance and nothing at a tight one. The tolerances where it helps are the ones where a hunt costs two evaluations a draw, and the crossing barely moves.