Every essay — page 13
What a block may vary
Two questions from two corners of the collection with the same answer. A walk over the admissible assignments that exchanges more than one unit per arm mixes faster and is refused more often, and the trade is exactly computable: the gain is a factor of six where a hunt is fifty times cheaper anyway, and nothing at all where the comparison is actually decided. And a variance ratio that drifts between blocks cannot be estimated inside the block it weights — but it can be modelled across them, once the bias in a log variance estimate is subtracted, because that bias depends on the degrees of freedom and the degrees of freedom alternate with the allocation.
A ratio that changes between blocks
A wrong weight costs width and a random weight costs level. The rule aimed at the quantity that actually varies is the only one that misses its own coverage, and the rule that models it across blocks recovers the whole of what knowing it is worth.
The bias that lands in the slope
The bias in a log variance estimate depends on nothing but its degrees of freedom, so it goes into the intercept — unless the degrees of freedom alternate with the design, which is exactly what a block-randomised trial makes them do.
The shape a dependence has
Estimating a covariance rather than naming it was priced on errors that really were a first-order autoregression, where generality can only cost. Here are four laws with the same first lag and nothing else in common: a moving average that stops, a memory that does not, a break in the middle. The cost of generality where it is not needed and the benefit of it where it is turn out to be the same size — and against a covariance that changes with position rather than with gap, one number, a window and an order are worth exactly the same as each other. Both tuning parameters are then chosen from the sample, and every criterion available for choosing them is aimed at something else.
A dependence with a shape
Four ways for errors to repeat, all with the same first lag and nothing else in common. A rule told the errors are a first-order autoregression finds the same number in all four, and is right about one of them.
Where the generality runs out
A covariance that changes half way through a sample is not one a window can estimate. One number, a window and an order are worth the same as each other on it — and letting the model change once, at a point nobody can locate, is worth as much again as all three.
The window a whitening wants
Every law here is best whitened by a window several times longer than its own memory, including the one whose memory ends at the fourth lag. The three ways of choosing it from the sample all land in the same place, and it is the wrong one.
The order the tail is drawn at
A fitted autoregression reproduces the sample exactly at the lags it was fitted on, so everything it says past them is extrapolation — and the order is the dial that decides how much of it there is.
A block, weighted inside itself
Every resampling in this collection attenuates the dependence it is trying to keep, and until now every one of those attenuations was inherited rather than chosen. The construction the long-run-variance literature actually uses weights the residuals down towards each block's own ends — and what that buys turns out to be exactly the squared value of the window at the two ends and nothing else about its shape. It is an asymptotic gain that has not arrived at any block length a hundred and twenty rows can afford, and at the lengths that are available a tapered block is very nearly a shorter plain one.
A block weighted inside itself
The triangle every block resample attenuates by is not a fact about blocks. It is the self-convolution of a rectangle, and a block weighted down towards its own ends has a different one — whose leading term is the squared value at the two ends and nothing else about the shape.
A taper and a critical value
Two constructions whose tapers visibly differ give the same critical value, and two that share a taper exactly do not. Adding a construction whose taper is a decision rather than an accident says which half of that is true.
What a dictionary buys and what it costs
A balancing rule is a list of functions, and the list decides two things that pull against each other. What it protects against is decided by parity: an interaction between two odd functions is even, an odd function is orthogonal to an even one at every correlation, and the two things every trial balances — a mean and a median split — are both odd, so their worst case is exactly zero however many of them are held. What it costs is the assignments it leaves, and past a certain thinness those stop being one set: the admissible assignments split into an arrangement and its mirror image, no sequence of admissible single swaps joins them, and the walk that samples them is uniform on half the reference distribution for ever.
A dictionary that is neither
A rule handed two median splits removes none of their interaction; a rule handed two covariates removes none of their product. Those were two results with two explanations, and they are one result with one — and finding it corrected the number underneath both.
What the extra function buys
A rule balancing the mean of each covariate has a worst case of exactly zero. Adding the median split — the other thing every trial balances — leaves it at exactly zero, and one square moves it.
The set a dictionary leaves
A rule constrained on six functions at a loose tolerance leaves a set as thin as one constrained on three at a tight one. Both sampling methods cross over at the same thinness, and the tolerance where that happens moves by a factor of three.
The walk that cannot cross
A thin enough admissible set is not one set. It splits into an assignment and its mirror image, no sequence of admissible single swaps joins them, and the walk that samples it is uniform on half the reference distribution for ever.
When a fixed width is reached
A fixed-width interval promises a precision rather than a sample size, so the trial stops when it has enough — and what the stopping rule is allowed to read decides whether the interval means anything. A rule that stops when its own interval is short enough is stopping on the spread it is about to quote, at a correlation of 0.938, and every weighting loses three to five points of coverage including the one told every true variance ratio. The width a set of weights will produce is predictable from the within-arm sums of squares alone, which are independent of every difference the interval is about; a rule that stops on the prediction is back at nominal, costs two and a half blocks, and gives up half of the width promise it was keeping.
A width the trial has to stop for
The weighting that covers at 94.9% on twelve blocks covers at 91.5% when the trial stops as soon as its interval is short enough — and so does the rule that is told every block's true variance ratio. The shortfall is the stopping, not the weights.
Stopping on the arms
The width a trial will report is predictable from quantities the interval is not about. A rule that stops on the prediction covers at 94.5% where one that stops on the interval covers at 91.5, and it costs two blocks and half of the width promise.
The trials that stopped early
A fixed-width trial that stops when its own interval is short enough covers 91.45% — an average of 78.2% among the 22.3% of runs that stop within eight blocks and 96% to 99% among those that run longer. Widening every interval by 17.1% brings the average to 95% and leaves the early stops at 85.6%, while 92.8% of runs now report an interval wider than the width they promised. Even doubling every interval leaves the early stops short.
A width rule on skewed outcomes
The blinded fixed-width rule rests on a within-arm spread being independent of the arm means, which only normal samples guarantee. On outcomes with a skewness of 4.75 the independence fails and the overall coverage barely notices — 93.60% to 94.70% across every shape counted, against 94.05% on normal outcomes. What skew moves is the runs that stop by twelve blocks, which cover about 90% with the skew in one arm, and the trial's length: a variance ratio corrected on normal theory lengthens it from 18.1 blocks to 26.0 with the skew in the first arm and shortens it to 14.2 with the skew in the second.
Fitted together, or fitted after
Every whitening in this collection is a two-step rule: fit a candidate, read the dependence off what is left over, whiten, fit again. The step nobody examined is the first one. A fit removes memory as well as signal, and how much is arithmetic rather than noise — the residuals of a hundred and twenty rows report a lag-one coefficient of 0.73 where the errors report 0.78 and the law says 0.80. Estimating the coefficient with the line instead of after it recovers most of that, and iterating the two-step rule to its fixed point recovers nearly as much — because a fixed point of the sum of squares is not a maximum of the likelihood, and the difference between them is real, one-sided, and worth nothing at all in the decision the number feeds. The order the same criterion picks moves by nearly two, which is worth a great deal more.
A dependence fitted with the line
Every whitening in this collection reads the dependence off a set of residuals, and residuals are not errors. Fitting the two together recovers most of what that costs, and changes almost nothing about the decision it feeds.
The fit that takes the memory out
A candidate's residuals report less dependence than its errors do, and how much less is arithmetic rather than noise. The rule used for a good reason reads the series that has lost the most.
Iterating is not maximising
Re-reading a correlation from the generalised residuals and refitting converges in seven steps. What it converges to solves the first-order condition of a sum of squares, and the likelihood has one term more than that.
A window for every candidate
The window and the order a whitening needs are chosen once, from the fullest candidate, on an argument that was made about an estimated covariance. A tuning parameter is not a covariance, and the two cost different amounts.
Paying for a search
A two-regime whitening chooses its change point by searching a profile, and then reads a criterion that counts parameters. Nothing charges for the search. Nothing can, in the usual way: under no break the two regimes share a coefficient and the break point is not identified at all, so the likelihood ratio is a supremum over a parameter that exists only under the alternative and no count of restrictions describes its distribution. Searching a hundred and twenty rows for a change point that is not there manufactures about five units of ratio where one parameter costs two — and a rule that counts the break as free splits a stationary sample on 99% of draws. A second break costs as much again, on a sample that has only ever had one.
A break that was looked for
A two-regime whitening finds its change point by maximising a profile, and then reads a criterion that counts parameters. Under no break there is no parameter to count, because every position describes the same model.
What a search costs in parameters
An information criterion's penalty is an estimate of the optimism a fit carries. For a break point the optimism can be measured and cannot be counted, and it comes to about two and a half parameters.
Choosing whether to break
Charging what the search manufactures takes a rule from splitting a stationary sample on 99% of draws to 16%. It also costs regret, because the two mistakes a rule can make are not the same size.
A second break on a flat profile
Searching a hundred and twenty rows for one change point manufactures five units of likelihood. Searching for a second manufactures four more, on a series that has at most one — and on a profile whose whole range is under seven.