The collection

Every essay — page 13

Essays 289 to 312 of 436, in the same order.

What a block may vary

Two questions from two corners of the collection with the same answer. A walk over the admissible assignments that exchanges more than one unit per arm mixes faster and is refused more often, and the trade is exactly computable: the gain is a factor of six where a hunt is fifty times cheaper anyway, and nothing at all where the comparison is actually decided. And a variance ratio that drifts between blocks cannot be estimated inside the block it weights — but it can be modelled across them, once the bias in a log variance estimate is subtracted, because that bias depends on the degrees of freedom and the degrees of freedom alternate with the allocation.

The shape a dependence has

Estimating a covariance rather than naming it was priced on errors that really were a first-order autoregression, where generality can only cost. Here are four laws with the same first lag and nothing else in common: a moving average that stops, a memory that does not, a break in the middle. The cost of generality where it is not needed and the benefit of it where it is turn out to be the same size — and against a covariance that changes with position rather than with gap, one number, a window and an order are worth exactly the same as each other. Both tuning parameters are then chosen from the sample, and every criterion available for choosing them is aimed at something else.

A block, weighted inside itself

Every resampling in this collection attenuates the dependence it is trying to keep, and until now every one of those attenuations was inherited rather than chosen. The construction the long-run-variance literature actually uses weights the residuals down towards each block's own ends — and what that buys turns out to be exactly the squared value of the window at the two ends and nothing else about its shape. It is an asymptotic gain that has not arrived at any block length a hundred and twenty rows can afford, and at the lengths that are available a tapered block is very nearly a shorter plain one.

What a dictionary buys and what it costs

A balancing rule is a list of functions, and the list decides two things that pull against each other. What it protects against is decided by parity: an interaction between two odd functions is even, an odd function is orthogonal to an even one at every correlation, and the two things every trial balances — a mean and a median split — are both odd, so their worst case is exactly zero however many of them are held. What it costs is the assignments it leaves, and past a certain thinness those stop being one set: the admissible assignments split into an arrangement and its mirror image, no sequence of admissible single swaps joins them, and the walk that samples them is uniform on half the reference distribution for ever.

When a fixed width is reached

A fixed-width interval promises a precision rather than a sample size, so the trial stops when it has enough — and what the stopping rule is allowed to read decides whether the interval means anything. A rule that stops when its own interval is short enough is stopping on the spread it is about to quote, at a correlation of 0.938, and every weighting loses three to five points of coverage including the one told every true variance ratio. The width a set of weights will produce is predictable from the within-arm sums of squares alone, which are independent of every difference the interval is about; a rule that stops on the prediction is back at nominal, costs two and a half blocks, and gives up half of the width promise it was keeping.

The stopping rule costs more than the weighting does. Coverage over 2000 runs of a trial whose variance ratio drifts by a factor of twenty, at three ways of deciding when to stop. Twelve blocks fixed in advance is the top line and reproduces what a trial of fixed length delivers. Stopping when the reported interval is short enough is the bottom line, and it costs between 3.0% and 5.5% of coverage — including for the rule that is told every block's true ratio, which is what says the shortfall belongs to the stopping and not to the weights. Stopping on a width predicted from the within-arm sums of squares is the middle line, and it is back at the fixed-length values. The standard error on each point is 0.49%.

A width the trial has to stop for

The weighting that covers at 94.9% on twelve blocks covers at 91.5% when the trial stops as soon as its interval is short enough — and so does the rule that is told every block's true variance ratio. The shortfall is the stopping, not the weights.

4 figures · Stopping, part 10
Two promises, and no rule here keeps both. A fixed-width procedure promises two things: that the interval covers at its nominal rate, and that it is no wider than the width asked for. Over 1500 runs of the modelled weighting, a rule that stops when the interval it will report is short enough keeps the width — only 2.0% of runs come out wider than 0.34 — and covers at 91.13% against a nominal 95%. A rule that stops on a width predicted from the within-arm sums of squares covers at 94.80% and comes out wider than promised on 42.3% of runs. The two promises are in conflict because keeping the second one exactly requires conditioning on the very quantity that has to be independent of the stopping time for the first.

Stopping on the arms

The width a trial will report is predictable from quantities the interval is not about. A rule that stops on the prediction covers at 94.5% where one that stops on the interval covers at 91.5, and it costs two blocks and half of the width promise.

4 figures · Width, part 8
What a fixed-width interval covers, by the number of blocks the trial ran before it stopped. Two thousand runs of each rule, the modelled weighting, a promise of 0.34. Reading its report: 4–8 blocks, 22.3% of runs, 78.2%; 9–12 blocks, 16.6% of runs, 90.4%; 13–16 blocks, 18.4% of runs, 96.2%; 17–20 blocks, 17.4% of runs, 96.0%; 21–28 blocks, 17.9% of runs, 96.4%; 29–36 blocks, 7.4% of runs, 99.3% — 91.45% overall. Reading the arms: 4–8 blocks, 0.0%, none; 9–12 blocks, 0.9%, 94.4%; 13–16 blocks, 30.4%, 95.6%; 17–20 blocks, 50.0%, 93.9%; 21–28 blocks, 18.0%, 94.4%; 29–36 blocks, 0.7%, 92.3% — 94.50% overall.

The trials that stopped early

A fixed-width trial that stops when its own interval is short enough covers 91.45% — an average of 78.2% among the 22.3% of runs that stop within eight blocks and 96% to 99% among those that run longer. Widening every interval by 17.1% brings the average to 95% and leaves the early stops at 85.6%, while 92.8% of runs now report an interval wider than the width they promised. Even doubling every interval leaves the early stops short.

6 figures · Width, part 9
The fixed-width trial's coverage when the outcomes are not normal, for both stopping rules. normal: stopping on the arms 94.05% after 18.1 blocks, on the report 89.95%; log-normal, skewness 0.95: stopping on the arms 94.70% after 18.5 blocks, on the report 90.80%; log-normal, skewness 2.26: stopping on the arms 94.15% after 19.3 blocks, on the report 90.25%; log-normal, skewness 4.75: stopping on the arms 94.45% after 18.7 blocks, on the report 90.50%; t, five degrees of freedom: stopping on the arms 94.35% after 18.3 blocks, on the report 90.30%; skewness 4.75, arm A only: stopping on the arms 93.80% after 26.0 blocks, on the report 89.90%; skewness 4.75, arm B only: stopping on the arms 93.60% after 14.2 blocks, on the report 89.95%; equal variances, normal: stopping on the arms 94.75% after 11.4 blocks, on the report 90.90%; equal variances, skewness 4.75: stopping on the arms 94.05% after 11.1 blocks, on the report 92.00%.

A width rule on skewed outcomes

The blinded fixed-width rule rests on a within-arm spread being independent of the arm means, which only normal samples guarantee. On outcomes with a skewness of 4.75 the independence fails and the overall coverage barely notices — 93.60% to 94.70% across every shape counted, against 94.05% on normal outcomes. What skew moves is the runs that stop by twelve blocks, which cover about 90% with the skew in one arm, and the trial's length: a variance ratio corrected on normal theory lengthens it from 18.1 blocks to 26.0 with the skew in the first arm and shortens it to 14.2 with the skew in the second.

6 figures · Width, part 10

Fitted together, or fitted after

Every whitening in this collection is a two-step rule: fit a candidate, read the dependence off what is left over, whiten, fit again. The step nobody examined is the first one. A fit removes memory as well as signal, and how much is arithmetic rather than noise — the residuals of a hundred and twenty rows report a lag-one coefficient of 0.73 where the errors report 0.78 and the law says 0.80. Estimating the coefficient with the line instead of after it recovers most of that, and iterating the two-step rule to its fixed point recovers nearly as much — because a fixed point of the sum of squares is not a maximum of the likelihood, and the difference between them is real, one-sided, and worth nothing at all in the decision the number feeds. The order the same criterion picks moves by nearly two, which is worth a great deal more.

Paying for a search

A two-regime whitening chooses its change point by searching a profile, and then reads a criterion that counts parameters. Nothing charges for the search. Nothing can, in the usual way: under no break the two regimes share a coefficient and the break point is not identified at all, so the likelihood ratio is a supremum over a parameter that exists only under the alternative and no count of restrictions describes its distribution. Searching a hundred and twenty rows for a change point that is not there manufactures about five units of ratio where one parameter costs two — and a rule that counts the break as free splits a stationary sample on 99% of draws. A second break costs as much again, on a sample that has only ever had one.

FieldsThreadsSeriesConceptsFigure librarySearch