When a fixed width is reached

A width rule on skewed outcomes

The blinded fixed-width rule rests on a within-arm spread being independent of the arm means, which only normal samples guarantee. On outcomes with a skewness of 4.75 the independence fails and the overall coverage barely notices — 93.60% to 94.70% across every shape counted, against 94.05% on normal outcomes. What skew moves is the runs that stop by twelve blocks, which cover about 90% with the skew in one arm, and the trial's length: a variance ratio corrected on normal theory lengthens it from 18.1 blocks to 26.0 with the skew in the first arm and shortens it to 14.2 with the skew in the second.

Worth reading first: The shortest interval is the one that misses · When the looking happens.

Stopping on the arms repaired a fixed-width trial by changing what its stopping rule reads. Instead of the interval it is about to report, the rule reads a width predicted from the first arm’s within-arm sums of squares, and under normality those sums of squares are independent of every arm mean — so the stopping time carries no information about the difference being estimated, and a trial that stops early is not a trial that got lucky. The blinded rule covers 94.50% where the rule that reads its report covers 91.45%.

That essay was explicit about the boundary. The independence “is a fact about normality, not about large samples. On skewed data the within-arm sums of squares and the arm means are correlated, the stopping time would carry information about the differences again, and the shortfall would return in some amount nothing here measures.” This essay measures it, on the same design: a variance ratio that drifts across thirty-six planned blocks, a half-width promised for a difference of 0.34, two thousand runs a row, with the outcomes drawn from log-normal distributions of increasing skewness and from a symmetric heavy-tailed t.

The fixed-width trial's coverage when the outcomes are not normal, for both stopping rules. normal: stopping on the arms 94.05% after 18.1 blocks, on the report 89.95%; log-normal, skewness 0.95: stopping on the arms 94.70% after 18.5 blocks, on the report 90.80%; log-normal, skewness 2.26: stopping on the arms 94.15% after 19.3 blocks, on the report 90.25%; log-normal, skewness 4.75: stopping on the arms 94.45% after 18.7 blocks, on the report 90.50%; t, five degrees of freedom: stopping on the arms 94.35% after 18.3 blocks, on the report 90.30%; skewness 4.75, arm A only: stopping on the arms 93.80% after 26.0 blocks, on the report 89.90%; skewness 4.75, arm B only: stopping on the arms 93.60% after 14.2 blocks, on the report 89.95%; equal variances, normal: stopping on the arms 94.75% after 11.4 blocks, on the report 90.90%; equal variances, skewness 4.75: stopping on the arms 94.05% after 11.1 blocks, on the report 92.00%.
Fig. 1 The coverage of the stopped interval for the rule that stops on the arms and the rule that stops on its report, for normal, log-normal and t outcomes, skew in one arm only, and a design with equal variances in the two arms.

What skew does to a sample

How strongly a skewed sample's mean and variance move together, against the skewness. Closed form γ₁ / √(γ₂ + 2n/(n − 1)) for a log-normal outcome. Counted over forty thousand samples of ten: skewness 0.95, 0.484 counted against 0.483; skewness 2.26, 0.651 counted against 0.639; skewness 4.75, 0.674 counted against 0.615. For a large sample: 0.497 at 0.95, 0.645 at 2.26, 0.616 at 4.75.
Fig. 2 The correlation between a log-normal sample’s mean and its variance, against the outcome’s skewness, for samples of three, of ten and for a large sample, with counted values for samples of ten.

For any outcome, the covariance between a sample’s mean and its variance is the third central moment divided by the sample size, and the correlation between them is the skewness divided by the square root of the excess kurtosis plus 2n/(n − 1). For a normal outcome both are zero and, more than that, the mean and the variance are independent. For a skewed one the correlation is not small, and it does not shrink as the sample grows: for large samples it is 0.497 at a skewness of 0.95, 0.645 at 2.26 and 0.616 at 4.75 — lower at the largest skewness because the excess kurtosis grows faster than the skewness does.

The closed form is checked by counting. Over forty thousand samples of ten, the correlation is 0.484 against 0.483 at a skewness of 0.95, and 0.651 against 0.639 at 2.26. At 4.75 the count is 0.674 against 0.615, and the gap is the tail: a log-normal of that skewness has an excess kurtosis above fifty, and a correlation estimated from forty thousand draws of it has not settled. The direction is not in doubt at any skewness.

So the premise the blinded rule rests on is simply false for skewed outcomes. A block whose arm came out with a small spread is, more often than not, a block whose arm came out with a small mean.

The stopping time leans on the error

How far the blinded rule's stopping time leans on the estimate's error, by the outcome's shape. normal: -0.009; log-normal, skewness 0.95: -0.070; log-normal, skewness 2.26: -0.057; log-normal, skewness 4.75: -0.025; t, five degrees of freedom: -0.003; skewness 4.75, arm A only: 0.127; skewness 4.75, arm B only: -0.402; equal variances, normal: -0.009; equal variances, skewness 4.75: 0.248.
Fig. 3 The correlation between the number of blocks a trial enrolls before stopping on the arms and the error of its estimated difference, by the shape of the outcome.

With normal outcomes the blinded rule’s stopping time and its estimate’s error are uncorrelated, −0.009. With skewness in both arms they stay nearly so: −0.070 at a skewness of 0.95, −0.057 at 2.26, −0.025 at 4.75. The symmetric heavy-tailed t, whose mean and variance are uncorrelated without being independent, gives −0.003.

That near-zero is not the independence surviving. It is two effects cancelling. Put the skew in the first arm alone and the correlation is +0.127: the rule reads that arm’s spread, a run whose first arm came out spread out runs longer, and a first arm that came out spread out came out high, so long runs carry positive errors. Put the skew in the second arm alone and the correlation is −0.402. The rule does not read that arm’s spread directly, but the weights read it, through the estimated variance ratio in each block, and a second arm that came out spread out came out high, lowering the difference and lengthening the run by lowering the weight the block contributes.

With both arms skewed in this design the two leans run in opposite directions and nearly cancel. The cancellation belongs to this design, not to skewed outcomes. In a design with equal variances in the two arms, skew in both arms gives a correlation of +0.248, and the arm-A effect wins.

Two thousand runs stopping on the arms, skewness 4.75, arm B only: the error of each estimate against when it stopped. The correlation between blocks and error is -0.402. The line is the mean error at each length reached by at least 25 runs: 0.244 at 9, 0.116 at 10, 0.101 at 11, 0.076 at 12, 0.080 at 13, 0.027 at 14, -0.002 at 15, -0.045 at 16, -0.039 at 17, -0.097 at 18, -0.105 at 19. Coverage 93.60% overall; by band of lengths 9 to 12 blocks, 494 runs, 90.3%; 13 to 16 blocks, 1145 runs, 94.8%; 17 to 20 blocks, 317 runs, 94.6%; 21 to 28 blocks, 36 runs, 88.9%. Errors beyond ±1.2 are drawn at the edge.
Fig. 4 Two thousand runs stopping on the arms with skew in the second arm only: the error of each estimate against the number of blocks enrolled, with the mean error at each length and, along the top, the coverage among the runs that stopped in each band of lengths.

Drawn run by run, the lean is not subtle. With skew in the second arm alone, the runs that stop at nine blocks have a mean error of 0.244, those at twelve 0.076, at fifteen −0.002 and at nineteen −0.105. A reader holding one of those trials, knowing when it stopped, knows which way its estimate is likely to be off — which is exactly the information the blinded rule was built to deny the stopping time.

The mechanism is visible directly in the stopping quantity. With normal outcomes, the first arm’s pooled spread at the stop correlates with the number of blocks at 0.528 and with the estimate’s error at −0.004. With skewness 4.75 in both arms, it correlates with the number of blocks at 0.771 and with the error at 0.311. The rule reads a quantity that now knows something about the answer.

And the coverage does not follow

With every shape counted, the blinded rule’s coverage stays where it was. Normal outcomes: 94.05%. Log-normal, skewness 0.95: 94.70%; 2.26: 94.15%; 4.75: 94.45%. The t: 94.35%. Skew in the first arm alone: 93.80%; in the second alone: 93.60%. Equal variances: 94.75% normal and 94.05% skewed. Each is counted over two thousand runs, with a standard error of about half a point. On the drifting design the largest gap from the normal-outcome value is 0.65 points, and on the equal-variance design it is 0.70.

The rule that reads its own report stays below the blinded rule in every case, between 89.90% and 92.00%, which is where it was on normal outcomes and near the 90% that an interval read after the stop it chose covers for the same reason. Skew neither rescues it nor sinks it.

So the prediction was half right. The stopping time does carry information about the differences again — the correlations say so, and the run-by-run picture shows it — and the shortfall does not return to the overall count. The earlier claim that the argument “stops where normality does” was the right thing to say without a measurement, and the measurement is more forgiving than the argument on average and less forgiving in two places the average does not show.

When a lean becomes a shortfall

A coverage count asks one question of each run: is the estimate within the reported half-width of the truth? The report-reading rule failed that question because it stopped on the half-width itself, so a run that stopped early had a half-width selected for being small, and the error was measured against an understated ruler.

The blinded rule under skew does something different. Its stopping time moves with the sign of the error — which way the estimate is off — not with the error’s size relative to the half-width. The half-width is still computed from the spread of the block differences, which the stopping rule never read, so a run that stopped early has an honest ruler and an estimate shifted a little in one direction. Averaged over every run, that is nearly free. With the skew in the second arm alone, 57.3% of runs stop at thirteen to sixteen blocks, where no length’s mean error is larger than 0.080 — under a quarter of the half-width — and only 1.3% stop at nine blocks, where the mean error is 0.244.

For the runs in which the shift is largest it is not free. An honest ruler laid around a shifted estimate misses on the side the estimate moved towards more often than it misses on the other, and the two do not balance. Among the 502 runs that stop within twelve blocks with the skew in the second arm, the interval covers 90.4%, and 44 of its 48 misses lie above the truth. With the skew in the first arm the lean reverses and so do the misses: the 139 runs that stop within twelve blocks cover 90.6%, and all 13 of their misses lie below. The runs that go on to thirteen blocks or more cover 94.66% and 94.04%. The first-arm count rests on few runs, with a standard error of two and a half points, so its direction is clearer than its size; the second-arm count rests on a quarter of all runs and sits more than three standard errors below 95%.

That is the conditional shortfall the trials that stopped early found in the rule that reads its own report, back at a smaller size and by another route. There the early stops covered 78.2% within eight blocks because their ruler was short. Here they cover about 90% within twelve because their estimate has moved and the ruler has not. The overall count hides it for the same reason in both: the later runs cover near nominal and outnumber the early ones. With both arms skewed the two leans cancel for the early stops as well — the 551 runs that stop within twelve blocks cover 94.7%, with 17 misses above the truth and 12 below.

So the one-arm cases hold the two lowest overall counts, 93.80% and 93.60%, and they are the two whose loss has a location. The average puts that loss on the edge of what two thousand runs can see; split by stopping time it is plain.

A weighting biased with no stopping at all

The estimate's bias at a fixed eighteen blocks, by weighting, with normal and skewed outcomes. oracle, normal: 0.0037 ± 0.0043; oracle, skewness 4.75: -0.0051 ± 0.0042; modelled, normal: 0.0031 ± 0.0043; modelled, skewness 4.75: 0.0255 ± 0.0043; pooled, normal: 0.0010 ± 0.0051; pooled, skewness 4.75: -0.0067 ± 0.0051; flat, normal: 0.0009 ± 0.0052; flat, skewness 4.75: -0.0077 ± 0.0053.
Fig. 5 The mean error of the estimated difference at a fixed eighteen blocks — no stopping rule at all — for four weightings, with normal outcomes and a skewness of 4.75, with two standard errors.

Skew leaves one mark that has nothing to do with stopping. Run the same design for exactly eighteen blocks, with no stopping rule, and read the mean error of each weighting’s estimate. On normal outcomes all four sit at zero within their standard errors. With a skewness of 4.75, three still do — the weighting told every true ratio at −0.0051 ± 0.0042, the single pooled ratio at −0.0067 ± 0.0051, equal weights at −0.0077 ± 0.0053 — and the modelled weighting does not: 0.0255 ± 0.0043, six standard errors from zero, against 0.0031 on normal outcomes.

The modelled weighting fits a curve to the log of each block’s variance ratio and weights each block by the fitted ratio. That makes it the one rule whose weights respond to each block’s own spreads block by block, and under skew a block’s spreads move with its means, so the weights lean towards blocks whose means came out a particular way. The pooled ratio averages the spreads over the whole trial before any block is weighted, and the oracle and equal weights read no spreads at all. That is the explanation that fits the four rows; the rows are what is measured.

The bias is small beside the half-width — 0.0255 against 0.34 — and the modelled weighting still covers 94.40% at eighteen blocks on skewed outcomes. But it is the one effect here that no stopping rule causes and none could remove, and it belongs to the weighting a ratio that changes between blocks recommended as the best.

The length a normal-theory correction sets

The half-width the blinded rule predicts on skewed outcomes, beside the one the estimate's spread calls for. normal: predicted 0.403, called for 0.408, modelled ratio ×0.998 the truth, 18.13 blocks when left to stop, 0.0% of runs at the cap; log-normal, skewness 0.95: predicted 0.407, called for 0.408, modelled ratio ×1.038 the truth, 18.55 blocks when left to stop, 0.1% of runs at the cap; log-normal, skewness 2.26: predicted 0.417, called for 0.408, modelled ratio ×1.163 the truth, 19.33 blocks when left to stop, 3.5% of runs at the cap; log-normal, skewness 4.75: predicted 0.427, called for 0.408, modelled ratio ×1.386 the truth, 18.69 blocks when left to stop, 10.1% of runs at the cap; t, five degrees of freedom: predicted 0.406, called for 0.414, modelled ratio ×1.059 the truth, 18.30 blocks when left to stop, 0.9% of runs at the cap; skewness 4.75, arm A only: predicted 0.495, called for 0.417, modelled ratio ×2.177 the truth, 25.96 blocks when left to stop, 33.4% of runs at the cap; skewness 4.75, arm B only: predicted 0.356, called for 0.403, modelled ratio ×0.636 the truth, 14.19 blocks when left to stop, 0.0% of runs at the cap.
Fig. 6 At a fixed eighteen blocks, the half-width the blinded rule predicts beside the half-width the estimate’s spread across runs calls for, by the outcome’s shape; at the right, the modelled variance ratio as a multiple of the true one and the number of blocks the rule enrols when it is left to stop.

The same weighting mis-states something else under skew, and this one decides when the blinded rule stops. The modelled weighting does not fit its curve to the raw log variance ratios. The logarithm of a sample variance sits below the logarithm of the true variance on average, and for a normal sample on k degrees of freedom the gap is a known constant, ψ(k/2) + log 2 − log k: −0.577 for a sample of three. The weighting subtracts that constant from each block’s log ratio before it fits — the repair a bias that lands in the slope calls for when the degrees of freedom alternate with the blocks — and on normal outcomes the fitted ratio at eighteen blocks comes out ×0.998 the true one.

The constant is a fact about normal samples. A log-normal sample of three with a skewness of 4.75 has an average log variance of −1.569, a whole unit lower than the correction allows for; a sample of ten sits 0.575 below it and a sample of twenty-seven 0.323 below. The correction removes the normal part of the gap and leaves the rest in the ratio.

Which way the rest goes depends on the arm, because the ratio is the second arm’s variance over the first’s, and the arithmetic can be done by hand. The lopsided blocks put three observations in the first arm and twenty-seven in the second, and the balanced blocks ten in each. With the skew in the first arm, the denominator is understated by about 0.99 in half the blocks and 0.575 in the other half, and the fitted ratio comes out ×2.177 the truth — a shift of 0.78 on the log scale. With the skew in the second arm the numerator is understated by 0.323 and 0.575, and the ratio comes out ×0.636. With both arms skewed the two partly cancel, ×1.386. At the two lower skewnesses it is ×1.038 and ×1.163, and on the t, ×1.059.

A ratio set too high gives each block too little weight, because the weighting believes the second arm noisier than it is, and the blinded rule predicts its width from those weights. At a fixed eighteen blocks with the skew in the first arm it predicts a half-width of 0.495 where the estimate’s spread across runs calls for 0.417; with the skew in the second arm, 0.356 against 0.403; on normal outcomes, 0.403 against 0.408.

So the rule stops at the wrong time, in the direction the ratio errs. With the skew in the first arm it enrols 25.96 blocks on average against 18.13, 33.4% of runs hit the cap at thirty-six, and the reported half-width averages 0.295 — narrower than the 0.34 promised, bought with nearly eight more blocks. With the skew in the second arm it enrols 14.19 blocks and reports 0.363 on average, and 56.7% of runs report an interval wider than promised, against 42.9% on normal outcomes.

This is a failure of the width promise and of the budget, and it leaves coverage alone for the reason the lean did: the reported half-width is computed from the block differences, never from the prediction, and the prediction decides only when to stop. It is also a failure the independence argument was never about. A blinded rule can be exactly independent of the effect and still be wrong about the width, whenever the quantity it predicts the width from is estimated with a bias. On normal outcomes that bias had been corrected; on skewed ones the correction is what is biased.

The independence was never the claim

It is worth separating what the blinded rule needs from what it was argued from. The argument was that the stopping time is independent of the differences, and that is enough for exact coverage. The measurement says it is more than enough. What coverage needs is weaker: that the stopping time not select the half-width. Under skew the blinded rule’s stopping time selects, a little, the direction of the error, and leaves the ruler alone — and a selected direction costs coverage only where the selection is strong, which is among the runs that stop earliest.

That is the same distinction blinded and still exact drew for re-estimating a variance ratio without seeing the effect, one level down: a rule is safe to the extent that what it reads cannot change the quantity it is judged on, and independence is the strongest way to guarantee that rather than the only one. Skew breaks the strongest guarantee and leaves a weaker one standing, and on this design the weaker one is the one the overall coverage turns out to need.

What a fixed-width protocol on skewed outcomes should say

That the independence is approximate, and by how much. The correlation between a sample’s mean and variance is computable from the outcome’s skewness and kurtosis before any data, and for a skewness of two it is above 0.6. A protocol that cites the blinded rule’s exactness should cite its premise too.

Which weighting is used, and whether it reads block-level spreads. The stopping rule survived skew; the modelled weighting’s estimate did not. A weighting that pools spreads across blocks before weighting kept its estimate unbiased on every shape counted.

The stopping length, as always. On skewed outcomes a trial that stopped within twelve blocks has an estimate that leans, and with the skew in one arm its interval covers about 90% rather than 95%. A reader combining it with other trials should know that trials stopping early and late are off in opposite directions — which averages out across trials and does not average out within one.

The half-width predicted at the stop beside the one reported. On normal outcomes the two agree to within what the degrees of freedom explain. A reported interval persistently wider than the prediction that stopped the trial, or narrower, is the visible sign that the variance ratio behind the prediction is off, and on skewed outcomes it is off towards the skewed arm.

What is counted here and what is closed

Closed. The correlation of a sample’s mean with its variance for a log-normal outcome, from the skewness and kurtosis; the skewness of each shape from its parameter; and the normal-theory bias of a log sample variance that the modelled weighting subtracts.

Counted. Every coverage, correlation and bias: two thousand runs a row, the modelled weighting unless named, a promise of 0.34 and a cap of thirty-six blocks; forty thousand samples for the counted correlations; a hundred thousand samples for each skewed log variance, and three thousand runs of eighteen blocks for each ratio multiple.

Particular to this design. Log-normal skew in one or both arms, a t on five degrees of freedom, the drifting design and one with equal variances. Discrete outcomes, skew of opposite signs in the two arms, and block sizes smaller than three are not measured, and the cancellation between the two arms’ leans is a fact about this design’s variance ratio rather than a general one.

Still open: a weighting that does not read the spreads it weights by, and a correction that is not normal theory

The one real damage skew did was to the modelled weighting’s estimate, through weights fitted to block-level spreads that move with block-level means. There is an obvious repair and it has a cost: fit the variance-ratio curve to spreads from which each block’s own contribution has been left out, so a block’s weight cannot depend on that block’s mean. Whether that removes the 0.0255, and what it costs in the precision the modelled weighting was chosen for, is a measurement this design can make and has not made.

The length has a repair of its own, and a harder one. The log-variance correction could be computed for the shape the outcomes actually have rather than for a normal one, using a kurtosis estimated from every block’s within-arm deviations pooled together, since a sample of three cannot estimate a kurtosis and thirty-six blocks of them can. Whether a pooled kurtosis settles fast enough to help a rule that may stop at nine blocks, and whether a corrected ratio then brings the skewed cases back to eighteen blocks without disturbing the early stops’ coverage, are both open. The simpler alternative is already on the page — the pooled ratio averages sums of squares before taking any logarithm and so has no log gap to correct — and its cost is the drift it cannot follow.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

BiasBlindingClosed formCoverageEstimated varianceFixed-width intervalIndependenceInverse variance weightingMonte CarloSkewnessStopping ruleVariance ratio