The bias that lands in the slope
Worth reading first: The shortest interval is the one that misses · What the exactness buys.
The rule that models a drifting variance ratio across blocks works, and the argument for it is one sentence: λ̂_b is λ_b times an F variate, so log λ̂_b is log λ_b plus a variate whose distribution depends on nothing but the degrees of freedom — and a bias that is the same in every block goes into the intercept and leaves the slope alone.
The first half of that sentence is exactly true. The second half is false in this design, and the reason it is false is the design’s most ordinary feature.
The bias is a digamma
For a variance estimate on k degrees of freedom,
which is negative for every k: the log of an unbiased variance estimate is biased low. At two degrees of freedom it is −0.5772, which is Euler’s constant arriving because ψ(1) = −γ; at nine it is −0.1152; at twenty-six, −0.0390.
The bias in log λ̂_b is the difference of two of these, one for each arm. It contains no unknown parameter, no variance, and nothing about the data: it is arithmetic on two integers, which is what makes it correctable exactly rather than estimable.
And the degrees of freedom alternate
A block-randomised trial in the corner this field is about does not have the same allocation in every block. In the design here the even blocks are twenty units split evenly — nine degrees of freedom in each arm, so the two biases cancel and the total is zero — and the odd blocks are thirty units split three to twenty-seven, giving two and twenty-six, so the total is
Half the blocks carry a bias of about half a unit on the log scale and half carry none. That would still be harmless if the two kinds were arranged symmetrically with respect to the covariate being fitted. They are not, and they cannot be: the blocks alternate, and any covariate that runs monotonically across the trial has the odd blocks sitting systematically on one side of its mean.
With twelve blocks and a covariate running from −1 to 1, the even blocks average −0.0909 and the odd ones +0.0909. The bias is correlated with the regressor, and a bias correlated with a regressor is a slope.
The arithmetic is exact enough to predict the damage before measuring it: the induced slope is 0.5383 × cov(odd, u)/var(u) = 0.5383 × 0.04545/0.3939 = 0.0621.
Which is what the fit does
The standard error of the fitted slope is 0.0074, so the uncorrected estimate is 8.0 standard errors above the truth and the corrected one is 0.3 below it. The predicted displacement of 0.0621 against the measured 0.0597 agrees to within the noise, which is the check that the digamma identity is the explanation rather than a plausible one.
The correction is one subtraction per block, applied before the regression, and it needs nothing that is not already known when the design is written down. It is also, unusually for anything in this collection, exactly free: there is no variance inflation, no degree of freedom spent, and no tuning constant — the two digammas are determined by the allocation, and the allocation is a design choice made before any unit is enrolled.
The factor is 3/26, and it says what to do about it
The 0.11538 in the prediction is the regression of the odd-block indicator on the covariate, and on an evenly spaced covariate it has a closed form worth having, because the closed form is the repair.
With B blocks alternating and the covariate running evenly from −1 to 1, the covariance between the indicator and the covariate is and the covariate’s variance is , so
At twelve blocks that is exactly 3/26 = 0.11538, and the induced slope is 0.5383 × 3/26 = 0.0621 — the number the section above measures, with the two integers it is made of now visible.
Two things follow that the single measurement does not show.
The damage falls like 1/B and does not vanish. Twenty-four blocks give 3/50 and a slope of 0.0323; a hundred give 3/202 and 0.0080. More blocks dilute the alignment between the allocation pattern and the covariate, because the odd blocks’ average covariate value moves towards the even blocks’, but they never break it — alternation guarantees that the odd blocks sit on one side, by exactly one spacing of the grid, at every B.
The alignment is the fixable half. The digamma bias is a fact about degrees of freedom and cannot be argued with; what turns it into a slope is that the two kinds of block are arranged in a pattern correlated with the regressor. Assign which blocks get the 3:27 allocation at random rather than by parity and the covariance has expectation zero: the bias is still there, still 0.5383 in half the blocks, and it goes into the intercept and the residual variance instead of into the slope.
That is worth stating as the choice it is. Either correct the bias exactly — it is arithmetic on two integers and needs nothing estimated — or make the design orthogonal to it. The one thing that does not work is the reasoning the field started from, which was that a bias depending on nothing but the degrees of freedom is safe. It is safe when the degrees of freedom are the same everywhere. Here they alternate, and alternating is a pattern, and a pattern is a regressor.
Why the log scale at all
There is a prior question the correction makes it natural to ask: why fit a model to log λ̂ rather than to λ̂ itself, given that the log is what introduces the bias in the first place.
Two reasons, and the second is the one that decides it.
A variance ratio is positive and multiplicative. A drift that doubles the ratio over the trial is the same kind of drift whether it starts at 2 or at 20, and only the log scale makes that a straight line. A model fitted to λ̂ on this design would be fitting a line through an exponential, which is a misspecification introduced by the choice of scale rather than by the world.
And the bias on the log scale is a constant per allocation, which is exactly what makes it correctable. On the raw scale E[λ̂_b] is λ_b times a factor that depends on the degrees of freedom — so the bias scales with the ratio, is therefore not a constant, and does not go into an intercept under any arrangement of the blocks. A line fitted to the ratios themselves reports a drift of 1.15 where the truth is 1.5, and no subtraction fixes it, because what is needed is a division and the model is additive.
So the log scale introduces a bias that can be removed exactly and the raw scale introduces one that cannot. That is the trade, and it is not close.
What it costs to get wrong, which is less than it should be
Here the honest reporting is that the refusal barely bites.
An uncorrected model covers at 94.40% against the corrected model’s 94.85% and a nominal 95% with a standard error of 0.345 points — one and a third standard errors low — and its interval is 1.9% narrower. So a rule with a slope eight standard errors out produces an interval whose coverage is a third of a point low.
The reason is worth having, because it is the same reason the pooled rule loses no coverage at all. A weighting is invariant to a common factor, so what a wrong slope does is tilt the weights slightly relative to each other rather than mis-state their level; and the calibration identity means a set of weights that is wrong but fixed still produces an interval measuring the variance of the estimator it computed. The uncorrected model over-weights the late blocks and under-weights the early ones by a few per cent, which is a mild efficiency loss and not a level failure.
So this is a refusal in the shape of the one about shopping over the variance ratio — a thing that is definitely wrong, priced, and found to be worth less than the reasoning would suggest. Recording that is as much of the point as recording the ones that break something: an estimator eight standard errors from the truth costing a third of a coverage point is a fact about how forgiving this construction is, and a reader who takes away only correct the bias has missed it.
The reason to correct it anyway is that the correction is free and its absence is not visible. Nothing in the output of an uncorrected fit says the slope is displaced; the residuals look fine, the fit looks fine, and the displacement is a deterministic function of the allocation pattern.
The same trap, in three other places
The shape is general enough to be worth naming, because the sentence that fails here is a good sentence and it fails wherever a nuisance’s precision is correlated with a regressor.
A meta-analysis regressing log effect sizes on study year, where later studies are larger: the bias in a log estimate depends on the study’s size, later studies have less of it, and the fitted trend acquires a slope from the sizes rather than from the effects.
A variance-components model with unbalanced groups regressed on a group-level covariate: the same arithmetic, with the group sizes playing the role of the allocation.
A sequential trial modelling a nuisance across looks, where the looks accumulate data: the degrees of freedom rise monotonically with the look number, which is the regressor, so the bias is perfectly correlated with it.
In every case the repair is the same and is arithmetic: the bias is a known function of the degrees of freedom, so subtract it. What makes the trap worth a measurement rather than a warning is that the reasoning which walks into it is correct as far as it goes — the bias really does depend only on the degrees of freedom, and that really is what usually makes it harmless.
A model that is wrong about the shape
The second thing the model could be wrong about is the shape of the drift rather than its slope, and the question is which way it degrades: towards the pooled rule, or towards the local one.
Bend the truth away from the line the rule fits — a quadratic in the block covariate, centred so that the linear part is unchanged — and the model keeps fitting a straight line through a curve.
At the strongest curvature measured, the modelled rule covers at 94.6%, the pooled rule at 95.3%, and the local rule at 92.5%. The model degrades towards the pooled rule and not towards the local one, and that is the answer that matters: a model that is wrong about the shape is still two numbers estimated from twelve blocks, so its error is nearly common across the weights and nearly cancels. A local estimate is twelve numbers estimated from one block each however right its shape is.
That separates the two properties that were running together in the argument for modelling. It is not the model is right about the drift that makes it work — it is the model estimates few quantities from many blocks. The first is a hope about the world and the second is a fact about the estimator.
Why a bias that depends on nothing is not a safe bias
The expression ψ(k/2) + log 2 − log k has a property that makes it look like something a fit can absorb: it depends on the degrees of freedom and on nothing else. Not on the true variance, not on the effect, not on the covariate, not on anything the trial is trying to estimate. A quantity like that reads as a shift, and a shift is what an intercept is for.
The reasoning is sound and the conclusion is false, and the gap between them is the useful part. A bias that is constant in a nuisance parameter is absorbed by an intercept only if it is constant across the units of the fit. The units of this fit are blocks. The blocks alternate between allocations by design — that is what makes the variance ratio drift in the first place — so their degrees of freedom alternate, so the bias alternates: 0.5383 where the allocation is lopsided and zero where it is even. Nothing about that alternation is random. It is scheduled, it runs in lockstep with block index, and block index is exactly what the model regresses on.
So the bias is not a level; it is a second regressor, perfectly known and perfectly correlated with the one being fitted. Its projection onto the slope is predictable in advance — 0.0621 — and measured at 0.0597, and it puts the uncorrected fit 8.0 standard errors from a truth that no amount of data would have moved it towards. More blocks make it worse in the only sense that matters: the standard error shrinks and the bias does not, so the fit becomes more confidently wrong.
And the fix is not a correction factor, it is subtraction of a known quantity. Two digammas per block, from the two degrees of freedom the design already fixed before the trial opened, and the slope lands at 1.4976 against a truth of 1.5. Nothing is estimated to do it; nothing is tuned. The correction was available at the design stage, which is the same place the alternation that necessitates it was decided.
What it costs to omit is a third of a coverage point, and that number is the honest one to quote. This is a refusal that barely bites, and it is recorded as one. But the shape is the transferable thing and the shape is not small: a nuisance bias enters an estimate as a regressor whenever the design makes the nuisance vary, and a design that alternates makes it vary in the most correlated way available.
What is claimed here, and what is not
This essay takes where the bias in a modelled variance ratio ends up. The claims are that E[log(s²/σ²)] is ψ(k/2) + log 2 − log k and so depends on nothing but the degrees of freedom; that in a design whose blocks alternate between allocations the bias alternates with them — 0.5383 in the lopsided blocks and zero in the even ones — and the alternation is correlated with any covariate running across the trial; that the induced slope is predictable in advance at 0.0621 and measured at 0.0597, putting the uncorrected fit 8.0 standard errors from the truth; and that correcting it is a subtraction of two digammas and costs nothing.
What stays out and is named as a decision: what the correction is worth on a design whose allocation is constant. There the two biases are the same in every block, they really do go into the intercept, and the correction is exactly unnecessary — which is the case most textbook treatments have in mind and is the reason the sentence this essay is written against is usually true.
The boundary against the essay that introduced the model is that it is about which rules hold their level and this one is about a bias inside the rule that does.
The checks, and the refusals that make them mean something
Two claims are gated. The digamma is required to be right at the three points where it is elementary — ψ(1) = −γ, ψ(½) = −γ − 2 log 2, and ψ(n) = −γ + Σ 1/k — because the correction is arithmetic rather than an estimate and an arithmetic error in it would be invisible. And the corrected slope is required to be within four standard errors of the truth, which at four thousand trials is a tolerance of three hundredths on a quantity the uncorrected fit misses by six.
The refusal is the sentence the essay is named for. A log model fitted without the degrees-of-freedom correction is rejected, on the ground that the bias is a constant per allocation rather than per trial, and the allocation alternates with the covariate being fitted. It is a refusal that costs a third of a coverage point, and it is recorded as such: the reasoning it refuses is wrong, and the interval it produces is very nearly fine.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A width the trial has to stop for — both name allocation ratio, blinding, blocking, coverage, estimated variance, fixed-width interval, interval width, variance ratio, weighted least squares
- Stopping on the arms — both name blinding, blocking, coverage, estimated variance, fixed-width interval, interval width, variance ratio, weighted least squares
- A block size that changes — both name blinding, blocking, coverage, degrees of freedom, fixed-width interval, nuisance parameter, weighted least squares
- A schedule that reads the mean — both name blinding, blocking, coverage, degrees of freedom, fixed-width interval, nuisance parameter
- What a schedule actually buys — both name blinding, blocking, coverage, degrees of freedom, fixed-width interval, nuisance parameter
- What the blindfold costs — both name blinding, coverage, degrees of freedom, fixed-width interval, interval width, nuisance parameter
Named objects
A flat tag is an object no other essay names yet.
Allocation ratioBiasBlindingBlockingCoverageDegrees of freedomEstimated varianceExperimental designFixed-width intervalInterval widthModel misspecificationNuisance parameterRegressionVariance ratioWeighted least squares