The shape the covariate enters by

The analysis and the shape

An unadjusted analysis after a rule that read the covariate is too cautious — by a third against a linear outcome, by nothing at all against a quadratic. And an adjustment for the wrong function recovers almost none of the precision the right one would.

Worth reading first: Randomisation is not balance · Balancing what is known in advance.

A trial assigned by a rule that reads the covariate leaves the analysis with a problem the analysis did not create: the assignment is not a coin, the standard error assumes it was, and the two disagree. The previous field measured that disagreement and reported it in one number — an unadjusted analysis after such a rule rejects 0.50% of true nulls at a nominal 5%, because its standard error is 1.384 times the actual spread of the estimate it is dividing.

That number is a fact about a linear outcome, and it does not survive a change of shape.

The conservatism, by shape

The same rule, the same trials, the same analysis, three outcomes: linear in the covariate, a threshold at 1, and a quadratic. All three standardised so the covariate matters equally.

the outcome depends on unadjusted adjusted for x its standard error, over its own spread
x, linearly 1.05% 4.85% 1.373
a threshold at 1 3.10% 5.45% 1.084
x², quadratically 5.45% 5.55% 0.976
What each analysis does at a true null, by shape. Four analyses of the same trials — 500 of them at each shape, 120 units, assigned by the rule that reads the covariate. Every rejection is false. The unadjusted analysis is the one that moves: 1.60% against a linear outcome, where the design removed a great deal that the standard error still prices, and 5.20% against a quadratic, where it removed nothing and the standard error is right. Adjusting holds the level in all three columns, and so does the design's own reference distribution, which needs to be told the rule and nothing else.
Fig. 1 Four analyses at a true null, by shape. The top row is the one that moves, and it moves across the whole range from a fifth of its nominal level to exactly it.

The mechanism is in the last column and it is arithmetic rather than statistics. The unadjusted standard error prices the variability a coin would have produced. Where the design removed a lot of that variability the price is too high — 1.373 times the estimate’s actual spread — and the test is conservative. Where the design removed nothing, because the outcome depends on a function the rule could not see, the price is right and the test is at its level.

So the previous field’s headline is not wrong, it is conditional on a shape nobody checked. The conservatism is not a property of covariate-adaptive allocation; it is a property of covariate-adaptive allocation working, and it disappears exactly when the design has stopped helping.

There is an uncomfortable corollary. A conservative test is a visible symptom of a design that worked, and a test at its nominal level is what a design that did nothing looks like. An analyst who checks the calibration of an unadjusted analysis and finds it correct has evidence that the trial was, in the way that matters, randomised by a coin.

Why the standard error is too big, in one line

The mechanism is worth writing out, because “the standard error prices a coin’s variability” is a sentence that can be read as hand-waving and is not.

An unadjusted difference in means has, over repeated trials assigned by a coin, a variance made of two parts: the outcome noise, and the part of the covariate’s contribution that the imbalance carries into the estimate. The usual standard error estimates the whole thing from the residual spread of the observed outcomes, which contains both.

A rule that balances the covariate removes the second part from the estimate’s variance and does nothing to the residual spread, because the residuals still contain the covariate’s contribution to each unit’s outcome. So the estimated standard error goes on describing a trial nobody ran. The factor between them — 1.373 against a linear outcome, 0.976 against a quadratic — is exactly the square root of how much of that second part the design removed, which is the share the rule can see arriving in the analysis.

That also says what would change the factor without changing the rule: more outcome noise dilutes it towards one, and a covariate that explains more of the outcome pushes it up. The 1.373 is a number about one signal-to-noise ratio as well as one shape.

The same four, against a real effect of 0.3. 400 trials of 120 patients allocated by minimisation at p = 0.8, with the prognostic factors carrying a real effect on the outcome and a treatment effect of 0.3. Two statistics, the plain difference and the same after adjusting for the balanced factors, each read against two reference distributions: a t table, and the set of allocations the rule could have produced from these covariates. The re-randomised rows give up almost nothing: 33.5% against 37.3%. The exactness is nearly free here, where the same construction cost nineteen points of power against a rule that read outcomes.
Fig. 2 Four analyses of one trial in the field that built them, against a real effect and after a rule that reads factor levels rather than numbers. Everything there is measured with the covariate entering the outcome linearly.

Adjusting holds the level, whatever the shape

The second column is the reassuring one. Adjusting for the covariate — the ordinary linear adjustment, with no knowledge of the outcome’s shape — rejects 4.85%, 5.45% and 5.55% at a nominal 5% across the three shapes, which is within noise of the level everywhere.

That is worth stating plainly because it is the one thing in this field that does not depend on the shape. Adjusting for a covariate that enters the outcome by some other function does not break the test: the adjustment removes what it can, the residual variance absorbs what it cannot, and the resulting standard error describes the estimate it is attached to. A mis-specified adjustment is inefficient, not invalid.

And recovers almost nothing

Efficiency is the other half, and it is where the shape returns. Give the trials a real effect and count how often each analysis finds it.

the outcome depends on unadjusted adjusted for x adjusted for the true shape
x, linearly 68.6% 90.1% 90.1%
a threshold at 1 66.4% 73.9% 90.1%
x², quadratically 64.8% 65.9% 89.8%
What each analysis finds, by shape. Four analyses of the same trials — 500 of them at each shape, 120 units, assigned by the rule that reads the covariate. The effect is real, so every number is power. Adjusting for the covariate gives the whole of what is available when the outcome is linear in it and about a third of it against a threshold; against a quadratic it recovers almost nothing, and only an analysis told the true shape reaches 87.6%.
Fig. 3 The same four analyses against a real effect. The third row is what the trial could have had; the second is what an analyst who adjusts for the covariate actually gets.

Against a linear outcome the ordinary adjustment is the right adjustment and gets all 21.5 points of the available power. Against a threshold it gets 7.5 of 23.7. Against a quadratic it gets 1.1 of 25.0 — the covariate is uncorrelated with the function the outcome uses, so regressing on it removes nothing at all and the adjustment is a wasted degree of freedom.

The design and the analysis therefore fail together, and for the same reason. Where the rule could not see the outcome’s structure, the linear adjustment cannot either, because both of them are projections onto the same column. An experimenter who worried about this at only one of the two stages has not covered themselves at either.

One more reading of the power table is worth making explicit, because it is the number an experimenter would actually be deciding with. Against a threshold outcome, the trial as designed and analysed delivers 73.9% power. The same trial, analysed by somebody who knew the shape, delivers 90.1%. Recovering those sixteen points by enlarging the trial instead would take roughly a third more units — so the cost of not knowing the shape, at this setting, is a third of the experiment.

Against a quadratic outcome it is worse: 65.9% against 89.8%, which is most of the difference between a trial that will probably work and one that will probably not.

Three analyses of the same trials, none of them wrong about the data. 320 trials at n = 60 with no treatment effect at all, so every rejection counted is a false one, and a covariate that drives the outcome with coefficient 1. The unadjusted comparison is at 5.94% after a coin — its level — and at 0.00% after the rule that reads the covariate: the design removed the imbalance and the analysis is still pricing it. Adjusting for the covariate gives 4.06%, and the rule's own reference distribution — hold the outcomes, re-run the rule 199 times, count — gives 3.13% against the 4.5% that 199 draws can deliver. The last of the three has to be told the assignment rule and nothing else, which is the one thing the experimenter certainly knows.
Fig. 4 The unadjusted analysis at a true null in the field that measured it, where the conservatism this essay takes apart was first counted.

The analysis that needs to be told the rule

There is a fourth analysis, and it is the one this site keeps arriving at when a design has done something an ordinary standard error cannot price: hold the outcomes fixed, re-run the same rule many times, and count how often it produces a statistic at least as large.

That test is exact by construction for any rule, because its reference distribution is the distribution the rule actually induces. It has to be told the rule; it does not have to be told anything about the outcome model, so nothing in this essay’s table can break it.

Measured at a real effect on a hundred units, it finds the effect 91.5% of the time, against 94.5% for the linear adjustment and 76.8% for the unadjusted analysis. Three points behind the adjustment, and it has assumed nothing about the shape at all.

That is an unusually good bargain and it comes with two ways to lose it, both measured.

The reference has to be the rule that ran. Re-randomising with a coin — which is what “randomisation test” means to most readers — after a trial assigned by a rule that reads the covariate prices imbalances the design could not have produced. The reference is far too wide, the test rejects a true null 0.5% of the time, and its power falls from 91.5% to 70.0%.

And the rule has to have some randomness left. A rule that always takes its preferred arm is a function of the covariate sequence: re-running it on the same units returns the same assignment every time, the reference distribution is a point mass at the observed statistic, and the p-value is 1 whatever the data say. Measured, that test finds a real effect 0.0% of the time.

The second of those is the sharp end of a trade the balancing field states in general terms — the more deterministic the rule, the better it balances and the less of the design’s own randomness is left to test with. At the endpoint it is not a loss of power. It is the absence of a test.

What a rule reads, against what the outcome uses. The variance of the treatment estimate relative to a coin's, for four things a rule might balance against three shapes the outcome might have, over 350 trials of 200 units. The diagonal is the easy part — a rule that reads the function the outcome uses removes about half the variance. What the table is for is the off-diagonal: reading the covariate alone is worth nothing against a quadratic (0.755), and reading all three is worth nearly as much against every shape as the matching rule is against its own (0.532, 0.493, 0.514).
Fig. 5 The design-stage table once more. An analyst who does not know which column the trial is in is choosing an adjustment against the same unknown the rule was choosing a basis against.

The first column and the last are the same number

The size table’s first and last columns are presented as symptom and mechanism, and they are close enough to being the same quantity that each can be recovered from the other — which is worth doing, because it is the check that says the mechanism really is the mechanism.

A statistic whose reported standard error is r times its actual spread is being compared against a threshold r times too far out, so it rejects about 2(1 − Φ(1.96r)) of the time. Running that backwards from each rejection rate gives r = 1.31 from the linear column’s 1.05%, 1.10 from the threshold’s 3.10%, and 0.98 from the quadratic’s 5.45%.

Against the reported 1.373, 1.084 and 0.976 those are right to within two per cent in two of the three cases and about five per cent in the first — and the first is the deepest tail, where a normal approximation to a t statistic’s rejection rate is worst and where five hundred trials resolve least. So the two columns are one measurement seen twice, and neither is telling the reader anything the other does not.

That is more useful than a consistency check, because the two columns are available in different circumstances. A trial can compute the ratio and cannot compute the rejection rate: the reported standard error is on the output, and the estimate’s actual spread over repeated trials can be obtained from the rule’s own re-randomisations without any outcome model. A single trial can therefore say how conservative its own unadjusted analysis is, in a number, before deciding what to report — which is the diagnostic the section above asks for, arriving with a scale attached.

The two ways of getting the reference wrong are one way

The essay’s two failure modes for the exact test look unrelated — a coin’s reference distribution after a rule that was not a coin, and a rule with no randomness left — and the first of them is not a new failure at all.

Re-randomising with a coin after a covariate-reading rule rejects a true null 0.5% of the time. The unadjusted analysis against a t table rejects 1.05%. Both are pricing the variability a coin would have produced against an estimate whose variability the design reduced, and they land in the same place because they are the same mistake: one gets its too-wide reference from a formula and the other from a simulation.

So “use a randomisation test” is not the repair; “use the rule’s own reference distribution” is. A randomisation test built on the wrong rule is not a weaker version of the right one — it reproduces, to within the noise of the count, the exact defect of the parametric analysis it was reached for instead of, and it costs 21.5 points of power to do it.

The second failure is of a different kind and the contrast is the point. A deterministic rule does not mis-price anything; it leaves nothing to price, and a reference distribution with one point in it finds a real effect 0.0% of the time. One failure is a wrong answer and the other is no answer, which is the distinction the analysis that has to know the rule draws between a test that misleads and a trial that is wasted.

What the trial cannot tell the analyst

Everything above depends on the outcome’s shape, and the shape is not observable in a useful sense at the moment the analysis has to be written down.

It is worth being precise about that, because the outcomes are observed and a determined analyst can look at them. What they cannot do is look at them and then choose the adjustment, without turning the analysis into a search over specifications — which this site has a whole field about, and whose price is a reference distribution nobody computes for an adjustment they chose after a scatterplot.

So the honest options are the ones above: commit to an adjustment in advance and accept that it may recover a third of what was available; commit to a richer basis at the design stage, before any outcome exists; or use the design’s own reference distribution, which is exact whatever the shape and requires no commitment about the outcome model at all.

There is one diagnostic that costs nothing and is not a search: compute the imbalance in a few candidate functions of the covariate and report them beside the balance table. A trial whose covariate is balanced to a hundredth and whose indicator at two standard deviations is imbalanced like a coin’s has said something true about itself, and it has said it without looking at a single outcome.

What an experimenter should take from the four rows

The table has a straightforward reading and it is worth writing down, because three of the four analyses are defensible and they are not interchangeable.

Do not report an unadjusted analysis after a rule that read the covariate. It is either conservative by a factor that depends on an unknown shape, or correctly calibrated for the wrong reason. Neither is a property to build a conclusion on, and the previous field reached the same conclusion from a single column of this table.

Report the imbalance in more than one function of the covariate. It costs nothing, it needs no outcome, and it is the only thing in the trial that distinguishes a design that worked from one that did not.

Adjust for the covariate, and expect it to be worth less than it looks. It holds the level whatever the shape, which is the important part, and it recovers all of the available precision only if the shape is the one it assumes.

If the shape is genuinely unknown, the design’s own reference distribution costs three points of power and assumes nothing — provided it is handed the rule that ran and the rule left something to re-randomise.

And enrich the basis at the design stage, which is the other lever and the only one available before any outcome exists.

The worst shape each rule is exposed to. For each basis a rule may read, the worst of its three variance ratios across the three shapes — the maximin reading of the same table, and the number an experimenter who does not know the shape is actually exposed to. Reading the covariate alone leaves a worst case of 0.755, which is no better than a coin. Reading three functions of it leaves 0.532. The insurance costs 12.4 points of variance against the shape the single-function rule was built for, which is the cheapest protection measured anywhere on this site.
Fig. 6 The design-stage lever, for comparison: what each basis a rule might read is exposed to. Two points of variance there buys more than any choice of analysis afterwards can.
What each analysis finds, by shapeFour analyses of the same trials — 500 of them at each shape, 120 units, assigned by the rule that reads the covariate. The effect is real, so every number is power. Adjusting for the covariate gives the whole of what is available when the outcome is linear in it and about a third of it against a threshold; against a quadratic it recovers almost nothing, and only an analysis told the true shape reaches 87.6%.linear in xa threshold at x = 1quadratic in xno adjustment66.0%64.2%65.2%adjusted for the covariate88.0%71.2%66.8%adjusted for the true shape88.0%87.8%87.6%the design's own reference85.0%70.2%63.8%the analysis500 trials of 120 units at each shapepower at an effect of 0.3
Fig. 7 The analyses once more, at a real effect. Drag it back to a true null and the third row stops being the best analysis and becomes the one that is merely correct, which is the difference between power and level in one picture.

One asymmetry between the two levers is worth naming before the summary. The design lever has to be pulled first and cannot be pulled again; the analysis lever can be pulled at any time and cannot recover what the design gave away. That ordering is why a field about balancing rules ends with an essay about analyses rather than the other way round: the analysis inherits whatever the design left, and the only question left for it is how much of that it can still use.

What is claimed here, and what is not

This essay takes what the analysis does once the design has finished, and the claims are three: the conservatism of an unadjusted analysis is a fact about the outcome’s shape and vanishes with it; a mis-specified adjustment is valid and nearly useless; and the design’s own reference distribution is exact whatever the shape and needs the rule to have randomness in it.

What stays out and is named as a decision: adjustment for a basis rather than for the covariate, which is the analysis-side version of the previous essay and would have the same shape as its answer; choosing the adjustment after seeing the outcomes, which is a search and would need its multiplicity priced; and the conditional versus unconditional question — whether the estimate should be reported against the imbalance that actually occurred or against the distribution of imbalances the rule produces — which is a genuine dispute this site has not taken a position in.

The boundary against the field before this one is the shape. That an unadjusted analysis after a covariate-reading rule is conservative, that adjusting repairs it, and that the rule’s own reference distribution is exact are all established there, on a linear outcome; what this essay adds is that the first of those three is the only one that depended on the linearity, and that it depended on it completely.

The checks, and the refusals that make them mean something

Two claims are gated in this field’s library. The conservatism is required to be present against a linear outcome and absent against a quadratic one, with the ratio of the reported standard error to the estimate’s actual spread required to be near 1.37 in the first case and within eight points of 1 in the second — the mechanism as well as the symptom, so that a change in one without the other would fail. And the adjustment is required to hold its level in all three columns while recovering most of the available power in one and almost none in another.

Two refusals stand behind them, and both are about the reference distribution rather than about the model. A reference generated by a coin after a trial that was not assigned by one takes the power from 91.5% to 70.0% and the rate at a true null to 0.5%. And a rule with no randomness left produces a reference distribution with one point in it, which finds a real effect 0.0% of the time — a check that cannot fail, arriving in the one place on this site where a procedure’s exactness is a theorem rather than a measurement.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Allocation ruleConditional inferenceCovariate adjustmentCovariate balanceDeterministic ruleError rateExact testModel misspecificationPowerRandomisation testReference distributionStandard errorThresholdTreatment effect