The shape the covariate enters by

Three functions of one number

A rule that balances the covariate is exposed to every shape the outcome might have. A rule that balances three functions of it costs two points of variance against the shape the first was built for and takes the worst case from a coin's to about half of it.

Worth reading first: A design is a number · The variance removed before the data.

The rule that reads a continuous covariate takes a column of numbers and maximises the information about the treatment effect given that column. Nothing in it requires the column to be the covariate. Hand it x and it balances x; hand it x and x² and it balances both; hand it three functions of one number and it balances all three, at the same cost per arrival and with the same arithmetic.

That is a one-line change to a rule, and it turns the exposure the first essay of this field measured into a choice.

Four things a rule might read

The covariate, which is what every rule in the two fields before this one reads.

The covariate and its square, which is what a rule reads if it suspects the outcome curves.

The covariate and an indicator, which is what it reads if it suspects a threshold at a stated place.

All three, which is a rule that balances three functions of one number using the same units it had before.

Against the three shapes an outcome might have, that gives twelve numbers, and the twelve are the whole essay.

What a rule reads, against what the outcome uses. The variance of the treatment estimate relative to a coin's, for four things a rule might balance against three shapes the outcome might have, over 350 trials of 200 units. The diagonal is the easy part — a rule that reads the function the outcome uses removes about half the variance. What the table is for is the off-diagonal: reading the covariate alone is worth nothing against a quadratic (0.755), and reading all three is worth nearly as much against every shape as the matching rule is against its own (0.532, 0.493, 0.514).
Fig. 1 What each rule is worth against each shape, as the variance of the treatment estimate relative to a coin’s. The diagonal is unsurprising. The off-diagonal is the field.
the rule reads linear in x a threshold at 1 quadratic in x
the covariate 0.488 0.745 0.966
and its square 0.526 0.674 0.531
and an indicator 0.518 0.507 0.738
all three 0.508 0.497 0.502

What the rule actually does with three columns

The change to the rule is smaller than the change to what it protects, and it is worth being concrete about it because “balances a basis” can sound like a new procedure.

The rule already computes, for each arrival and each arm, what the information about the treatment effect would be if the arrival went there: S_aa − S_ax′S_xx⁻¹S_ax, with x the columns it has been given. With one column that correction term is a scalar; with three it is a quadratic form in a three by three matrix. The arithmetic per arrival is a few dozen multiplications either way, the running sums are the same running sums, and nothing about the rule’s shape changes.

What changes is the model the criterion is one over the variance of. With three columns it is

y=α+τa+β1x+β2x2+β31{x>c}+ε,y = \alpha + \tau a + \beta_1 x + \beta_2 x^2 + \beta_3 \mathbf{1}\{x > c\} + \varepsilon,

and the rule is now optimal for that model rather than for the straight line. It is still optimal for exactly one model — this is not a rule that has stopped assuming a shape — but the model it assumes now contains the three shapes the table measures as special cases, which is the whole of the protection.

Three ways one covariate can matter. Each curve is a standardised function of the same covariate: the covariate itself, an indicator for exceeding 1, and a quadratic. All three are scaled to unit variance, so an outcome built from any of them is explained to the same degree — what differs is only the shape. A rule that balances the covariate's mean is balancing the first curve exactly and the others only through their correlation with it: 0.6623 for the threshold and 0.000 for the quadratic, which is zero by symmetry.
Fig. 2 The three shapes the table is scored against, standardised so that how much the covariate matters is held fixed and only its shape differs.

The diagonal is the easy part

Each rule does well against the shape it was told about: 0.488, 0.531 and 0.507 in the three cells where the function the rule balances is the function the outcome uses. That is the same measurement the first essay of this field makes, arriving three times, and it says that the machinery works — a rule that balances the right function halves the variance of the estimate whatever that function is.

Two details in the diagonal are worth having. The gain is nearly identical in all three cases, which is what “standardised to unit variance” was for: how much a covariate matters is held fixed, so the only thing that could differ is the shape, and it does not. And it is worth about half, which is what the first field measured for a linear covariate — the number does not depend on the shape once the rule has been told it.

The off-diagonal is the exposure

A rule that reads only the covariate is at 0.966 against a quadratic outcome, which is a coin, and at 0.745 against a threshold. A rule that reads the covariate and an indicator at 1 is at 0.738 against a quadratic. A rule that reads the covariate and its square is at 0.674 against a threshold — better than the mean-only rule, because a square is not orthogonal to an indicator at 1 even though it is orthogonal to x.

Every off-diagonal cell is worse than its diagonal and none of them is a catastrophe, which is the useful shape for a decision: the exposure is real and bounded, and the question is what to pay to avoid it.

The worst shape each rule is exposed to. For each basis a rule may read, the worst of its three variance ratios across the three shapes — the maximin reading of the same table, and the number an experimenter who does not know the shape is actually exposed to. Reading the covariate alone leaves a worst case of 0.755, which is no better than a coin. Reading three functions of it leaves 0.532. The insurance costs 12.4 points of variance against the shape the single-function rule was built for, which is the cheapest protection measured anywhere on this site.
Fig. 3 The maximin reading of the same table: the worst of each row. The rule that reads the covariate alone has a worst case indistinguishable from a coin, and the rule that reads three functions has one at about half.

Nineteen times the exchange rate of the other field

The comparison with the design field’s maximin argument is worth doing in one currency, because the two answers differ by more than the essay’s “so different in size” suggests.

Here the insurance costs 2.0 points — 0.508 against 0.488 on the shape the one-column rule was built for — and buys 45.8 points of worst case, from 0.966 to 0.508. That is an exchange rate of 22.9 points bought per point paid.

In the design field the robust design costs 31.6 points at the guess and buys 38.6 at the worst corner, an exchange rate of 1.22.

A factor of nineteen between the two. Both are the same construction — take the worst case over a set of things that might be true, rather than the best case at one of them — and one of them is nearly free while the other is a genuine decision.

What separates them is what the protection is spent out of. A design for a non-linear model has a fixed number of runs and buys robustness by moving runs to settings that are wrong if the guess is right, so the cost is first-order in the budget. A balancing rule buys robustness by changing which of two arms each arrival goes to, and the number of units never moves; what it gives up is only that its criterion is aimed at a slightly larger model than the true one. That is a second-order cost, and second-order costs are what two points look like.

The practical form is a rule about when to bother arguing. Where the robust construction spends the budget, it is a decision; where it spends only the criterion, it is not. An experimenter facing the second kind should take the insurance without measuring the premium, because the premium is bounded by how much a slightly over-specified model costs and that is a couple of per cent.

The second column is the expensive one and the third is free

The diagonal has a non-monotonicity in it that the maximin reading passes over.

Against a linear outcome the one-column rule gives 0.488, the two-column rules give 0.526 and 0.518, and the three-column rule gives 0.508. The three-column rule is better than either two-column rule against the shape all of them are over-specified for.

So the cost of columns is not monotone: the second column costs 3.0 to 3.8 points and the third gives one to two of them back. That is not a rounding — the three cells differ by more than the whole premium the maximin reading is priced at.

The reason is that a criterion aimed at a three-function model is aimed at a space rather than at two particular functions, and the rule’s assignments are decided by a quadratic form in a well conditioned three-by-three matrix rather than a two-by-two one whose second column is a poorly chosen single alternative. Adding the third function does not add a constraint the linear shape has to pay for; it completes a basis, and a complete basis is cheaper to be optimal for than half of one.

The practical reading is the same as the field’s, sharpened. The premium is two points against one column, and it is negative against the natural intermediate choice — so a rule that has already been given a second function has no reason at all not to be given the third.

The maximin reading, and what the insurance costs

Take the worst cell in each row, which is what an experimenter who does not know the shape is actually exposed to:

  • the covariate alone: 0.966
  • and its square: 0.674
  • and an indicator: 0.738
  • all three: 0.508

The rule that reads three functions of the covariate has a worst case across the three shapes of 0.508 — nearly as good as the best any rule achieves against any single shape.

And the price is two points. Against a linear outcome, reading three functions gives 0.508 where reading the covariate alone gives 0.488. Twenty thousandths of the variance of the treatment estimate, in exchange for taking the worst case from a coin’s to half of one.

This is the maximin argument of the design field arriving in a completely different subject, and the two are worth comparing because the answers are so different in size. There, protecting a fourfold rectangle of parameter values costs 32 points of efficiency at the guess — the robust design is 68.4% efficient where the design built at the guess is 100% — and buys a worst case of 45.9% against 7.3%. Here the insurance costs two points and buys nearly everything.

The reason for the difference is worth stating, because it is not that one field is luckier. A design for a non-linear model has a fixed number of runs and has to spend them at settings; protecting a range means spending them at settings that are wrong for every parameter value in it. A balancing rule spends nothing: it has the same units either way, and reading a second function of the covariate costs only the degree to which the two functions compete for the same assignments. When protection is bought with allocations rather than with runs, it is nearly free.

The worst shape each rule is exposed toFor each basis a rule may read, the worst of its three variance ratios across the three shapes — the maximin reading of the same table, and the number an experimenter who does not know the shape is actually exposed to. Reading the covariate alone leaves a worst case of 1.062, which is no better than a coin. Reading three functions of it leaves 0.627. The insurance costs 5.8 points of variance against the shape the single-function rule was built for, which is the cheapest protection measured anywhere on this site.the covariate1.062and its square0.667and an indicator0.792all three0.627a coinwhat the rule is allowed to read350 trials of 400 units, worst of three shapesinsurance at 5.8 points
Fig. 4 The same worst cases in a trial twice the size. Drag it: the ordering does not move, because it is a statement about which functions are being balanced rather than about how well any of them is.
What balancing several numbers at once costs each of them. The criterion generalises without a word changing — the covariate imbalance becomes a vector and the correction a quadratic form — so the question is what it is worth rather than whether it can be done. At n = 200 with 200 trials per point, a rule balancing one covariate leaves 12.7% of a coin's imbalance in it; balancing eight leaves 23.2% in each. The assignment has a fixed amount of freedom and every covariate added takes a share of it. The rule degrades rather than failing: at eight covariates it is still four times better balanced than a coin, and the eight are being held simultaneously rather than in turn.
Fig. 5 The competition between constraints, measured on several covariates rather than several functions of one. Three functions of one number compete less, because they are correlated with each other.

The threshold’s position, one field along

The indicator in the basis has a place in it, and the previous essay says how much that matters: an indicator at 1 is correlated 0.6623 with the covariate, one at 2 only 0.3621. A basis containing an indicator at 1 therefore adds less new information than one containing an indicator at 2, because more of what it says is already said by x.

That gives the basis choice an unexpected shape. The threshold worth putting in the basis is the one furthest from the median, because it is the one a mean-balancing rule protects against least — and the one an experimenter is most likely to guess wrong, since a rare cutoff is a rare cutoff. Putting an indicator at the median into a basis that already contains x is nearly redundant: it adds a function that is 80% explained by one already there.

That is the same arithmetic as the field’s opening identity, read as advice rather than as a measurement, and it is the only piece of guidance here that does not need a table.

Where the free lunch stops

Three functions is not the end of the sequence, and it is worth saying what stops it being.

Every function added is a constraint the assignment has to satisfy with the same units, and the constraints compete. With n units the rule is choosing among 2ⁿ assignments to make k quantities small at once, and each additional quantity takes something from the others — the same trade measured for several covariates in the field before this one, where balancing four independent covariates costs each of them a measurable amount.

Here the competition is milder because the functions are functions of one number: x, x² and an indicator are correlated with each other, so balancing one partly balances the others, and the third constraint costs less than the second. That is why the cost above is two points rather than ten.

It is also why the sequence has to stop somewhere: a basis rich enough to cover every shape the outcome might have is a basis with as many functions as units, and a rule that balances n quantities with n units has no freedom left and is not a randomisation at all. Somewhere between three functions and n there is a point where the constraints exhaust the design, and the rule stops being able to randomise long before it stops being able to balance — which is the same endpoint that a deterministic rule reaches for a different reason.

What balancing several numbers at once costs each of them. The criterion generalises without a word changing — the covariate imbalance becomes a vector and the correction a quadratic form — so the question is what it is worth rather than whether it can be done. At n = 200 with 200 trials per point, a rule balancing one covariate leaves 12.7% of a coin's imbalance in it; balancing eight leaves 23.2% in each. The assignment has a fixed amount of freedom and every covariate added takes a share of it. The rule degrades rather than failing: at eight covariates it is still four times better balanced than a coin, and the eight are being held simultaneously rather than in turn.
Fig. 6 The competition between constraints in the field that measured it: balancing several covariates at once costs each of them. Three functions of one number are much more correlated than three separate covariates, which is why the cost here is smaller.
What each analysis does at a true null, by shape. Four analyses of the same trials — 500 of them at each shape, 120 units, assigned by the rule that reads the covariate. Every rejection is false. The unadjusted analysis is the one that moves: 1.60% against a linear outcome, where the design removed a great deal that the standard error still prices, and 5.20% against a quadratic, where it removed nothing and the standard error is right. Adjusting holds the level in all three columns, and so does the design's own reference distribution, which needs to be told the rule and nothing else.
Fig. 7 What the analyses do at a true null, which is the other lever and the subject of the next essay. The top row is the one that moves with the shape.

What the analysis is doing while this happens

One objection deserves a straight answer before the practical part: why bother with the design at all, when the analysis can adjust for whatever it likes after the fact?

Because the two do different things and only one of them can be done late. Adjustment removes the part of the outcome the adjusted-for function explains, and it does so whatever the imbalance was — so an analysis that knows the shape recovers the precision regardless of the design. What it cannot do is recover the precision when it does not know the shape, and an analyst who did not know the shape at the design stage does not know it at the analysis stage either.

The measurements in the last essay of this field put numbers on that: against a threshold outcome, adjusting for the covariate recovers about a third of the available power and adjusting for the true function recovers all of it. So the design and the analysis are exposed to the same unknown, and enriching the basis is the only one of the two moves that can be made before anybody has seen an outcome.

There is also a reason to prefer the design that has nothing to do with precision. A basis fixed in advance is a commitment; a functional form chosen after the outcomes have been seen is an analysis with a search inside it, and the search has to be priced. Balancing three functions costs two points of variance and no multiplicity at all.

What a practitioner would have to decide

The basis has to be written down before the trial, and it is a statement about which shapes are worth protecting against rather than a guess at which one is right.

The measurements say three things about that decision. Adding a square is worth more than adding an indicator at a stated place, because a square protects against curvature anywhere while an indicator protects against a threshold at one point. An indicator at the wrong place is worth something anyway — the rule that reads an indicator at 1 is at 0.738 against a quadratic, better than the mean-only rule’s 0.966 — because any second function breaks the symmetry that leaves a mean-balancing rule helpless. And the marginal value of the third function is small, which is the usual shape of an insurance premium.

A fourth thing the table says is easy to miss: the row that reads the covariate and its square is better against a threshold than the row that reads the covariate alone, at 0.674 against 0.745, even though nobody told it about any threshold. A square is not orthogonal to an indicator at 1, so protecting against curvature protects partly against a cut, and any enrichment of the basis is worth something against shapes it was not chosen for.

What none of the measurements can say is which shapes to include, because that is a statement about the world. What they do say is that the current default — one function, the covariate itself — is the one choice with a worst case no better than tossing a coin.

What a rule reads, against what the outcome uses. The variance of the treatment estimate relative to a coin's, for four things a rule might balance against three shapes the outcome might have, over 350 trials of 400 units. The diagonal is the easy part — a rule that reads the function the outcome uses removes about half the variance. What the table is for is the off-diagonal: reading the covariate alone is worth nothing against a quadratic (1.062), and reading all three is worth nearly as much against every shape as the matching rule is against its own (0.627, 0.590, 0.593).
Fig. 8 The same twelve cells in a trial of four hundred units. Nothing about the table’s shape depends on the trial size: these are ratios of variances, and both variances fall at the same rate.
Four rules, three shapes, and no ordering that survives. The variance of the unadjusted treatment estimate under each rule, as a fraction of the variance a coin gives, over 450 trials of 200 units each. Against a covariate that enters linearly the rule that reads the number nearly halves it. Against a threshold at 1 it removes about a fifth. Against a quadratic every rule here is at or worse than a coin — they are all optimising a criterion that is one over the variance of an estimate in a model this outcome does not obey, and a constraint that helps nothing still costs something. Nothing in a trial says which column it is in.
Fig. 9 The default, measured across the four rules the previous field compares. Every row of that table is a rule reading one function of the covariate, which is why every row has the same exposure.

What is claimed here, and what is not

This essay takes balancing a basis rather than a mean, and the claims are the twelve-cell table, the maximin reading of it, and the two-point cost of the insurance.

What stays out and is named as a decision: bases with more than three functions, and the point at which the constraints exhaust the design — named above and not measured, because it needs a trial size sweep against a basis size sweep and the interesting part of it is where the randomisation dies rather than where the balance stops improving; splines and other bases that are not polynomials or indicators; and the analysis, which is a separate lever and the subject of the last essay of this field.

The boundary against the design fields is the currency. A maximin design spends runs at settings and pays for protection in efficiency; a maximin rule spends assignments and pays for protection in a few points of variance. Everything about why a worst case is the right thing to optimise, and why its optimum sits on a tie, is established there and used here without being re-derived.

The checks, and the refusals that make them mean something

Two claims are gated in this field’s library. Reading three functions is required to cost less than six points against the shape one function was built for, and to have the best worst case across the three shapes by a wide margin — the two halves of the insurance argument, which would fail in opposite directions if the trade were being described wrongly. And the rule that reads the mean alone is required to have a worst case no better than a coin’s, which is the exposure this essay exists to repair and is measured on the same trials.

The refusals for this field are the two in the essays on either side. The one that bears on this essay is the rule reported on the covariate it reads: every rule in the table above removes nearly all of the imbalance in x, including the ones with a worst case at a coin’s, so the balance table that a trial reports is exactly as flattering for the best row as for the worst.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Allocation ruleBasis functionsCovariate balanceDesign criterionEfficiencyImbalanceInformation matrixMaximin designModel misspecificationNuisance parameterRobustnessThresholdTreatment effectVariance reduction