Three functions of one number
Worth reading first: A design is a number · The variance removed before the data.
The rule that reads a continuous covariate takes a column of numbers and maximises the information about the treatment effect given that column. Nothing in it requires the column to be the covariate. Hand it x and it balances x; hand it x and x² and it balances both; hand it three functions of one number and it balances all three, at the same cost per arrival and with the same arithmetic.
That is a one-line change to a rule, and it turns the exposure the first essay of this field measured into a choice.
Four things a rule might read
The covariate, which is what every rule in the two fields before this one reads.
The covariate and its square, which is what a rule reads if it suspects the outcome curves.
The covariate and an indicator, which is what it reads if it suspects a threshold at a stated place.
All three, which is a rule that balances three functions of one number using the same units it had before.
Against the three shapes an outcome might have, that gives twelve numbers, and the twelve are the whole essay.
| the rule reads | linear in x | a threshold at 1 | quadratic in x |
|---|---|---|---|
| the covariate | 0.488 | 0.745 | 0.966 |
| and its square | 0.526 | 0.674 | 0.531 |
| and an indicator | 0.518 | 0.507 | 0.738 |
| all three | 0.508 | 0.497 | 0.502 |
What the rule actually does with three columns
The change to the rule is smaller than the change to what it protects, and it is worth being concrete about it because “balances a basis” can sound like a new procedure.
The rule already computes, for each arrival and each arm, what the information about the treatment effect would be if the arrival went there: S_aa − S_ax′S_xx⁻¹S_ax, with x the columns it has been given. With one column that correction term is a scalar; with three it is a quadratic form in a three by three matrix. The arithmetic per arrival is a few dozen multiplications either way, the running sums are the same running sums, and nothing about the rule’s shape changes.
What changes is the model the criterion is one over the variance of. With three columns it is
and the rule is now optimal for that model rather than for the straight line. It is still optimal for exactly one model — this is not a rule that has stopped assuming a shape — but the model it assumes now contains the three shapes the table measures as special cases, which is the whole of the protection.
The diagonal is the easy part
Each rule does well against the shape it was told about: 0.488, 0.531 and 0.507 in the three cells where the function the rule balances is the function the outcome uses. That is the same measurement the first essay of this field makes, arriving three times, and it says that the machinery works — a rule that balances the right function halves the variance of the estimate whatever that function is.
Two details in the diagonal are worth having. The gain is nearly identical in all three cases, which is what “standardised to unit variance” was for: how much a covariate matters is held fixed, so the only thing that could differ is the shape, and it does not. And it is worth about half, which is what the first field measured for a linear covariate — the number does not depend on the shape once the rule has been told it.
The off-diagonal is the exposure
A rule that reads only the covariate is at 0.966 against a quadratic outcome, which is a coin, and at 0.745 against a threshold. A rule that reads the covariate and an indicator at 1 is at 0.738 against a quadratic. A rule that reads the covariate and its square is at 0.674 against a threshold — better than the mean-only rule, because a square is not orthogonal to an indicator at 1 even though it is orthogonal to x.
Every off-diagonal cell is worse than its diagonal and none of them is a catastrophe, which is the useful shape for a decision: the exposure is real and bounded, and the question is what to pay to avoid it.
Nineteen times the exchange rate of the other field
The comparison with the design field’s maximin argument is worth doing in one currency, because the two answers differ by more than the essay’s “so different in size” suggests.
Here the insurance costs 2.0 points — 0.508 against 0.488 on the shape the one-column rule was built for — and buys 45.8 points of worst case, from 0.966 to 0.508. That is an exchange rate of 22.9 points bought per point paid.
In the design field the robust design costs 31.6 points at the guess and buys 38.6 at the worst corner, an exchange rate of 1.22.
A factor of nineteen between the two. Both are the same construction — take the worst case over a set of things that might be true, rather than the best case at one of them — and one of them is nearly free while the other is a genuine decision.
What separates them is what the protection is spent out of. A design for a non-linear model has a fixed number of runs and buys robustness by moving runs to settings that are wrong if the guess is right, so the cost is first-order in the budget. A balancing rule buys robustness by changing which of two arms each arrival goes to, and the number of units never moves; what it gives up is only that its criterion is aimed at a slightly larger model than the true one. That is a second-order cost, and second-order costs are what two points look like.
The practical form is a rule about when to bother arguing. Where the robust construction spends the budget, it is a decision; where it spends only the criterion, it is not. An experimenter facing the second kind should take the insurance without measuring the premium, because the premium is bounded by how much a slightly over-specified model costs and that is a couple of per cent.
The second column is the expensive one and the third is free
The diagonal has a non-monotonicity in it that the maximin reading passes over.
Against a linear outcome the one-column rule gives 0.488, the two-column rules give 0.526 and 0.518, and the three-column rule gives 0.508. The three-column rule is better than either two-column rule against the shape all of them are over-specified for.
So the cost of columns is not monotone: the second column costs 3.0 to 3.8 points and the third gives one to two of them back. That is not a rounding — the three cells differ by more than the whole premium the maximin reading is priced at.
The reason is that a criterion aimed at a three-function model is aimed at a space rather than at two particular functions, and the rule’s assignments are decided by a quadratic form in a well conditioned three-by-three matrix rather than a two-by-two one whose second column is a poorly chosen single alternative. Adding the third function does not add a constraint the linear shape has to pay for; it completes a basis, and a complete basis is cheaper to be optimal for than half of one.
The practical reading is the same as the field’s, sharpened. The premium is two points against one column, and it is negative against the natural intermediate choice — so a rule that has already been given a second function has no reason at all not to be given the third.
The maximin reading, and what the insurance costs
Take the worst cell in each row, which is what an experimenter who does not know the shape is actually exposed to:
- the covariate alone: 0.966
- and its square: 0.674
- and an indicator: 0.738
- all three: 0.508
The rule that reads three functions of the covariate has a worst case across the three shapes of 0.508 — nearly as good as the best any rule achieves against any single shape.
And the price is two points. Against a linear outcome, reading three functions gives 0.508 where reading the covariate alone gives 0.488. Twenty thousandths of the variance of the treatment estimate, in exchange for taking the worst case from a coin’s to half of one.
This is the maximin argument of the design field arriving in a completely different subject, and the two are worth comparing because the answers are so different in size. There, protecting a fourfold rectangle of parameter values costs 32 points of efficiency at the guess — the robust design is 68.4% efficient where the design built at the guess is 100% — and buys a worst case of 45.9% against 7.3%. Here the insurance costs two points and buys nearly everything.
The reason for the difference is worth stating, because it is not that one field is luckier. A design for a non-linear model has a fixed number of runs and has to spend them at settings; protecting a range means spending them at settings that are wrong for every parameter value in it. A balancing rule spends nothing: it has the same units either way, and reading a second function of the covariate costs only the degree to which the two functions compete for the same assignments. When protection is bought with allocations rather than with runs, it is nearly free.
The threshold’s position, one field along
The indicator in the basis has a place in it, and the previous essay says how much that matters: an indicator at 1 is correlated 0.6623 with the covariate, one at 2 only 0.3621. A basis containing an indicator at 1 therefore adds less new information than one containing an indicator at 2, because more of what it says is already said by x.
That gives the basis choice an unexpected shape. The threshold worth putting in the basis is the one furthest from the median, because it is the one a mean-balancing rule protects against least — and the one an experimenter is most likely to guess wrong, since a rare cutoff is a rare cutoff. Putting an indicator at the median into a basis that already contains x is nearly redundant: it adds a function that is 80% explained by one already there.
That is the same arithmetic as the field’s opening identity, read as advice rather than as a measurement, and it is the only piece of guidance here that does not need a table.
Where the free lunch stops
Three functions is not the end of the sequence, and it is worth saying what stops it being.
Every function added is a constraint the assignment has to satisfy with the same units, and the constraints compete. With n units the rule is choosing among 2ⁿ assignments to make k quantities small at once, and each additional quantity takes something from the others — the same trade measured for several covariates in the field before this one, where balancing four independent covariates costs each of them a measurable amount.
Here the competition is milder because the functions are functions of one number: x, x² and an indicator are correlated with each other, so balancing one partly balances the others, and the third constraint costs less than the second. That is why the cost above is two points rather than ten.
It is also why the sequence has to stop somewhere: a basis rich enough to cover every shape the outcome might have is a basis with as many functions as units, and a rule that balances n quantities with n units has no freedom left and is not a randomisation at all. Somewhere between three functions and n there is a point where the constraints exhaust the design, and the rule stops being able to randomise long before it stops being able to balance — which is the same endpoint that a deterministic rule reaches for a different reason.
What the analysis is doing while this happens
One objection deserves a straight answer before the practical part: why bother with the design at all, when the analysis can adjust for whatever it likes after the fact?
Because the two do different things and only one of them can be done late. Adjustment removes the part of the outcome the adjusted-for function explains, and it does so whatever the imbalance was — so an analysis that knows the shape recovers the precision regardless of the design. What it cannot do is recover the precision when it does not know the shape, and an analyst who did not know the shape at the design stage does not know it at the analysis stage either.
The measurements in the last essay of this field put numbers on that: against a threshold outcome, adjusting for the covariate recovers about a third of the available power and adjusting for the true function recovers all of it. So the design and the analysis are exposed to the same unknown, and enriching the basis is the only one of the two moves that can be made before anybody has seen an outcome.
There is also a reason to prefer the design that has nothing to do with precision. A basis fixed in advance is a commitment; a functional form chosen after the outcomes have been seen is an analysis with a search inside it, and the search has to be priced. Balancing three functions costs two points of variance and no multiplicity at all.
What a practitioner would have to decide
The basis has to be written down before the trial, and it is a statement about which shapes are worth protecting against rather than a guess at which one is right.
The measurements say three things about that decision. Adding a square is worth more than adding an indicator at a stated place, because a square protects against curvature anywhere while an indicator protects against a threshold at one point. An indicator at the wrong place is worth something anyway — the rule that reads an indicator at 1 is at 0.738 against a quadratic, better than the mean-only rule’s 0.966 — because any second function breaks the symmetry that leaves a mean-balancing rule helpless. And the marginal value of the third function is small, which is the usual shape of an insurance premium.
A fourth thing the table says is easy to miss: the row that reads the covariate and its square is better against a threshold than the row that reads the covariate alone, at 0.674 against 0.745, even though nobody told it about any threshold. A square is not orthogonal to an indicator at 1, so protecting against curvature protects partly against a cut, and any enrichment of the basis is worth something against shapes it was not chosen for.
What none of the measurements can say is which shapes to include, because that is a statement about the world. What they do say is that the current default — one function, the covariate itself — is the one choice with a worst case no better than tossing a coin.
What is claimed here, and what is not
This essay takes balancing a basis rather than a mean, and the claims are the twelve-cell table, the maximin reading of it, and the two-point cost of the insurance.
What stays out and is named as a decision: bases with more than three functions, and the point at which the constraints exhaust the design — named above and not measured, because it needs a trial size sweep against a basis size sweep and the interesting part of it is where the randomisation dies rather than where the balance stops improving; splines and other bases that are not polynomials or indicators; and the analysis, which is a separate lever and the subject of the last essay of this field.
The boundary against the design fields is the currency. A maximin design spends runs at settings and pays for protection in efficiency; a maximin rule spends assignments and pays for protection in a few points of variance. Everything about why a worst case is the right thing to optimise, and why its optimum sits on a tie, is established there and used here without being re-derived.
The checks, and the refusals that make them mean something
Two claims are gated in this field’s library. Reading three functions is required to cost less than six points against the shape one function was built for, and to have the best worst case across the three shapes by a wide margin — the two halves of the insurance argument, which would fail in opposite directions if the trade were being described wrongly. And the rule that reads the mean alone is required to have a worst case no better than a coin’s, which is the exposure this essay exists to repair and is measured on the same trials.
The refusals for this field are the two in the essays on either side. The one that bears on this essay is the rule reported on the covariate it reads: every rule in the table above removes nearly all of the imbalance in x, including the ones with a worst case at a coin’s, so the balance table that a trial reports is exactly as flattering for the best row as for the worst.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- Where the guarantee is exactly zero — both name basis functions, covariate balance, design criterion, maximin design, model misspecification, robustness, threshold
- A zero that was an assumption — both name basis functions, covariate balance, design criterion, imbalance, maximin design
- Two contrasts, one split — both name allocation rule, efficiency, maximin design, treatment effect, variance reduction
- What the extra function buys — both name basis functions, covariate balance, design criterion, maximin design, robustness
- The arm whose variance is its answer — both name allocation rule, efficiency, treatment effect, variance reduction
- A covariance with no parameter in it — both name efficiency, model misspecification, nuisance parameter
Named objects
A flat tag is an object no other essay names yet.
Allocation ruleBasis functionsCovariate balanceDesign criterionEfficiencyImbalanceInformation matrixMaximin designModel misspecificationNuisance parameterRobustnessThresholdTreatment effectVariance reduction