Concept

Basis functions — where it appears

The functions of a covariate a rule is handed, whose span — not the list itself — is what the rule can balance. What a rule removes of an outcome shape is that shape's squared multiple correlation on the span, so two lists spanning the same space are the same rule.

Named by 25 essays across 10 fields — each of them below, with the objects they name alongside it.

The six best bases of 2 functions, and what each protects. Every cell is R²(g | span B) — the share of the imbalance in that shape a rule balancing that basis removes — computed from exact inner products between Hermite functions and indicators, with nothing simulated. The rows are ordered by their worst cell, which is the number an experimenter who does not know the shape is exposed to. The best row here guarantees 26.8% against every shape in the list, and the worst of the six guarantees 15.1%: the difference between them is entirely which subspace was picked, at the same cost per arrival.

A basis is a subspace

A balancing rule cannot tell one basis from another with the same span, so choosing what to hand it is choosing a subspace — and then what it removes of any outcome shape is a projection, computable exactly, with no trial anywhere in it.

basis · Criterion
The expansion that never terminates. The Hermite coefficients of a median split, in magnitude, against the reference j to the power −3/4, anchored at the first one. Every even order is exactly zero because sign is an odd function, and every odd order is not, so no truncation is exact — where a polynomial of degree d is exact at any order past d. Summed, the tail past J falls like 1/√J: sixty orders still leave 6.6% of the variance outside. That statement is what made a cut dictionary's geometry unavailable in closed form, and it is a statement about the function against itself. What it is not is the accuracy of an inner product between two correlated variables, where every term past J carries a factor of ρ^m as well.

A cut is not a polynomial, and it does not have to be

A threshold's expansion never terminates, which is why a balancing dictionary's geometry was closed for powers and taken to draws for cut points. Conditioning on the second variable closes it for both.

splits · Blocking
The rule is parity, and it runs both ways. At a correlation of 0.5, four combinations of a dictionary and an outcome shape. The joint sign flip (X, Y) → (−X, −Y) leaves the bivariate normal alone at every correlation, so a function that changes sign under it is orthogonal to one that does not. A product of two odd functions is even; a product of an odd and an even one is odd. So an odd dictionary removes exactly none of the first and something of the second, and an even dictionary does the reverse — which it does, to machine precision, in both of the two rows that should be zero. This is one rule where there had been two: that a median split's square is constant, and that a polynomial dictionary contains the products a correlation generates.

A dictionary that is neither

A rule handed two median splits removes none of their interaction; a rule handed two covariates removes none of their product. Those were two results with two explanations, and they are one result with one — and finding it corrected the number underneath both.

dict · Criterion
One zero is arithmetic and one is a symmetry. What two balancing rules remove of the interaction they are aimed at, on five joint laws of the ranks matched at a Spearman correlation of 0.4, with a normal covariate throughout. A rule holding a median split of each covariate removes exactly nothing of the product of the splits under every one of them, including the two that are not symmetric under reflection — and the reason is not a symmetry at all: a centred median split takes the values ±½, so its square is a quarter identically, and the interaction is orthogonal to both main effects whatever the joint law is. A rule holding the mean of each removes exactly nothing under the three radially symmetric copulas and 7.71% under the two that are not. Bars at the floor are exact zeros; the axis cannot draw 9e-32.

A zero that is arithmetic

A median split's exact zero was explained by a symmetry of the latent normal. It holds under a Clayton copula, which has no such symmetry, because a centred median split squares to a quarter identically.

copula · Criterion
One zero holds and one does not. Three rules, at a correlation of 0.5, against the skewness of the covariate. A rule balancing the mean of each covariate removes exactly nothing of their product when the marginal is symmetric — including the heavy-tailed symmetric one at skewness zero, which is what says the guarantee needs symmetry rather than normality — and removes up to 29.7% when it is not. A rule balancing a median split of each removes exactly nothing of the product of the splits under every marginal here, to 1e-30: both sides are functions of the sign of the latent normal, and a monotone transformation moves neither. A rule balancing a threshold at a value on the covariate's own scale removes between 4.9% and 22.5% — it never had a zero to lose, under any marginal at all.

A zero that rests on a symmetry

A balancing rule removes exactly none of an interaction between two odd functions, at every correlation. The argument needs the joint sign flip to preserve the law, and no real covariate is symmetric about anything.

skew · Criterion
The fourth-order expectation, by two routes. Every inner product in an eight-term slice of the dictionary at ρ = 0.5, computed from the linearisation and Mehler's formula and counted from two hundred thousand draws of a correlated pair. The entries that matter are the ones off the main effects: ⟨f(X)u(Y), g(X)v(Y)⟩ is a fourth-order expectation, which the independent-covariate field could not write down. The worst departure is 1.99 standard errors over 36 pairs, measured in each pair's own error because the entries differ in size by two orders of magnitude.

The fourth moment that was missing

Mehler's formula makes the main effects exact at any correlation and stops there, because the interactions need an expectation of four Hermite functions rather than two. A linearisation turns the four into two, and the whole geometry becomes closed again.

joint · Blocking
What is left of a probe after the rule has had it. The share of each dictionary function a rule balancing x, x2, x3, cut0 has already taken, on trials of 14 units, averaged over 100 designs. Four of the eight functions are the basis, so their share is exactly one: a randomisation test run on one of them is asking about a quantity the rule forced to zero, and one of them is the default probe of the field this measurement comes from. The four that are not still read 0.919, 0.873, 0.903, 0.832 — between 0.832 and 0.919 of them is inside the span — against closed-form removed shares of 0.000, 0.692, 0.590, 0.692. At 14 units a rule with four functions in it takes most of anything it is shown.

The part the rule already took

A diagnostic that reports on what a balancing rule was not handed is run through a column that is 92% inside the span the rule balanced — because orthogonality in the population is not orthogonality on fourteen units.

aimed · Randomisation
What each rung is made of. Each pair of searches, over 300 draws, split into the two effects its excess is the difference of. The overlap is what the second search loses by having the first already run at its own answer; the interaction is what the joint search finds by moving the first off it. They subtract to the excess exactly, on every draw, because the pinned supremum cancels. Two disjoint dictionaries of independent columns read an excess of 0.000011 and are made of 0.000514 and 0.000503. A break paired with a dictionary of step columns has an interaction of exactly 0 and is all overlap. And a break paired with an independent column has an overlap of -0.004395 against an interaction of 0.002364, which is what puts its excess below zero.

Two effects in one number

How much two searches over one sample share is measured as the net of two things — ground both of them find, and configurations only the joint search reaches. One extra supremum per draw separates them exactly.

separate · Break point
How far apart the two components are, on each probe. The median separation between the two components of the admissible set — the difference in their mean probe values, over the spread inside a component — over the 100 of 200 designs whose set is enumerated and found split. The separating direction carries 10.565 and needs the enumeration. The fourth power as the earlier fields use it carries 1.543; projected off the span the rule balances, 5.080. The design's own leverage, which uses no dictionary and no outcome, carries 3.836. A random direction in the same subspace carries 0.942, and a direction chosen by looking for concentrated structure carries 0.543 — below random, and the one heuristic here that is worse than not choosing at all.

A probe chosen from the design

The design's own leverage aligns with the separating direction four times better than a random direction in the same subspace. The concentrated direction the argument invites is worse than random.

aimed · Randomisation
What the second search finds, alone and afterwards. For four of the pairs, what the second search removes on its own and what it removes once the first has already run. The gap between the two is the overlap in absolute terms. Where the searches share nothing the two readings are the same: an independent column removes 0.0261 alone and 0.0260 afterwards. Where one contains the other they are 0.1387 and exactly zero. The pair the earlier field measured sits between: a whitening window removes 0.5033 alone and 0.3120 after a break search has run. This is the earlier field's own reading of its pair, on the share scale rather than in log-likelihood units, and it is the number a rule that runs both searches actually has to charge for.

A search that is already the other

A break search shifts every coefficient after a row, so a step column is one of the directions it can move in. Paired with a dictionary of them it reads exactly one, on every draw, and that fixes the top of the scale.

apart · Criterion
Where the two kinds of cut sit. Six covariates, each a monotone transformation of the same latent normal. The vertical line at zero is where every median split sits, on every one of them, because a monotone map preserves order: the median of the covariate is the image of the median of the latent normal. The marks on the curves are where a threshold at 1 on the covariate's scale falls — 1.000, 0.881, 0.875, 0.783, 0.713, 0.337 — and none of them is at zero. That is the whole of the difference. A function of the sign of the latent normal is odd, and a rule made of odd functions removes exactly nothing of an interaction between two of them; a threshold anywhere else is neither odd nor even and removes something.

A split survives what a mean does not

The two things every trial balances come apart on a skewed covariate. A median split is a function of the sign of the latent normal whatever the marginal is; a mean is not, and its exact zero is gone at a skewness of one.

skew · Criterion
The zero was a fact about independence. What a balancing rule handed every main effect of both covariates removes of a pure interaction, as the covariates are allowed to move together. At ρ = 0 it is exactly nothing — at machine precision, at any number of main effects — which is the independent-covariate result and is correct. It is not small anywhere else: the product of the two covariates loses 64.0% of itself by ρ = 0.5, because h₁h₁ = h₀ + √2·h₂ and Mehler pairs h₂ with h₂ at ρ². Four interactions are drawn and none of them keeps the zero.

A zero that was an assumption

A rule handed every main effect of both covariates removes exactly none of a pure interaction. That is true at machine precision, it is a fact about independence, and it dies as the square of the correlation.

joint · Criterion
Sheppard's arcsine, by two routes. Corr(sign X, sign Y) as the covariates' correlation runs from zero to one, drawn twice. One route is a sixty-four-node quadrature of the orthant probability over the correlation — the general construction, which works at any pair of cut points; the other is (2/π) arcsin ρ, which is elementary and works only at the median. They agree to 3.3e-16 at every one of 81 correlations, which is what licenses the quadrature everywhere else. The curve is above the diagonal at small ρ and below it at large: two signs agree with probability ½ + arcsin(ρ)/π, so a correlation of 0.5 gives exactly ⅓ and a correlation of 0.8 gives 0.5903.

The arcsine that closes it, and the error that was overstated

Two median splits of a correlated pair agree with probability ½ + arcsin(ρ)/π, exactly. And the truncation the field was avoiding falls geometrically in the correlation, not algebraically in the order.

splits · Routes
The mean's zero is the copula's symmetry. Five copulas, each at a Spearman rank correlation of 0.4, with a normal covariate throughout — so nothing here is about the marginal, which is the whole of the earlier field. Horizontally: how far the copula's density is from its own reflection through the centre of the unit square, measured rather than read off the family's name. Vertically: what a rule balancing the mean of each covariate removes of their product. The three copulas at zero on the horizontal axis remove exactly nothing, to thirty decimal places. The two that are not symmetric remove 7.71%. A guarantee that held for six marginals turns out to have needed something the marginals could not have told anybody about.

The symmetry the marginals could not show

A mean's interaction zero needs the covariate to be symmetric and the copula to be symmetric under reflection. Six marginals could only ever test one of those, and the other is broken by the commonest kind of dependence there is.

copula · Criterion
Both halves grow; the difference does not. The control pair's two components and their difference, against how much each of its two searches can find, over 1200 draws at each dictionary size. Two disjoint sets of independent columns are additive at every size — the excess stays inside a standard error or two of zero throughout — and it is not because there is nothing there. The overlap grows from 0.000112 at two columns to 0.000870 at ten, a factor of 7.76, and the interaction grows with it, staying within a factor of two of the overlap at every size. Two searches competing for one residual sum share ground and find configurations neither has alone, in almost equal measure, and their difference is what the earlier field's scale calls zero.

What a zero is made of

Two disjoint dictionaries of independent columns read an excess of 0.000116 and are made of an overlap of 0.000583 and an interaction of 0.000467. The control the whole scale is anchored on reads zero because two effects cancel.

separate · Break point
What the worst case is worth, one function at a time. The smallest share each dictionary removes, over seven outcome shapes, at a correlation of 0.5. A rule balancing the mean of each covariate has a worst case of exactly zero — against the square, and against both products. Adding a median split to it, which is the second thing every trial balances, leaves the worst case at exactly zero, because a median split is odd and so is a mean. Adding the square instead moves it to 6.8%, and the extra functions after that move it to 7.4%. The worst case is decided by which parities the dictionary contains rather than by how many functions are in it.

What the extra function buys

A rule balancing the mean of each covariate has a worst case of exactly zero. Adding the median split — the other thing every trial balances — leaves it at exactly zero, and one square moves it.

dict · Criterion
The guarantee, as the basis is allowed more functions. The lower line is the best worst case over the six named shapes for a basis of each size, found by scoring every subset of the dictionary — an exact answer, since the problem is finite. One function guarantees 2.3%, which is nearly nothing; three guarantee 59.0% and the basis that does it is the covariate, its square and its cube, with no indicator in it. The upper line is the same problem with the basis drawn rather than fixed, which is worth 2.09 times as much at two functions and 1.32 at three. The two lines converge because a basis large enough to protect everything has nothing left to randomise over.

Which shapes are worth protecting

Choosing a basis by its worst case is a finite problem with an exact answer. The answer has no tie in it, which a maximin optimum is supposed to have — and the tie comes back, along with twice the guarantee, when the basis is drawn rather than chosen.

basis · Blocking
The split depends on the order. How much two searches share, measured both ways round, over 300 draws. Pinning the first search at its own answer and searching the second gives one overlap; pinning the second and searching the first gives another. A break paired with an independent column reads -0.004395 one way and 0.002163 the other, at 8.70 paired standard errors and on opposite sides of zero. The excess the two components subtract to is the same in both orders by construction, so what changes is only how it is attributed. There is no order-free way to say which of two searches found ground both can reach, and the two orders bracket it.

A split that depends on the order

Run the second search first and pin that instead, and the same draw gives a different overlap and a different interaction — with the same difference. And one pair has no second order at all.

separate · Break point
Whichever dial made the set thin, the crossing is at the same thinness. Each curve is one dictionary, swept over eight tolerances at two hundred units; a point above the line is a set thin enough that walking beats hunting. The curves lie nearly on top of one another, which is the answer to whether the crossing is a fact about the tolerance or about the thinness it produces: the crossings sit between one admissible assignment in 176 and one in 268 for dictionaries of 3 to 6 functions. The mechanism is that a hunt costs exactly 1/p and a walk costs almost the same everywhere — between 82 and 394 evaluations per usable draw across the whole table — so the crossing is wherever 1/p reaches a number that does not move.

The set a dictionary leaves

A rule constrained on six functions at a loose tolerance leaves a set as thin as one constrained on three at a tight one. Both sampling methods cross over at the same thinness, and the tolerance where that happens moves by a factor of three.

dict · Assignment
The guarantee that survives a correlation, and the one that does not. What a balancing rule handed both main effects removes of the pure interaction between them, as the covariates become dependent. For median splits it is exactly zero at every correlation, because sign(x)² = 1: the interaction sign(X)sign(Y) is orthogonal to sign(X) and to sign(Y) whatever ρ is. For the product of the raw covariates it is 4ρ²/(1+ρ²)² — 64.00% by ρ = 0.5, rising to all of it at perfect correlation. A cut away from the median sits between them and is not small: 23.01% at a cut of one. The zero is not a fact about interactions. It is a fact about a dictionary whose functions square to a constant, which a polynomial one does not.

The zero that survives a cut

A rule holding both main effects removes half of a pure interaction between correlated powers and exactly none between correlated median splits. The guarantee that a correlation destroyed was never about interactions.

splits · Criterion
What a rule reads, against what the outcome uses. The variance of the treatment estimate relative to a coin's, for four things a rule might balance against three shapes the outcome might have, over 350 trials of 200 units. The diagonal is the easy part — a rule that reads the function the outcome uses removes about half the variance. What the table is for is the off-diagonal: reading the covariate alone is worth nothing against a quadratic (0.755), and reading all three is worth nearly as much against every shape as the matching rule is against its own (0.532, 0.493, 0.514).

Three functions of one number

A rule that balances the covariate is exposed to every shape the outcome might have. A rule that balances three functions of it costs two points of variance against the shape the first was built for and takes the worst case from a coin's to about half of it.

shape · Criterion
A class that is a subspace has no guarantee below its own dimension. Each cell is the worst case over every unit-variance function in a class of dimension m, for a rule reading k functions: the smallest squared principal-angle cosine between the two subspaces. Wherever k is less than m the number is zero to machine precision, and that is not a weak guarantee but the absence of one — some direction of the class is orthogonal to the entire basis, and against an outcome in that direction the rule does exactly what a coin does. An experimenter who declines to name the shapes and asks instead to be protected against everything smooth is asking for the cells above the diagonal.

Where the guarantee is exactly zero

An experimenter who declines to name the shapes, and asks instead to be protected against anything in a class, is asking for a number that is not small but zero. Bounding the class is unavoidable, and the two ways of doing it choose different bases.

basis · Randomisation
A guarantee that stops being a number. The worst case of each dictionary over six outcome shapes, at a correlation of 0.5, against the skewness of the covariate. Under a symmetric marginal every rule made of odd functions has a worst case of exactly zero, and the rule holding a mean and a median split of each covariate — the two things every trial balances — is one of them. Under skew that zero becomes 0.74%, 1.83%, 2.24%, 2.49%: small numbers, each of which depends on a marginal nobody stated. The guarantee has not improved by becoming positive. It has stopped being a guarantee, because it can no longer be written down without the covariate's distribution in it.

Balancing a skewed covariate

The worst case of the rule every trial runs goes from exactly zero to somewhere between a quarter of a per cent and two and a half. Which is small, and is a number that cannot be stated without the covariate's distribution in it.

skew · Criterion
A thin enough set is not one set. Every admissible set of 14 units this table can enumerate, by how much of the assignment space it admits and how many pieces it falls into under single swaps. A walk is uniform on the piece it starts in and never leaves it. The pieces are not fragments: at 522 admissible assignments the set splits into 3 halves of exactly 520 each, and every assignment's complement is in the other half — no sequence of admissible single swaps takes an assignment to its own mirror image. Two-swap proposals reconnect four of the six disconnected sets here — the two they do not are the thinnest, where a two-unit move rarely lands anywhere admissible either — which makes a bigger proposal a correctness repair rather than the speed dial it was measured as.

The walk that cannot cross

A thin enough admissible set is not one set. It splits into an assignment and its mirror image, no sequence of admissible single swaps joins them, and the walk that samples it is uniform on half the reference distribution for ever.

dict · Randomisation
Where the constraints exhaust the randomisation. At 16 units there are 12,870 equal splits, so the ones meeting a stated tolerance can be counted rather than estimated. With each of the first k standardised imbalances required to be within 0.4 of a coin's own spread, the admissible count runs 3874 → 1006 → 314 → 0 → 0 → 0 — and at 4 functions there is no admissible assignment at all. The count is the number of distinct answers a randomisation test can give: at 3 functions its finest attainable p-value is 1 in 314. Balance improves with every constraint and the reference distribution shrinks with it, and the two run out at different rates.

When the constraints run out

Every function added to a basis is a constraint the assignment has to satisfy with the same units. At sixteen units and a stated tolerance the admissible assignments run 3,874, then 1,006, then 314, then none — and the count is exact, because the assignment space is finite.

basis · Allocation

Named alongside it

The objects these essays reach for when they reach for this one.

Covariate balanceOrthogonalityProjectionInteractionClosed formHermite polynomialsExperimental designRerandomisationThresholdCovariate adjustmentCorrelationDesign criterion

All concepts