Choosing what the rule reads

Where the guarantee is exactly zero

An experimenter who declines to name the shapes, and asks instead to be protected against anything in a class, is asking for a number that is not small but zero. Bounding the class is unavoidable, and the two ways of doing it choose different bases.

Worth reading first: Randomisation is not balance · A design is a number.

The maximin basis is a statement of the form “against any of these six shapes, at least this much”, and the six shapes are the part a statistician cannot supply. The natural response from anybody who has to sign the protocol is to ask for something less arbitrary: protect me against anything smooth, or anything in the span of these, or anything at all.

Each of those is a well-posed question and this essay is what the answers are. The first one is zero.

A class that is a subspace has no guarantee below its own dimension. Each cell is the worst case over every unit-variance function in a class of dimension m, for a rule reading k functions: the smallest squared principal-angle cosine between the two subspaces. Wherever k is less than m the number is zero to machine precision, and that is not a weak guarantee but the absence of one — some direction of the class is orthogonal to the entire basis, and against an outcome in that direction the rule does exactly what a coin does. An experimenter who declines to name the shapes and asks instead to be protected against everything smooth is asking for the cells above the diagonal.
Fig. 1 The worst case over every unit-variance function in a class of dimension m, for a rule reading k functions. Above the diagonal the number is zero, and it is zero rather than small.

A class that is a subspace

Suppose the class of outcomes to be protected against is not a list but a subspace: every unit-variance function in the span of m stated functions. This is what “protect me against any combination of these” means, and it is a natural request — an experimenter who is willing to say the outcome depends on the covariate, its square and a threshold, but not to say which, is asking for exactly this.

The worst case is a principal-angle problem with a closed form. The class’s Gram matrix whitens it, the basis’s projection is a matrix in the same coordinates, and the guarantee is the smallest eigenvalue of the projected Gram against the full one. There is a direction of the class the basis explains least, and that direction is the worst outcome.

Where the basis has fewer functions than the class has dimensions, that eigenvalue is exactly zero. Not small — zero, to machine precision, because some direction of an m-dimensional space is orthogonal to a k-dimensional one whenever k < m. Against an outcome in that direction, the rule does precisely what a coin does.

The table has the shape of a staircase. A two-dimensional class is protected at 1.0000 by two of the right functions and at 0 by any one; a three-dimensional class at 1.0000 by three and at 0 by two; and so on down the diagonal.

Six shapes, and the span of six shapes

The consequence for the previous essay’s answer is sharper than it looks, and it is worth stating as a sentence somebody might otherwise say without noticing.

The maximin basis of two functions guarantees 26.8% against each of the six named shapes. Against the span of those same six shapes it guarantees zero.

Those are two different promises and they sound like the same promise. “Protected against a threshold, a quadratic, a cubic and three cut points” is true. “Protected against anything built out of a threshold, a quadratic, a cubic and three cut points” is worth nothing at all, because the combination that defeats the basis is in the span and nothing in the list rules it out.

An experimenter who believes the outcome is one of six named things is in the first situation. An experimenter who believes it is some mixture of six named things is in the second, and their guarantee has evaporated without any of the six changing.

The six best bases of 2 functions, and what each protects. Every cell is R²(g | span B) — the share of the imbalance in that shape a rule balancing that basis removes — computed from exact inner products between Hermite functions and indicators, with nothing simulated. The rows are ordered by their worst cell, which is the number an experimenter who does not know the shape is exposed to. The best row here guarantees 26.8% against every shape in the list, and the worst of the six guarantees 15.1%: the difference between them is entirely which subspace was picked, at the same cost per arrival.
Fig. 2 The list the guarantee is over. Six numbers per row, all of them above zero for the best row, and the guarantee over their span is not on this page because it is not a positive number.

How big the blind spot is, not just that it exists

“Some direction is orthogonal” understates it. When the basis reads k functions and the class has dimension m, the directions against which the rule does exactly what a coin does form a subspace of dimension m − k — not one unlucky shape but a whole subspace of them, and every unit-variance combination inside it.

So the fraction of the class carrying no guarantee at all is (mk)/m(m-k)/m. Six functions against a class of dimension twelve leaves half the class unprotected; against dimension seven it leaves a seventh.

The worst case falls off a cliff and the average does not

The same arithmetic prices the maximin framing against the alternative nobody states, which is what a rule delivers against a typical member of the class rather than the worst one.

Take the outcome to be a uniformly random unit-variance function in the m-dimensional class, with the basis spanning k of its directions. The expected share of it the rule explains is exactly k/mk/m.

Set the two side by side at six functions:

  • m = 6: worst case 1.0000, average 1.0000.
  • m = 7: worst case 0, average 0.857.
  • m = 12: worst case 0, average 0.500.

One extra dimension in the class takes the worst case from one to zero and the average from one to 0.857. The maximin guarantee is discontinuous in the size of the class being protected against; the average is not, and it degrades gently.

That is the whole reason the staircase looks so alarming and is not, by itself, an argument against the rule. It is an argument about which of the two questions has been asked. An experimenter who genuinely faces an adversary choosing the outcome after seeing the design should read the zeros. One who faces an outcome they merely cannot name in advance is being told a number about a direction nobody has any reason to be in.

Which is why “anything smooth” cannot be answered

The same expression settles the request the essay opens with, and it settles it twice over.

A smoothness class is infinite-dimensional, so m=m = \infty. The worst case is zero for every finite k, which is the essay’s answer — and the average k/mk/m is zero as well, for every finite k.

So the failure there is not the maximin framing being pessimistic. Against a class that large, a finite basis explains nothing on average as well as nothing in the worst case, and the request has to be changed rather than answered. That is what a norm ball does: it stops asking about a subspace and starts asking about a set whose members’ roughness is bounded, and a bounded roughness is what makes a finite basis worth anything at all.

Where a class of the same dimension is not the problem

The staircase has an interior, and it is worth seeing that the guarantee is a real number there rather than only ever zero or one.

Take the covariate, a cut at 1 and a cut at 2 as the class — three dimensions — and hand the rule the covariate, its square and its cube, which is also three. The three eigenvalues are 0.1295, 0.7008 and 1.0000: one direction of the class is nearly reached, one is well reached, and one is barely touched at all. The guarantee is 12.95%.

Take three cut points as the class instead and the eigenvalues are 0.0311, 0.6188 and 0.8920 — worse, because a polynomial basis reaches cut points last.

Swap one class member for something the basis spans — the covariate, its square and a cut at 2 — and the eigenvalues become 0.3242, 1.0000 and 1.0000: two directions are exactly reached and the guarantee rises to 32.4%.

So a class of the same dimension as the basis gives a genuine number, and the number is a statement about angles between two three-dimensional spaces rather than about either one of them. What it is not is close to one. Matching the dimension is not the same as spanning the class, and it is easy to think it is.

Why zero and not merely small

The exact zero deserves one more paragraph, because “the guarantee is zero” is the kind of statement that sounds like an artefact of an idealisation and is not.

The class is every unit-variance function in the span. The basis reaches a k-dimensional subspace of the m-dimensional class. A subspace of dimension k inside a space of dimension m > k has an orthogonal complement inside that space of dimension m − k, and every unit vector in the complement is a member of the class with R² exactly zero. There is nothing approximate anywhere in that sentence.

What is an idealisation is that the rule balances its span perfectly, and the identity’s own essay measures the departure: a real rule on four hundred units leaves about 1.0539 of a coin’s imbalance where the projection says one, because a constraint that removes nothing still competes for the assignments. So on a finite trial the worst case is not zero but slightly worse than a coin — which is not an improvement on the statement.

The distinction that matters is between the guarantee is weak and the guarantee is that the design does nothing, and it changes what an experimenter should do. A weak guarantee is an argument for a larger trial. A zero guarantee is an argument for naming the shapes, because no amount of data moves it.

What the projection predicts and what a rule on real units leaves. Each line is one basis: four shapes scored against it, with the closed-form R² on the horizontal axis and the imbalance a rule on 400 units actually leaves on the vertical. The identity says the points should sit on 1 − R², and they sit on a line from α at the left to β at the right instead. Both departures are the finite sample: α is above one because a constraint that removes nothing still competes for the assignments (1.05, 1.15, 1.08, 1.14 as the basis grows), and β is above zero because the rule balances its own columns approximately (-0.032, 0.020, 0.024, 0.039). The slope is the geometry and it is the same in every row.
Fig. 3 The finite-trial version of the same statement: at R² = 0 the rule is at or slightly worse than a coin, and no trial size moves that point towards one.

The other way to bound a class

If a subspace gives zero and a list feels arbitrary, the third option is a class that is infinite but bounded by decay: every function whose Hermite coefficients satisfy Σ j²ˢ c_j² ≤ 1, which says the outcome may have content at any order provided the high orders are small.

This has a closed form too, and a very simple one. The worst function puts all of its mass on the cheapest excluded order, so the worst unremoved variance is the largest j⁻²ˢ over the orders the basis omits — which means the maximin basis is the first k Hermite functions, for every k and every s, and the guarantee is (k + 1)⁻²ˢ.

At s = 1 the guarantee runs 0.2500, 0.1111, 0.0625 and 0.0400 as the basis grows from one function to four. At s = 2 it runs 0.0625, 0.0123, 0.0039 and 0.0016.

Two things follow that are worth having.

The polynomial basis arrives without naming a single shape. The previous essay found the covariate, its square and its cube as the best triple against six named shapes; here the same basis falls out of a smoothness bound with no shape list anywhere. Two quite different arguments reaching the same answer is the strongest thing this field has to say about polynomials.

And the guarantees are small. 6.25% at three functions and s = 1 is a much weaker promise than the 59.0% the shape list gives, and it should be: it is a promise about a vastly larger class. A guarantee that covers more covers each thing less.

The other way to bound the class, and what it does not cover. A smoothness class — every outcome whose Hermite coefficients satisfy Σ j²ˢcⱼ² ≤ 1 — has a worst case with a closed form: the largest j⁻²ˢ among the orders the basis omits. So the maximin basis is the first k Hermite functions for every k and every s, and the guarantee is (k + 1)⁻²ˢ: 25.0% at k = 1, 11.1% at k = 2, 6.3% at k = 3 for s = 1. That is a polynomial basis chosen without naming a single shape — and the lines here are why it is not the end of the argument. A threshold's own smoothness functional does not converge, so a threshold is in no smoothness class at all and the guarantee says nothing about it.
Fig. 4 The smoothness functional of a threshold, as more Hermite orders are counted. It does not converge, which is why a smoothness class does not contain a threshold.

The shape a smoothness class does not contain

Here is where the second route stops being a substitute for the first, and it is the cleanest result in this essay.

A threshold is not smooth in this sense. Its Hermite coefficients decay far too slowly: at the median they are 0.7979, 0, −0.3257, 0, 0.2185, 0, −0.1686 — an alternating sequence over the odd orders that shrinks like a small power of j rather than exponentially. So Σ j² c_j² does not converge. Counted over the first five orders it is 2.79; over ten, 5.74; over twenty, 15.70; over forty, 43.63; over eighty, 122.28. It keeps growing, and it grows without bound.

The same is true of a threshold at 1, whose partial sums run 2.85, 6.73, 19.40, 50.01 and 140.11.

So a smoothness class does not contain a threshold at any radius, and a basis chosen to be maximin over a smoothness class carries no guarantee at all about the shape a mean-balancing rule is famously helpless against. The two bounded classes are not nested and neither dominates: the shape list covers thresholds and says nothing about anything not in it, and the smoothness class covers everything smooth and says nothing about a step.

A quadratic’s smoothness functional is 4. A threshold’s, counted to eighty orders, is 122.28 and climbing. That ratio is what the word “smooth” is doing in the sentence.

The other way to bound the class, and what it does not cover. A smoothness class — every outcome whose Hermite coefficients satisfy Σ j²ˢcⱼ² ≤ 1 — has a worst case with a closed form: the largest j⁻²ˢ among the orders the basis omits. So the maximin basis is the first k Hermite functions for every k and every s, and the guarantee is (k + 1)⁻²ˢ: 25.0% at k = 1, 11.1% at k = 2, 6.3% at k = 3 for s = 1. That is a polynomial basis chosen without naming a single shape — and the lines here are why it is not the end of the argument. A threshold's own smoothness functional does not converge, so a threshold is in no smoothness class at all and the guarantee says nothing about it.
Fig. 5 Both cut points, and both diverging. The lower line is the threshold at 1 counted over fewer orders, which is the same statement about a different cut.

How much of a threshold a finite basis can ever reach

The divergence has a companion measurement that says the same thing in the units the rest of this field uses.

Summing the squared Hermite coefficients of a median split over the first eighty orders accounts for 0.9432 of its variance, and over two hundred orders 0.9610. So a basis of the first eighty Hermite functions — a basis nobody could balance, on a trial nobody could run — would still leave nearly six per cent of a median split unreached.

A basis of three reaches 0.7427 of it. Going from three functions to eighty buys about twenty points, and the last four per cent is not available at any finite basis size.

That is the sense in which a step is a hard shape for this whole apparatus: not that it is unreachable, but that its tail is thick enough that the return on additional functions falls away long before the shape is covered.

The six best bases of 3 functions, and what each protectsEvery cell is R²(g | span B) — the share of the imbalance in that shape a rule balancing that basis removes — computed from exact inner products between Hermite functions and indicators, with nothing simulated. The rows are ordered by their worst cell, which is the number an experimenter who does not know the shape is exposed to. The best row here guarantees 59.0% against every shape in the list, and the worst of the six guarantees 35.5%: the difference between them is entirely which subspace was picked, at the same cost per arrival.linearquadraticcubiccut at 1cut at 2median cutx + x² + x³1.001.001.000.660.590.740.590x² + x³ + 1{x>0}0.711.001.000.430.561.000.430x² + 1{x>0} + 1{x>2}0.721.000.450.411.001.000.411x + x³ + 1{x>2}1.000.391.000.461.000.740.390x + x³ + 1{x>−1}1.000.391.000.550.380.760.3811{x>0} + 1{x>1} + 1{x>2}0.780.410.361.001.001.000.355worstthe rule readsclosed-form projections, no simulationordered by the worst cell, which is the guarantee
Fig. 6 Triples again, with the thresholds’ columns visibly the hard ones. Three polynomial orders reach between 59% and 74% of the three cut points and all of the three polynomial shapes. Drag the basis size and watch which columns are the last to fill.

Where the two bounds disagree about what to buy next

The two bounded classes do not merely give different guarantees; they give different advice, and the disagreement is sharpest at the margin.

Under a smoothness bound the next function to add is always the next Hermite order, because the guarantee is the largest omitted j⁻²ˢ and the largest omitted order is the smallest one. The sequence is fixed in advance and does not depend on anything about the problem: x, then x², then x³, then x⁴, for ever.

Under the shape list it is not. Going from three functions to four, the maximin basis adds an indicator at 2 to the three polynomials rather than a fourth polynomial order — because the binding shape at three functions is the threshold at 2, at 59.0%, and a fourth polynomial order moves it much less than the indicator itself does. The guarantee goes to 69.8% and the basis is no longer a polynomial basis.

So an experimenter following the smoothness prescription and an experimenter following the shape list agree exactly up to three functions and diverge at four, and the divergence is in the direction the smoothness class is blind in. That is the concrete form of “the two classes are not nested”: they recommend the same thing for a while and then one of them stops noticing the shape the other was worried about.

The other way to bound the class, and what it does not cover. A smoothness class — every outcome whose Hermite coefficients satisfy Σ j²ˢcⱼ² ≤ 1 — has a worst case with a closed form: the largest j⁻²ˢ among the orders the basis omits. So the maximin basis is the first k Hermite functions for every k and every s, and the guarantee is (k + 1)⁻²ˢ: 25.0% at k = 1, 11.1% at k = 2, 6.3% at k = 3 for s = 1. That is a polynomial basis chosen without naming a single shape — and the lines here are why it is not the end of the argument. A threshold's own smoothness functional does not converge, so a threshold is in no smoothness class at all and the guarantee says nothing about it.
Fig. 7 And the reason, once more, in the only picture that shows it: the shape the smoothness class cannot contain, with its own bound running away.

What an experimenter has to choose between

The three options are now all measurable and the choice between them is a choice about what kind of promise is wanted.

A list. Strongest guarantee, at 59.0% for three functions, and it holds only for what is on the list. Requires subject-matter judgement and makes it explicit.

A subspace. No guarantee at all below the class’s own dimension, and a real one at or above it. The honest use of this is not as a class to protect but as a warning: if the plausible outcomes really do span a space larger than the basis, no basis of that size is doing anything for the worst case.

A smoothness bound. A guarantee over an infinite class, at 6.25% for three functions, choosing the polynomial basis with no judgement required — and silent about thresholds, which is the shape the whole exposure was discovered on.

None of them dominates and the arithmetic cannot pick. What it can do is stop the third being offered as a way of avoiding the first, which is the move this essay exists to price.

A class that is a subspace has no guarantee below its own dimension. Each cell is the worst case over every unit-variance function in a class of dimension m, for a rule reading k functions: the smallest squared principal-angle cosine between the two subspaces. Wherever k is less than m the number is zero to machine precision, and that is not a weak guarantee but the absence of one — some direction of the class is orthogonal to the entire basis, and against an outcome in that direction the rule does exactly what a coin does. An experimenter who declines to name the shapes and asks instead to be protected against everything smooth is asking for the cells above the diagonal.
Fig. 8 And the staircase once more, because it is the one fact here that is easy to state and easy to mis-state: below the diagonal, a real number; above it, zero.
The guarantee, as the basis is allowed more functions. The lower line is the best worst case over the six named shapes for a basis of each size, found by scoring every subset of the dictionary — an exact answer, since the problem is finite. One function guarantees 2.3%, which is nearly nothing; three guarantee 59.0% and the basis that does it is the covariate, its square and its cube, with no indicator in it. The upper line is the same problem with the basis drawn rather than fixed, which is worth 2.09 times as much at two functions and 1.32 at three. The two lines converge because a basis large enough to protect everything has nothing left to randomise over.
Fig. 9 The guarantee a list gives, for comparison with the smoothness bound’s 25%, 11.1%, 6.25% and 4%.

What is claimed here, and what is not

This essay takes what happens when the class is not a list, and the claims are the exact zero below the diagonal, the interior eigenvalues for three classes of dimension three, the closed form (k + 1)⁻²ˢ for the smoothness guarantee, and the divergence of a threshold’s smoothness functional.

What stays out and is named as a decision: an ellipsoid class centred somewhere other than the origin, which is what a prior guess about the outcome shape would give and which turns the whole problem into a Bayes one; classes defined by a bound on a derivative rather than on Hermite coefficients, which are the usual nonparametric objects and do not have a closed form against this criterion; and any statement about how large s should be, which is the same unanswerable question as which shapes to list wearing different clothes.

The boundary against the previous essay is what is being maximised over. There the class is finite and the optimisation is over bases; here the class is infinite and the optimisation is over both, with the worst case computed rather than enumerated. What both essays share is that the answer is a projection, which is the field’s opening identity.

The checks, and the refusals that make them mean something

Two claims are gated in this field’s library. A basis of fewer functions than the class it is asked to protect is required to guarantee exactly zero — to machine precision rather than to a tolerance, since “very small” and “zero” are different claims and only one of them is true. And the smoothness guarantee is required to equal (k + 1)⁻²ˢ at six combinations of k and s, with the threshold’s smoothness functional required to keep growing as more orders are counted, which is what says the two bounded classes are genuinely different rather than one being a version of the other.

The refusal for this essay is a guarantee over a list offered as a guarantee over the list’s span. The best pair guarantees 26.8% over six named shapes and 0 over the space they span, and no examination of the design tells the two apart — only the sentence describing what it protects does.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Basis functionsClosed formCovariate balanceDesign criterionEigenvalueHermite polynomialsMaximin designModel misspecificationNonparametricOrthogonalityProjectionRobustnessSmoothnessThreshold