Every essay — page 14
What a chain cannot report
A thin admissible set falls into an arrangement and its mirror image, and the walk that samples it is uniform on half the reference distribution for ever while every diagnostic passes. That was found by enumerating fourteen units, and enumeration stops at about twenty-four. The test that does not enumerate is two chains — one from an assignment, one from its complement — compared on a statistic that changes sign under the complement, with a third chain from the same starting point to say whether a large reading is a fact about the set or about the length of the run. It agrees with the enumeration at every tolerance where the answer is known. At two hundred units it reports something else: the walk reaches the whole set at the tolerances a trial would use, and past a point the diagnostic stops agreeing with itself.
A test rather than a survey
A thin admissible set falls into an arrangement and its mirror image, and the walk that samples it is uniform on half the reference distribution for ever. That was found by enumerating fourteen units, and enumeration stops at twenty-four.
The statistic that changes sign
A test for an unreachable half needs a quantity that tells one half from the other. Every symmetric reading of a mirror pair is identical, and a magnitude is the natural thing to reach for.
A defect that is about size
The admitted share of a rerandomisation barely moves with the number of units. The number of admissible neighbours grows like the square of it, and that is what decides whether the walk can go everywhere.
The diagnostic at two hundred
Pointed at a trial size no enumeration reaches, the test gives three answers rather than one — and past a certain thinness it stops agreeing with itself, which is the honest reading and the one nothing could give before.
A guarantee that needed a symmetry
What a balancing rule can remove of an interaction is decided by parity, and parity is a statement about a symmetry of the law rather than about the covariate. Keep the dependence and change the marginals — a Gaussian copula under a monotone transformation — and the two things every trial balances come apart. A median split is a function of the sign of the latent normal whatever the marginal is, so its exact zero survives every transformation to the last digit. A mean is odd only when the marginal is symmetric, and its zero is gone at a skewness of one. The separating case is a marginal that is heavy-tailed and symmetric, where every zero holds exactly: it is not normality the guarantees needed. And a threshold at a value on the covariate's own scale never had one at all.
A zero that rests on a symmetry
A balancing rule removes exactly none of an interaction between two odd functions, at every correlation. The argument needs the joint sign flip to preserve the law, and no real covariate is symmetric about anything.
A split survives what a mean does not
The two things every trial balances come apart on a skewed covariate. A median split is a function of the sign of the latent normal whatever the marginal is; a mean is not, and its exact zero is gone at a skewness of one.
The cut that is not a quantile
A protocol that says split the covariate at a threshold and one that says split it at the median read the same and are different rules. One has an exact guarantee under every marginal and the other has none under any.
Balancing a skewed covariate
The worst case of the rule every trial runs goes from exactly zero to somewhere between a quarter of a per cent and two and a half. Which is small, and is a number that cannot be stated without the covariate's distribution in it.
Where a taper's case begins
A tapered block was measured against a plain one on the exact bias each implies, and the answer was that the taper's advantage has not arrived at any block length a hundred and twenty rows can afford — with the crossing at ℓ = 20 and a difference there of three tenths of a point, too small for the critical-value table to resolve. Two of those three statements are about the wrong quantity. What a sample reports at ℓ = 20 is four points rather than three tenths, because the autocovariances the window is applied to are themselves attenuated and the window that discards the long lags loses less of them; and the comparison is made at a shared block length where each window has its own best one. Read at each window's own setting, on the error rather than on the bias, the ordering reverses at a hundred and twenty rows.
The gap a sample shows
The exact difference between two block windows at a block length of twenty is three tenths of a point. What a hundred and twenty rows report is four and a third, because the autocovariances the window is applied to are attenuated too.
Bias is not the whole of it
A window that reaches zero at its ends attenuates less and uses less of each block. The block length that minimises its bias is not the one that minimises its error, and comparing two windows at one length compares one of them mis-tuned.
Measuring a variance rather than a quantile
A resample's implied long-run variance can be computed from the sample with no resampling in it at all. A critical value cannot, and the difference is a factor of three in the draws before any of the resampling is counted.
The error no window repairs
Every block window's best estimate of a long-run variance is wrong by about forty per cent at a hundred and twenty rows, and the largest part of that is not a bias at all. Choosing the window moves a twentieth of it.
A covariance with no parameter
A regression's coefficients and its errors' dependence can be fitted together when the dependence is one number. When it is an estimated covariance there is nothing for “jointly” to mean — until a family is named, and then the family's own width is the parameter. Three things the measurement says, two of them the opposite of the guess: the plug-in's shortfall is the taper's rather than the data's and is nothing under two of four laws; nothing in the likelihood chooses the width, because nested families buy about a unit of it a lag and that is what a criterion charges; and a fit-only objective does not run away, because a unit diagonal fixes the trace.
A family before a fit
A regression's coefficients and one correlation can be maximised together. Replace the correlation with an estimated covariance and there is nothing left for "jointly" to mean — until a set of covariances is named, and the set turns out not to contain the truth.
The plug-in and the maximum
A tapered covariance estimate sits five and a half log-likelihood units below the maximum of the likelihood it is substituted into. Four fifths of that is what the optimiser would have found if nothing were missing.
Nothing in the fit picks the width
A wider band is always a better fit, and it is better by about one unit of log-likelihood a lag — which is the order of what a criterion charges for a parameter. Three defensible rules choose widths a factor of three apart.
What fitting them together buys
Maximising over the coefficients and the covariance together beats the two-step under one of four dependences and ties under the other three. It is the one the band family contains, and the likelihood said so before any coefficient was compared.
How long the list is
A tuning parameter chosen per candidate rather than once for the table costs something, and the window's figure was measured while the order's was not — because their lists are different lengths and matching them changes what each rule is. Matched at every length, the two cost the same. What separated them was not the list at all: a per-candidate whitening is a different error model for every candidate, the Gaussian likelihood has a term that says so, and the sieve's rule in this collection never carried it.
The comparison that was not made
Choosing a whitening's window separately for every candidate costs 0.00401 of regret. The same question about an order was named and left, because the two lists are different lengths. The order's answer is 0.00360, and matching the lists changes almost nothing.
The volume a whitening moves
A sieve's whitening has a determinant and this collection's criterion for it never carried one. Shared across a table the term cancels exactly, which is why nothing ever noticed; used per candidate it is worth more than a parameter and the whole comparison turns on it.
A list is not a rule
How often five candidates disagree about a tuning parameter runs from nothing at two values on the list to two draws in five at thirteen. What the disagreement costs does not move at all.
Two searches over one sample
A break point that has been looked for costs more than a count of parameters says. So does a window chosen from a list, and nothing here had charged for it. A rule that does both is running two searches over one sample, and the charges do not add: the pair manufactures a fifth less than the two apart. What follows is sharper than a correction to a sum — most of what a break search finds under correlated errors is the correlation, so once a whitening has been chosen from the same sample the break's own charge more than halves, and carrying the published one across switches the test off entirely.
Two searches, one sample
A searched break in a regression manufactures 34.7 of likelihood ratio where a count of coefficients says 11.1. A searched window manufactures 84.0. The two together manufacture 99.4, not 118.7.
The charge that is not a sum
Charging two searches what each costs on its own is conservative, and conservative here means the test never fires. At the largest break measured it declares nothing, on every draw, while a calibrated threshold reaches 29%.
A charge that depends on the rule
The break search's charge is 34.7 on its own and 15.4 once a window has been chosen from the same sample. Most of what a break search finds under correlated errors is the correlation, and a whitening has taken it already.
The diagnostic after the trial
The test for whether a balanced-assignment walk can reach the whole admissible set is run on a covariate function, before any outcome exists. Run instead on the difference in arm means it is the same test — under the sharp null the outcome is a fixed column — and it is about the statistic the p-value is actually built from. Two findings: an outcome is a probe nobody chose, and on a set that is genuinely split three in ten of them see nothing at all; and the published two-sided p-value is exactly right on half a reference distribution, while a one-sided one is not.
The statistic the p-value is about
The test for whether a balanced-assignment walk reaches its whole set is run on a covariate function chosen before the trial. Run on the difference in arm means it is the same test, and it is about the number the trial publishes.
A probe nobody chose
On a set that is definitively in two pieces, seven of twenty-four outcomes report nothing at all. Every covariate probe reports it. What separates them is not accuracy — it is that one of them can be chosen and the other is what happened.