Theme

The thread: Reversals that are nobody's mistake

Simpson's paradox, regression to the mean and the winner's curse all arise from correct arithmetic applied honestly. Each is shown as a region of a parameter space rather than one famous table, so the questions of how often and how large have answers.
The six best bases of 2 functions, and what each protects. Every cell is R²(g | span B) — the share of the imbalance in that shape a rule balancing that basis removes — computed from exact inner products between Hermite functions and indicators, with nothing simulated. The rows are ordered by their worst cell, which is the number an experimenter who does not know the shape is exposed to. The best row here guarantees 26.8% against every shape in the list, and the worst of the six guarantees 15.1%: the difference between them is entirely which subspace was picked, at the same cost per arrival. Choosing what the rule reads

A basis is a subspace

A balancing rule cannot tell one basis from another with the same span, so choosing what to hand it is choosing a subspace — and then what it removes of any outcome shape is a projection, computable exactly, with no trial anywhere in it.

Four blocks, and only the ends matter. The weights a block carries, normalised so that the resample keeps the residuals' variance and only their covariances are attenuated. What separates these shapes, for everything that follows, is the value at the two ends and nothing about the middle: the first-order attenuation is −(w(0)² + w(1)²)/(2∫w²), which is -1.0000 for the rectangle, -0.3896 for the trapezoid cut off at half height, and exactly zero for both windows that reach the axis. The half-height trapezoid is in the table to be the case that separates a shape from a boundary value: it is smooth, it is tapered, and it buys none of the order the other two buy. A block, weighted inside itself

A block weighted inside itself

The triangle every block resample attenuates by is not a fact about blocks. It is the self-convolution of a rectangle, and a block weighted down towards its own ends has a different one — whose leading term is the squared value at the two ends and nothing else about the shape.

The profile a break point is chosen from. One sample of 120 rows under a break in the persistence, fitted as two first-order regimes at every admissible break point. The maximum is at row 78, where the true break is at 60. The shaded band is every break point within two log-likelihood units of the best one — 8 of the 73 positions searched, which is 11% of the range. The horizontal line is the one-regime fit the search is compared against; the whole profile is above it, at every position, which is the point: a maximum over 73 candidates is above the null by construction and not by evidence. Paying for a search

A break that was looked for

A two-regime whitening finds its change point by maximising a profile, and then reads a criterion that counts parameters. Under no break there is no parameter to count, because every position describes the same model.

The price of each thing the rule is not told. What each rule gives up against the best model available, at a persistence of 0.85 on a fifteen-candidate table, over 400 draws. Reading down: least squares with the ordinary penalty; the whitening at the true ρ; the same at a ρ̂ estimated per candidate; that rule with the term the Gaussian likelihood carries and it omits; a Bartlett-tapered Ω̂ estimated once from the fullest candidate at L = 8; the same estimated per candidate; and the truncated Ω̂, which exists on only 45.0% of draws and is averaged over those. Knowing ρ recovers 89.9% of what counting rows gives up, estimating it 83.4%, and estimating a whole covariance 75.6%. Estimating the dependence, not naming it

A covariance with no parameter in it

The whitening that repairs a criterion is told the dependence is a first-order autoregression and left to find one number. A real dependence is not one number, and the obvious estimate of it is not a covariance matrix.

The expansion that never terminates. The Hermite coefficients of a median split, in magnitude, against the reference j to the power −3/4, anchored at the first one. Every even order is exactly zero because sign is an odd function, and every odd order is not, so no truncation is exact — where a polynomial of degree d is exact at any order past d. Summed, the tail past J falls like 1/√J: sixty orders still leave 6.6% of the variance outside. That statement is what made a cut dictionary's geometry unavailable in closed form, and it is a statement about the function against itself. What it is not is the accuracy of an inner product between two correlated variables, where every term past J carries a factor of ρ^m as well. A cut point, at a correlation

A cut is not a polynomial, and it does not have to be

A threshold's expansion never terminates, which is why a balancing dictionary's geometry was closed for powers and taken to draws for cut points. Conditioning on the second variable closes it for both.

The dependence, at four removes. Under AR(1) at 0.8, four different sequences all called the dependence. The top line is the law. The middle line is what a sample of 120 errors reports on average — computable exactly, because the expectation of a sample autocovariance is arithmetic once the covariance is known. The lower line is what a candidate's residuals report, which is what every two-step rule in this collection actually reads: a fit removes variance, and it removes more of the persistent part than of the rest. At the first lag the three are 0.800, 0.7773 and 0.7338. The dots are counted from draws and share no arithmetic with the line they sit on; the worst departure is 1.2 standard errors. Fitted together, or fitted after

A dependence fitted with the line

Every whitening in this collection reads the dependence off a set of residuals, and residuals are not errors. Fitting the two together recovers most of what that costs, and changes almost nothing about the decision it feeds.

Four dependences a single parameter cannot tell apart. Every law here is standardised to a lag-one autocorrelation of 0.8, so a rule told the errors are a first-order autoregression finds the same number in all four and has no way of seeing what separates them. The geometric decay is the world in which estimating a covariance rather than naming it was priced, and found to cost. The five-period moving average has 0.200 at the fourth lag and exactly nothing past it, where the geometric law says 0.328 at the fifth. Long memory at d = 4/9 is still at 0.576 by the twentieth lag, where the geometric law has reached 0.012. The break has no autocorrelation function at all: what is drawn for it is the average over the pairs at each gap, which is what a stationary estimate converges to. The shape a dependence has

A dependence with a shape

Four ways for errors to repeat, all with the same first lag and nothing else in common. A rule told the errors are a first-order autoregression finds the same number in all four, and is right about one of them.

Two covariates make the dictionary an outer product. Four functions of each covariate, and everything a balancing rule may be handed. The margins are the 8 main effects and the block between them is the 16 interactions, which are 66.7% of the dictionary. Every inner product in it is closed form — ⟨f₁g₁, f₂g₂⟩ = ⟨f₁,f₂⟩⟨g₁,g₂⟩ when the covariates are independent — so nothing about the geometry gets harder. What gets harder is the counting: choosing k of 24 is C(24, k), which is 10,626 at four and 735,471 at eight. When the set is too large to walk

A dictionary that is a product

Two covariates make what a balancing rule may read an outer product — eight main effects and sixteen interactions — and every inner product in it is still closed form. What a rule holding all eight main effects removes of a pure interaction is not small. It is zero.

The rule is parity, and it runs both ways. At a correlation of 0.5, four combinations of a dictionary and an outcome shape. The joint sign flip (X, Y) → (−X, −Y) leaves the bivariate normal alone at every correlation, so a function that changes sign under it is orthogonal to one that does not. A product of two odd functions is even; a product of an odd and an even one is odd. So an odd dictionary removes exactly none of the first and something of the second, and an even dictionary does the reverse — which it does, to machine precision, in both of the two rows that should be zero. This is one rule where there had been two: that a median split's square is constant, and that a polynomial dictionary contains the products a correlation generates. What a dictionary buys and what it costs

A dictionary that is neither

A rule handed two median splits removes none of their interaction; a rule handed two covariates removes none of their product. Those were two results with two explanations, and they are one result with one — and finding it corrected the number underneath both.

The stopping rule costs more than the weighting does. Coverage over 2000 runs of a trial whose variance ratio drifts by a factor of twenty, at three ways of deciding when to stop. Twelve blocks fixed in advance is the top line and reproduces what a trial of fixed length delivers. Stopping when the reported interval is short enough is the bottom line, and it costs between 3.0% and 5.5% of coverage — including for the rule that is told every block's true ratio, which is what says the shortfall belongs to the stopping and not to the weights. Stopping on a width predicted from the within-arm sums of squares is the middle line, and it is back at the fixed-length values. The standard error on each point is 0.49%. When a fixed width is reached

A width the trial has to stop for

The weighting that covers at 94.9% on twelve blocks covers at 91.5% when the trial stops as soon as its interval is short enough — and so does the rule that is told every block's true variance ratio. The shortfall is the stopping, not the weights.

One zero is arithmetic and one is a symmetry. What two balancing rules remove of the interaction they are aimed at, on five joint laws of the ranks matched at a Spearman correlation of 0.4, with a normal covariate throughout. A rule holding a median split of each covariate removes exactly nothing of the product of the splits under every one of them, including the two that are not symmetric under reflection — and the reason is not a symmetry at all: a centred median split takes the values ±½, so its square is a quarter identically, and the interaction is orthogonal to both main effects whatever the joint law is. A rule holding the mean of each removes exactly nothing under the three radially symmetric copulas and 7.71% under the two that are not. Bars at the floor are exact zeros; the axis cannot draw 9e-32. The other half of the dependence

A zero that is arithmetic

A median split's exact zero was explained by a symmetry of the latent normal. It holds under a Clayton copula, which has no such symmetry, because a centred median split squares to a quarter identically.

One zero holds and one does not. Three rules, at a correlation of 0.5, against the skewness of the covariate. A rule balancing the mean of each covariate removes exactly nothing of their product when the marginal is symmetric — including the heavy-tailed symmetric one at skewness zero, which is what says the guarantee needs symmetry rather than normality — and removes up to 29.7% when it is not. A rule balancing a median split of each removes exactly nothing of the product of the splits under every marginal here, to 1e-30: both sides are functions of the sign of the latent normal, and a monotone transformation moves neither. A rule balancing a threshold at a value on the covariate's own scale removes between 4.9% and 22.5% — it never had a zero to lose, under any marginal at all. A guarantee that needed a symmetry

A zero that rests on a symmetry

A balancing rule removes exactly none of an interaction between two odd functions, at every correlation. The argument needs the joint sign flip to preserve the law, and no real covariate is symmetric about anything.

Four rules, three shapes, and no ordering that survives. The variance of the unadjusted treatment estimate under each rule, as a fraction of the variance a coin gives, over 450 trials of 200 units each. Against a covariate that enters linearly the rule that reads the number nearly halves it. Against a threshold at 1 it removes about a fifth. Against a quadratic every rule here is at or worse than a coin — they are all optimising a criterion that is one over the variance of an estimate in a model this outcome does not obey, and a constraint that helps nothing still costs something. Nothing in a trial says which column it is in. The shape the covariate enters by

Balanced on the wrong function

A rule that reads a covariate's numbers halves the variance of the treatment estimate, if the covariate enters the outcome as a straight line. If it enters as a threshold the rule is worth a fifth of that, and if it enters as a curve every rule here is worse than a coin.

Which allocations reverse the overall comparison. The per-group success rates are held fixed; only the split of each group between treatment and control changes. 32% of the allocations reverse, and the worst reverses by 13.1 percentage points. Reversals that are not errors

Simpson's reversal is a region, not a table

The treatment wins in both groups and loses overall. That is normally shown with one famous table, which cannot answer the two questions a reader has — how often, and how large. Swept, it turns out to occupy 31% of the allocation space.

The same study read three ways. At time 2 the truth is 0.497. Kaplan–Meier gives 0.532; dropping the censored subjects gives 0.180; treating the censoring time as the event time gives 0.392. Both naive readings understate survival, because the subjects they mishandle are the ones doing well. When the data stops early

The data that stops early

A subject still event-free when a study ends is not missing and not observed. It is known to exceed something, which is a third state most tools have no slot for — and the two obvious ways of forcing it into one are wrong by 31 and 13 percentage points.

A design that is right once, and one that is never wrong by much. Three designs for the Michaelis–Menten model, scored at every true value of K across a 16-fold range. The peaked curve is the two-point local design built at K = 1: 100% there and 66.7% at the worst point of the range. The curve just under it is the design that averages the criterion over a uniform prior on the same range, which is barely different — 67.9% at worst — because averaging is dominated by the middle of the range where the local design is already good. The flat line is the maximin design: 3 settings, never above 80.8% and never below 78.8%. Its worst case is 12.0 points better, and what it gives up is the 21.2 points at the one value the local design was built for. A design that assumes less

The design for the worst case

A design for a non-linear model is optimal at a guess about the answer. Averaging over a prior repairs that on average; protecting the worst value in a range is a different problem, with a different answer, and it needs a third setting to reach it.

The eighth was not a constant. The probability that a per-candidate tuning list changes which candidate the table selects, at each of 5 separations between the candidates, over 800 draws apiece. The earlier field reports this flat at about an eighth across list length, on a table and a world it never varies. Vary how much the omitted coefficients are worth — one multiplier, with the table, the list, the law and the sample size all held — and it runs from 17.6% to 1.5%, a factor of 11.75. The world in which every candidate is true is the world in which the tuning list decides most; the world in which one candidate dominates is the world in which it decides nothing. What decides whether a tuning list decides

The eighth that was not a constant

How often a per-candidate tuning list changes which candidate wins is reported flat at about an eighth across list length. Vary how far apart the candidates are instead and it runs from 17.6% to 1.5%.

What a sample shows, and what the algebra does. The difference between a rectangular block's implied long-run variance and a trapezoidal one's, as a share of the truth. Above the axis the rectangle is less biased and below it the trapezoid is. The heavy line is exact — computed from the law's own autocovariances — and it crosses at 19.2. The others are what samples of 120, 240, 480, 960 rows report, and every one of them exaggerates whichever window is ahead: at ℓ = 20, where the exact difference is 0.28 points, a sample of 120 rows shows 4.31 points — 15 times larger. That is the number the earlier reading of this comparison was missing: three tenths of a point is what the algebra says and not what a hundred and twenty rows report. Where a taper's case begins

The gap a sample shows

The exact difference between two block windows at a block length of twenty is three tenths of a point. What a hundred and twenty rows report is four and a third, because the autocovariances the window is applied to are attenuated too.

One curve, three designs, and the disagreement is the weights. The compartmental response exp(−θ₁t) − exp(−θ₂t) at the guess θ = (0.2, 1.2), with three designs underneath it. Each mark is a setting and its height is the share of the experiment spent there. The D-optimal design splits the runs equally between two settings — that is the determinant's answer and it is equal for every model of this kind. The design for the first parameter alone puts 1.9% of the experiment at its early setting and the rest at its late one, because the early runs are there to identify the nuisance and nothing more. The design that protects a 4-fold rectangle needs 3 settings and spends 9.1% at the earliest of them. What the procedure may not read

The guess with two numbers in it

Every optimal design for a non-linear model is optimal at a guess. Where the model has one parameter that moves the settings, that guess is a number and everything about it comes out in closed form; where it has two, three constants become functions and one of them becomes zero.

A pair pulled back at 20% of the gap per step. Above, the two series. Below, the difference between them. The gap is pulled back towards zero by 20% of itself each step, so it stays inside a band of 14.3 while the series themselves travel much further. Nothing here is stationary except the difference. The faint line below is the gap for two free walks from the same seed, drawn for comparison. Series that move together

The regression that is not spurious

Two random walks regressed on each other are called significantly related three times in four, so the time-series field ends in a warning. The exception it names and does not measure is here — and when the pair is genuinely tied, the fitted relation converges at rate 1/n rather than the usual 1/√n.

What a mean split leaves, with both halves varying. The share of a mean split's interaction that survives the rule balancing it, at every copula and every marginal, matched at a Spearman correlation of 0.40. The three radially symmetric copulas leave exactly nothing with a symmetric covariate and rise steeply with the skew. The two asymmetric ones start at 7.707% and go opposite ways: the lower-tail copula falls to 0.002% at a skewness of 0.95 — the two failures cancel almost exactly, and a guarantee both fields report as broken is restored — while the upper-tail one climbs to 40.288%. And the heavy-tailed symmetric covariate, which leaks exactly nothing on its own, doubles what the asymmetric copulas leak: 14.229% against 7.707%. Both halves of the dependence at once

Two failures that cancel

A mildly skewed covariate under a lower-tail copula leaks 0.002% of an interaction where each failure alone leaks eight and seven per cent. Turn the copula over and the same pair compounds.

Two searches find some of the same luck. What each search reports on a sample with no break in it, and what the two report together, on four dependences. The dashed line is the sum of the two — what a rule charging each search separately would levy — and the two together always come in below it: 19.30, 15.44, 14.24, 26.64 short, on 100%, 99%, 99%, 100% of draws. The shortfall is not a rounding. Under AR(1) at 0.8 it is 19.30 of the 34.70 the break search manufactures on its own, which is more than half of it. Two searches over one sample are looking at the same noise, and the second one has less left to find. Two searches over one sample

Two searches, one sample

A searched break in a regression manufactures 34.7 of likelihood ratio where a count of coefficients says 11.1. A searched window manufactures 84.0. The two together manufacture 99.4, not 118.7.

How much of one search the other has already found. Five pairs of searches on one sample, on a scale whose zero and one are both fixed by construction. Zero is two searches over disjoint sets of independent columns: they remove shares of the residual sum that add, at 0.8 standard errors from exactly additive, and they read 0.004. One is a break search paired with a step column it contains, which reads exactly one on every draw because the step adds nothing at all. Between them: two dictionaries of step columns cut a few rows apart read 0.125, and the pair the earlier field measured — a break and a whitening window, both reading the same residual series — reads 0.762, three quarters of the way to one search containing the other. And below zero, a break paired with a search over independent columns reads -0.306: the joint search finds configurations neither half of it contains, so charging the two separately under-charges. Two searches over different features

Two searches that share nothing

Two searches over independent columns remove shares of the residual sum that add exactly. On the scale a chi-square point is quoted on they look super-additive by a fifth of a unit, and none of it is overlap.

The same covariate, three ways round. Three worlds over a treatment, an outcome and a covariate, joined by the same three edges at the same three strengths — 0.90, 0.50 and 0.70 — differing only in which way the two edges touching the covariate point. In the first the covariate causes both and adjusting for it recovers the effect of 0.50 exactly. In the second the treatment causes the covariate, the effect is 1.13, and adjusting returns 0.50 — the direct edge alone, with the part that travels through the covariate deleted. In the third the treatment and the outcome both cause the covariate, the effect is 0.50, and adjusting returns -0.087. The regression that produces those three numbers is one formula, and nothing in the data says which panel it is being run in. What conditioning on a variable does

One arithmetic, three decisions

A covariate beside a treatment and an outcome can be a common cause of both, a step on the path between them, or an effect of both. The regression that includes it is the same arithmetic in all three, and it is right in one — returning 0.5000, deleting 0.6300 of the effect, and turning 0.5000 into −0.0872.

One wrong model, four designs, four slopes. The slope a straight line converges to when the truth is a quadratic, under four covariate distributions, by two routes: the population projection in closed form, and the mean of 2500 fitted slopes at 200 rows apiece. The even spread over [0, 2] gives 1.6000 and the same spread moved to [1, 3] gives 2.6000, while widening it to [0, 4] gives 2.6000 — the same number as the shifted one, because a symmetric design's target is the truth's tangent slope at the design's own mean and does not read the spread at all. An exponential spread with the SAME mean as the first gives 2.6000. So two studies of one world, each fitting the same wrong model, honestly report slopes 1.0000 apart, and neither is making an error. A standard error for a model that is wrong

What a wrong model estimates

A straight line fitted to a curved truth converges on the tangent at its own design's mean. Two honest studies of one world, fitting the same wrong model, report 2.600000 and 1.600000, and neither is in error.

Twenty points and one more, at leverage 0.74. Without the distant point the slope is 0.495; with it the slope is -0.389. Its leverage is 0.737 and its Cook's distance is 24.1, against a conventional threshold of 1. Regression, and what the summary hides

The line that one point drew

A single observation among twenty-one reverses the sign of a fitted relationship. Its leverage is known from its x value before the outcome is looked at, so this is a property of the design rather than a surprise in the data.

How often a randomised trial reports the reversal, advantage 6 points. Simple randomisation against randomisation stratified by group, 4,000 trials at each size. The simple design reverses on 3.40% of trials at its worst size and 0.50% at 1280 units; the stratified design reverses on none of them, at any size. When the stratified answer and the pooled one disagree

The reversal a coin cannot prevent

Randomisation removes Simpson's reversal in expectation, which is not the same as removing it. A correctly randomised trial of eighty units, on a population where the treatment helps in both groups, reports it losing overall on 3.40% of trials — and stratifying the randomisation takes that to zero at every size.

One relationship at five designs, residual spread 1.00. Every panel has the same slope of 1, the same intercept of 0 and the same residual standard deviation of 1.00. Only the range of x differs. R-squared runs from 0.021 to 0.849, and the estimated residual spread is 0.9932 in all five. What a summary of a scatter is a property of

R² is a property of the design

One line, one residual spread, five studies that differ only in how far apart they placed their x values. R² runs from 0.021 to 0.849 and the estimated residual spread is 0.993 in every one of them. Nothing about the relationship changed.

What a second positive is worth, prevalence 0.10%. A 90% sensitive, 95% specific test. One positive gives 1.77%. Two independent positives give 24.49%, which is what multiplying the likelihood ratios says. At a correlation of 0.1 between the tests' errors it is 10.16%, and at 0.5 it is 3.16%. Two tests, a threshold, and the rate they are read against

The second test that is not a second opinion

Two positives from a 90/95 test on a one-in-a-thousand condition give a 24.49% chance of disease if the tests are independent. At a correlation of 0.1 between their errors it is 10.16%, and at 0.5 it is 3.16% — barely more than the 1.77% one positive was worth.

Coverage of the Wilson interval against the expected count, at 10, 30, 100 and 1,000 trials. Read against the expected number of successes the four sample sizes draw the same curve near the boundary. The worst coverage is 83.50% at n = 10, 83.71% at n = 30, 83.79% at n = 100, 83.81% at n = 1000, each at an expected count near 0.177, and the limiting depth is e^(−0.1765) = 83.82%. A proportion's interval near the boundary, and the coin

A hole no sample size fills

Wilson's interval is the recommended repair for a proportion, and away from the boundary it wobbles a point or two around 95%. Near zero it has a hole: at an expected count of 0.1765 its coverage is 83.50% at ten trials, 83.79% at a hundred and 83.81% at a thousand, and it never climbs past e to the minus 0.1765, which is 83.82%. The hole is where the interval built on one success stops containing the truth, and it belongs to the count rather than to the sample size.

Which groups partial pooling serves, standard error 1 population width. Pooling's expected squared error for a group, divided by its own mean's, against how far the group truly sits from the centre. It is ×0.25 at the centre and crosses ×1 at 1.732 population widths, beyond which 8.33% of a normal population lies; capping the shift at one standard error holds every group under ×2. What partial pooling does to one group, to the set, and to a ranking

A group from the population's own tail

Partial pooling halves the total squared error when a group's own standard error equals the spread between groups. Every group whose true effect sits more than 1.73 population widths from the centre — 8.33% of a perfectly normal population — does worse than it would have with its own mean, and its loss grows without bound. Among eight groups with the spread estimated, the most extreme is worse off in 61.6% of datasets. Capping the shift at one standard error keeps the total at 0.528 of the unpooled error and holds every group under twice it.

How far apart the two components are, on each probe. The median separation between the two components of the admissible set — the difference in their mean probe values, over the spread inside a component — over the 100 of 200 designs whose set is enumerated and found split. The separating direction carries 10.565 and needs the enumeration. The fourth power as the earlier fields use it carries 1.543; projected off the span the rule balances, 5.080. The design's own leverage, which uses no dictionary and no outcome, carries 3.836. A random direction in the same subspace carries 0.942, and a direction chosen by looking for concentrated structure carries 0.543 — below random, and the one heuristic here that is worse than not choosing at all. A probe chosen rather than picked

A probe chosen from the design

The design's own leverage aligns with the separating direction four times better than a random direction in the same subspace. The concentrated direction the argument invites is worse than random.

What the second search finds, alone and afterwards. For four of the pairs, what the second search removes on its own and what it removes once the first has already run. The gap between the two is the overlap in absolute terms. Where the searches share nothing the two readings are the same: an independent column removes 0.0261 alone and 0.0260 afterwards. Where one contains the other they are 0.1387 and exactly zero. The pair the earlier field measured sits between: a whitening window removes 0.5033 alone and 0.3120 after a break search has run. This is the earlier field's own reading of its pair, on the share scale rather than in log-likelihood units, and it is the number a rule that runs both searches actually has to charge for. Two searches over different features

A search that is already the other

A break search shifts every coefficient after a row, so a step column is one of the directions it can move in. Paired with a dictionary of them it reads exactly one, on every draw, and that fixes the top of the scale.

Where the two kinds of cut sit. Six covariates, each a monotone transformation of the same latent normal. The vertical line at zero is where every median split sits, on every one of them, because a monotone map preserves order: the median of the covariate is the image of the median of the latent normal. The marks on the curves are where a threshold at 1 on the covariate's scale falls — 1.000, 0.881, 0.875, 0.783, 0.713, 0.337 — and none of them is at zero. That is the whole of the difference. A function of the sign of the latent normal is odd, and a rule made of odd functions removes exactly nothing of an interaction between two of them; a threshold anywhere else is neither odd nor even and removes something. A guarantee that needed a symmetry

A split survives what a mean does not

The two things every trial balances come apart on a skewed covariate. A median split is a function of the sign of the latent normal whatever the marginal is; a mean is not, and its exact zero is gone at a skewness of one.

The same copula, turned over. A Clayton copula and its reflection, at the same Spearman correlation of 0.40 and the same Kendall tau of 0.275, against the covariate's marginal. With a symmetric covariate the two are the same number to nine decimals — 7.707% apiece — because the leak then depends on how much asymmetry the copula has and not on which way it points. Skew the covariate and they come apart: at a skewness of 2.26 the lower-tail copula leaves 3.431% and the upper-tail one 36.213%, a factor of 10.6. Both halves of the dependence are asymmetries and an asymmetry has a direction; a lower-tail copula concentrates the dependence where a right-skewed marginal is compressed and the two distortions partly undo each other, and an upper-tail one concentrates it where the marginal is stretched. Both halves of the dependence at once

A symmetry that was not enough

A heavy-tailed symmetric covariate has a skewness of zero and leaks exactly nothing under three copulas. Under the two asymmetric ones it doubles the leak, from 7.707% to 14.229%.

How far each reference distribution's 95% point falls short. Seven constructions on rows that repeat each other, at a block length of 5 and 200 draws, against the statistic's own 95% point of 3.0224 computed from three thousand draws of the same world. Reading down: a multiplier on every row keeps no dependence at all and is 44% short; a multiplier shared along a block keeps the triangle; a fixed block keeps the same triangle and is 8% closer, which is the pair that says a taper is not what decides this; the stationary bootstrap; the two tapered blocks, both further short than the untapered one at this block length; and errors generated from a fitted model, which is the only construction here not bounded by what the residuals report. A block, weighted inside itself

A taper and a critical value

Two constructions whose tapers visibly differ give the same critical value, and two that share a taper exactly do not. Adding a construction whose taper is a decision rather than an accident says which half of that is true.

The zero was a fact about independence. What a balancing rule handed every main effect of both covariates removes of a pure interaction, as the covariates are allowed to move together. At ρ = 0 it is exactly nothing — at machine precision, at any number of main effects — which is the independent-covariate result and is correct. It is not small anywhere else: the product of the two covariates loses 64.0% of itself by ρ = 0.5, because h₁h₁ = h₀ + √2·h₂ and Mehler pairs h₂ with h₂ at ρ². Four interactions are drawn and none of them keeps the zero. When the two are not independent

A zero that was an assumption

A rule handed every main effect of both covariates removes exactly none of a pure interaction. That is true at machine precision, it is a fact about independence, and it dies as the square of the correlation.

Which window is better depends on who chose the block length. The margin between a rectangular block and a tapered one, on 400 samples of 120 rows, under four rules for choosing the block length. At the length that would actually have been best on each draw the taper is ahead by 2.12 points of a 35.8% error, at 22.1 paired standard errors; at a length estimated from the sample's own persistence it is ahead by 1.89. At the length this field's own figures use — eight — the rectangle is ahead by 2.04, and at the rule of thumb by 4.54. Every rule sees the same draws. What separates them is the length: the two rules that lose to the rectangle pick 4.00 and 8.00 where the best available is 24.57, and a tapered window at a quarter of the right length has thrown away most of what it was weighting. A block length chosen from the data

An ordering that depends on the rule

The tapered block beats the rectangular one at the best available block length and at one estimated from the data. At a length written into a protocol, and at the rule of thumb, the rectangle wins — at every sample size measured.

What a longer block buys and what it costs. A trapezoidal block at 120 rows, with the error split into the two things it is made of. The bias falls with the block length, because a longer block attenuates less, and it flattens at 23.2% because the sample's own autocovariances are short whatever window is applied to them. The spread rises with it, because a longer block means fewer of them. Their sum in quadrature has a minimum at ℓ = 16, which is not where either of the two has one. The faint line is the rectangle's total error, for scale: it is above the trapezoid's from ℓ = 12 onwards. Where a taper's case begins

Bias is not the whole of it

A window that reaches zero at its ends attenuates less and uses less of each block. The block length that minimises its bias is not the one that minimises its error, and comparing two windows at one length compares one of them mis-tuned.

Four datasets, slope 0.50, R² 0.67. Every one of these fits reports the same slope to two decimals and the same R². Only the first is a linear relationship with noise: the second is a curve, the third is a line with one outlier, and the fourth has its slope set by a single point. Regression, and what the summary hides

Four datasets, one summary

Four datasets agree on slope, intercept and R² to two decimals. One is a linear relationship, one is a curve, one is a line with an outlier, and one has its slope set by a single point. The summary cannot tell them apart and neither can any other summary.

Coverage against sample size, true proportion 0.15. Coverage does not improve monotonically. n = 19 covers 93.8% while the larger n = 20 covers 81.9%. The sample space is discrete, so the endpoints jump as n changes. Intervals, counted

More data is not monotonically better

Coverage of an interval for a proportion does not improve smoothly as the sample grows. It oscillates, and there are larger samples that cover materially worse than smaller ones — a sample of twenty covers twelve points worse than a sample of nineteen.

The one candidate an effective sample size is right about. n/n_eff with the finite-sample inflation Σ(1 − |k|/n)ρ^|k| is not an approximation to tr(HΩ) for a fit with only an intercept — it is that trace, to machine precision, because the hat matrix of a constant column is 1/n everywhere and its trace against Ω is the mean of Ω. The quoted limit form n(1 − ρ)/(1 + ρ) is not even right about that one. And the average correction the table's fifteen candidates actually need is 3.318 per parameter, well below the scalar, so applying it to all of them over-charges every one. Counting what is independent

One number for a table of candidates

An effective sample size is a real quantity, it is exactly right about one thing, and that thing is a mean. Substituted into Akaike's criterion it changes nothing at all, because the penalty it is meant to fix has no sample size in it.

Stationary is not the same as convergent. How far each k-swap walk is from uniform after t steps, started at the least balanced admissible assignment of 410. Every one of these chains has a symmetric proposal and rejects by standing still, so every one of them is doubly stochastic and every one preserves the uniform distribution exactly. Only five of the six get there. Exchanging all six units of each arm is a single proposal — the complement — and the admissible set is closed under complement, so the walk takes it every time and oscillates between two assignments for ever: after 160 steps it has visited 1 state and sits 0.9976 from uniform. Its stationary distribution is a fact about the matrix; its limit does not exist. What a block may vary

Stationary is not convergent

A walk that exchanges every unit in each arm preserves the uniform distribution exactly and never gets near it. Every doubly stochastic matrix has the same stationary distribution; only some of them have a limit.

Sheppard's arcsine, by two routes. Corr(sign X, sign Y) as the covariates' correlation runs from zero to one, drawn twice. One route is a sixty-four-node quadrature of the orthant probability over the correlation — the general construction, which works at any pair of cut points; the other is (2/π) arcsin ρ, which is elementary and works only at the median. They agree to 3.3e-16 at every one of 81 correlations, which is what licenses the quadrature everywhere else. The curve is above the diagonal at small ρ and below it at large: two signs agree with probability ½ + arcsin(ρ)/π, so a correlation of 0.5 gives exactly ⅓ and a correlation of 0.8 gives 0.5903. A cut point, at a correlation

The arcsine that closes it, and the error that was overstated

Two median splits of a correlated pair agree with probability ½ + arcsin(ρ)/π, exactly. And the truncation the field was avoiding falls geometrically in the correlation, not algebraically in the order.

Every candidate is behind by what its parameter count says. Each dot is one of the fifteen subsets of four predictors, fitted on a rolling window of 80 rows and scored against the benchmark out of sample over 60 origins, at a null where every one of them contains the truth. The line is σ²(q₀/(R − q₀ − 1) − q/(R − q − 1)), which is arithmetic on two integers and a window length. Most of these pairs are not nested — a subset of two predictors and a different subset of two share neither model — and the closed form does not care: the displacement is a statement about how many coefficients each side estimates. The candidates of the benchmark's own dimension sit at zero. A search with no fixed point

The displacement is a parameter count

A nested variant is behind its benchmark out of sample before anything is searched for. The closed form for how far turns out to have nothing about nesting in it — only two integers and a window length — and it prices a table where no candidate contains any other.

What a rule gives away by being predictable. Minimisation run at every probability from a coin to fully deterministic, over 500 cohorts of 120 at each. The upper line is the share of assignments an investigator who knows the rule and the enrolled patients can name in advance: 49.7% at p = 0.5, which is a coin and cannot be beaten, and 87.6% at p = 1 — short of everything only where the two arms tie and the rule falls back on a coin. The lower line is what that is worth: an investigator who enrols a patient 0.5 of a standard deviation better than average whenever they predict their favoured arm produces a treatment effect of 0.75 where the truth is zero. Nothing about the randomisation was broken; the bias entered through who was enrolled, which is the one thing an allocation rule cannot control. The dashed line is the closed form 2δ(2g − 1). Balancing on what was recorded first

The rule that can be guessed

A balancing rule improves as it becomes more deterministic, and a deterministic rule can be worked out in advance from information the person enrolling the patient already has. At full determinism 87.6% of assignments are guessable, and an investigator who acts on the guess produces a treatment effect of three quarters of a standard deviation where the truth is zero.

A window with two wrong ends. The regret of a rule whitened by a Bartlett-tapered Ω̂, as the window widens, against three rules that need no window at all. At L = 0 the estimate is the identity and the rule is exactly least squares — 0.08556, the same number to five places. It falls to 0.01754 at L = 20 and rises again by L = 30, because a quarter of the sample's lags are then being estimated from it. The automatic bandwidth a practitioner would reach for, 4(n/100) to the power 2/9, which at this sample size is 4, gives 0.02963 — 69% above the best available window. The rule told the dependence is an AR(1) sits at 0.01440 throughout, which is the price of not knowing the form. Estimating the dependence, not naming it

The window that has to be chosen, and the term that was dropped

An estimated covariance has a bandwidth in it, and both ends of the dial are wrong for different reasons. The rule a practitioner would reach for is two thirds worse than the best window there is.

Four designs, and what each of them guarantees. The worst Ds-efficiency each design achieves anywhere in a rectangle of parameter values 4 times wide in each coordinate. The design built at the guess guarantees 7.3%: it is perfect where it was built and nearly useless at one corner. The D-optimal design at the same guess guarantees 25.6% — it answers the wrong question everywhere and is therefore not concentrated on being right anywhere. The third is the one this field exists to measure: maximin over the first parameter's whole range, with the second held at its guess. It guarantees 7.0%, which is no better than the design that protects nothing. Protecting both is worth 45.9%, and it needs 3 settings to do it. What the procedure may not read

The worst case in two directions

A design that protects a range of one parameter is robust. Protect the range of one parameter while holding the other at a guess and the design is still robust, still has a guarantee, and guarantees no more than a design that protects nothing at all.

A near-perfect cancellation, at one correlation. A covariate skewed at 0.75 under a lower-tail copula, at each of 7 rank correlations. The earlier field measures this cell at a Spearman of 0.4 and reads 0.002% where adding the two halves' own leaks gives 16.579% — a cancellation so near exact that it is that field's headline. Across the sweep the same cell reads 0.0084%, 0.0182%, 0.0097%, 0.0015%, 0.0864%, 0.4693%, 1.5327%. Its smallest value is at 0.4, in the interior, and by 0.7 it is 1007.40 times larger. The near-zero is where two curves cross, and they cross beside the one correlation that was measured. The same table at seven correlations

The zero that was a crossing

A cell that leaks 0.002% where adding its two halves gives 16.6% is a field's headline. On a finer grid it passes through zero at a rank correlation of 0.38 — two hundredths from where it was measured.

Two factors, opposite directions. The two factors the cost of a per-candidate tuning parameter is a product of, as the candidates are pulled apart, over 800 draws at each of 5 separations. How often the candidates disagree about the tuning parameter rises from 31.8% to 88.8%; the share of those disagreements that change which candidate the table selects falls from 51.6% to 1.7%. So the setting where the candidates quarrel most about the tuning parameter is the setting where the quarrel matters least, and a sweep that reads the rate and stops has read the factor pointing the wrong way. What decides whether a tuning list decides

Two factors pointing opposite ways

As the candidates on a table are pulled apart, they quarrel about the tuning parameter three times as often and the quarrel decides the winner thirty times less often. A sweep that reads the first factor has read the one pointing the wrong way.

Two independent random walks, 100 steps. Nothing connects these two series: each is generated from its own independent draws. Regressing one on the other gives a slope with t = -10.9, R² = 0.55 and p = 0.0e+0 — a result that would be reported as a finding by any standard output. When the observations repeat each other

Two walks and a finding

Regress one random walk on another, independently generated, and the slope is significant 76.7% of the time with a median R² of 0.17. Nothing connects the two series, nothing in the output says so, and more data makes it worse.

What a positive test means, sensitivity 90%, specificity 95%. At a prevalence of one in a thousand, 98 of every hundred positives are false. At one in 10, 33 are. The test has not changed. Reversals that are not errors

What a positive test is worth

A test that is 90% sensitive and 95% specific sounds accurate. For a condition affecting one person in a thousand, 98% of its positive results are wrong, and a worse test on a commoner condition beats a better test on a rare one.

What a search manufactures, law by law. The average likelihood ratio a search over 120 rows reports, on each of the four laws, over 400 draws. Three of them have no break at all and report 5.080, 4.839 and 4.748; the fourth has one and reports 8.757, so the real break is worth only 3.677 beyond what the search would have found anyway. The dashed line is 2, which is what an information criterion charges for one extra parameter. A search costs about two and a half of them, and the number is a measurement rather than a count. Paying for a search

What a search costs in parameters

An information criterion's penalty is an estimate of the optimism a fit carries. For a break point the optimism can be measured and cannot be counted, and it comes to about two and a half parameters.

Both halves grow; the difference does not. The control pair's two components and their difference, against how much each of its two searches can find, over 1200 draws at each dictionary size. Two disjoint sets of independent columns are additive at every size — the excess stays inside a standard error or two of zero throughout — and it is not because there is nothing there. The overlap grows from 0.000112 at two columns to 0.000870 at ten, a factor of 7.76, and the interaction grows with it, staying within a factor of two of the overlap at every size. Two searches competing for one residual sum share ground and find configurations neither has alone, in almost equal measure, and their difference is what the earlier field's scale calls zero. Overlap and complementarity, separated

What a zero is made of

Two disjoint dictionaries of independent columns read an excess of 0.000116 and are made of an overlap of 0.000583 and an interaction of 0.000467. The control the whole scale is anchored on reads zero because two effects cancel.

What studentising costs. How much wider the studentised interval is than the percentile one, cell by cell, over 300 draws, with what each cell gains in coverage beside it. Averaged over the eight cells the interval is 2.09 times as wide and covers 0.46 points better. At the two rules that choose short blocks the two intervals are within a fifth of each other; at the oracle's length, where a resample holds two or three whole blocks, the studentised interval is 4.37 and 5.04 times as wide. A repair that doubles the width and buys half a point is not one a reader could not have had by widening the interval it replaced. The interval, studentised

What studentising costs

Averaged over eight cells the studentised interval is 2.09 times as wide as the percentile one and covers 0.46 points better. At the block lengths the rules choose, the scale it divides by rests on two or three numbers.

What the worst case is worth, one function at a time. The smallest share each dictionary removes, over seven outcome shapes, at a correlation of 0.5. A rule balancing the mean of each covariate has a worst case of exactly zero — against the square, and against both products. Adding a median split to it, which is the second thing every trial balances, leaves the worst case at exactly zero, because a median split is odd and so is a mean. Adding the square instead moves it to 6.8%, and the extra functions after that move it to 7.4%. The worst case is decided by which parities the dictionary contains rather than by how many functions are in it. What a dictionary buys and what it costs

What the extra function buys

A rule balancing the mean of each covariate has a worst case of exactly zero. Adding the median split — the other thing every trial balances — leaves it at exactly zero, and one square moves it.

A test between nested models, under a null that is true. 1000 comparisons: an AR(1) truth, forecast by a fitted AR(1) and by a fitted AR4 whose extra coefficients are zero. In population the two forecasts are identical, so every rejection is false. The larger model's mean squared error is 1.1663 against 1.0583 — worse, by exactly the noise in estimating coefficients that are not there — and the ordinary test therefore declares the smaller model significantly better 67.2% of the time. Read one-sided in the direction anybody asks about, it finds the larger model better 0.0% of the time. Adding the squared difference between the two forecasts back into the loss differential puts the level at 4.9%. Comparing two forecasters

When one model contains the other

The comparison a forecaster most often wants is between a model and the same model with one more term. That is exactly the comparison the standard test cannot make — and it fails by declaring the smaller model significantly better, more confidently the more data it is given.

One curve is a binomial coefficient and the other is a line. The number of subsets a maximin over this dictionary would have to score, against the number the exchange algorithm actually scores. At three functions the walk is 2,024 subsets and is the honest answer; at eight it is 735,471 and the exchange algorithm has looked at 421. The warrant for the second curve is the four sizes where both exist and agree, which is a weak warrant — it says the algorithm has not yet been wrong, not that it cannot be — and it is the only one available past the point the first curve leaves the page. When the set is too large to walk

Where the enumeration stops

A maximin over an eight-function dictionary is a walk over seventy subsets. Over twenty-four it is 735,471 at eight functions, and the exchange algorithm that replaces the walk scores 421. What licenses the second curve is four sizes where both exist and agree, which is a weaker warrant than it looks.

Generality in the wrong direction buys nothing. Regret on a sample whose persistence changes from 0.95 to 0.65 at row 60, over 200 draws. The three stationary rules — told one number, told a window, told an order — are within 0.4 standard errors of each other, and all three stop in the same place: they are general in the lag direction, and the departure is in the other one. Letting the model change once, at a point estimated from the same residuals, is worth 0.05021 more at 4.5 paired standard errors — about as much again as the whole of the first repair. Being told where the break is adds 0.01926, and being told the entire covariance adds 0.02465. The shape a dependence has

Where the generality runs out

A covariance that changes half way through a sample is not one a window can estimate. One number, a window and an order are worth the same as each other on it — and letting the model change once, at a point nobody can locate, is worth as much again as all three.

The crossing is in the dependence, not in the split. Regret of each rule as the design and the errors are made persistent at the same coefficient, scored on fresh rows because the closed form assumes exactly what is being taken away. An optimism theorem counts rows; when the rows repeat each other there are fewer of them than there are rows, the penalty is too small for the fit it is correcting, and the criterion starts buying coefficients it should not — its average winner grows from 3.31 coefficients to 3.90. The hold-out never used the theorem and overtakes at ρ ≈ 0.81. Schwarz's criterion, worst of the three on independent rows, is best on repeating ones — its heavier penalty is right for the wrong reason. Scoring a search without spending data

Where the two searches cross

The obvious dial between a criterion and a hold-out is how much of the sample to hold out, and moving it never changes the answer. The dial that does is one nobody chooses — how much each row repeats the one before it — and the two rules change places at about 0.81.

Three sets of weights, five designs, and no estimator that is exact everywhere. Coverage of the same interval under three weightings. h_b is the inverse variance when the arms share a variance or the allocation is constant; equal weights are right when every block has the same two counts; the estimated precision weights are right in the limit and exact nowhere, because the decomposition needs the weights to be the constants they are only estimating. In the corner — two variances, changing sizes, changing allocation — the two exact estimators are the ones that miss, at 98.45% and 95.65%, and the one with no theorem behind it is at 95.05%. That is the whole statement: there is an exact estimator under either condition, and none under both. A promise about two arms

Which weights are the inverse variances

There is an exact estimator when the two arms share a variance and another when every block has the same two counts, and between them they cover every trial anybody designs on purpose. In the corner where neither holds, both cover 98.45% instead of 95%, and the only estimator at its level is the one with no theorem behind it.

Unrepresentative in every respect but the one that matters. Three properties of the complete cases as the chance of being observed leans harder on the regressor, in closed form, at 35.0% of outcomes missing throughout. The mean of the regressor among the rows kept climbs from 0.0000 to 0.5528 against a population mean of zero, and the mean of the outcome from 0.0000 to 0.3980 above its own. The bias in the fitted slope is exactly zero at every one of the ten settings, because selection acting on the regressor alone leaves the conditional law of the outcome given the regressor untouched and least squares conditions on exactly that. The sample is wrong about almost everything and right about the one quantity being estimated. The value that is not there

Dropping the incomplete rows

Push the missingness until the rows that survive have a covariate mean of 0.543905 against a population zero and a variance of 0.5041 against one, and the fitted slope is still exactly right. Where the rule reads the outcome instead, the same sweep takes coverage to 2.42% at eight hundred rows.

What each error is a claim about, and what the claim comes out as. Each variance estimate's average over 20000 draws, divided by the variance the slope actually has across those same draws, at 80 rows with the error variance leaning towards the edges of the design (γ = 0.8). One is a standard error that is right. The model-based estimate reads 0.6081 of the spread, so its standard error is 77.98% of the one it should report; the four robust corrections read 0.9576, 0.9821, 0.9961, 1.0362. Two further routes agree with the count and share none of its arithmetic: n times the counted variance is 4.8905 against a population sandwich of 4.9200, and the counted ratio of the two standard errors is 1.2799 against a closed form of 1.2806. A standard error for a model that is wrong

The bread and the filling

The robust standard error is not a safety margin. At one setting of the error variance it is 1.2806 times the model-based one and at another it is 0.8246 times it, and the sign of a single dial decides which.

One odds ratio in every stratum, and a different one marginally. Five strata with baseline risks from 5% to 85%, a treatment allocated by a coin in each, and a conditional odds ratio of exactly 2.5 throughout. The marginal odds ratio is 1.789. Nothing is confounded; the odds ratio is simply not a weighted average of odds ratios. When the stratified answer and the pooled one disagree

The change that is not confounding

Five strata, a treatment allocated by a coin in every one, and an odds ratio of exactly 2.5 in all five. The odds ratio computed on the pooled table is 1.789. Nothing is confounded — an odds ratio is not a weighted average of odds ratios, and the risk difference, on the same table, is exactly its own stratum value.

What a search costs is not a property of that search. The likelihood ratio a searched break in the regression reports, two ways, on every law. On its own — the whole rule being a split of the sample, at no whitening — it averages 34.70 under AR(1) at 0.8, against the 11.07 a chi-square on the five coefficients a split adds would use as a threshold. Inside a rule that also chooses a window from a list of eight, the same search adds only 15.40 — less than half. Most of what a break search finds under correlated errors is the correlation, and a whitening chosen from the same sample has taken it already. A charge measured for one search, carried into a rule that makes two, is not conservative in some harmless direction: it is measuring a different quantity. Two searches over one sample

A charge that depends on the rule

The break search's charge is 34.7 on its own and 15.4 once a window has been chosen from the same sample. Most of what a break search finds under correlated errors is the correlation, and a whitening has taken it already.

Three copulas that break nothing, and a factor of two between them. The three radially symmetric copulas, at a matched Spearman correlation of 0.40, against the covariate's marginal. All three leave exactly nothing with a symmetric covariate — that is the guarantee, and it holds to twenty decimal places. What they do to a skewed covariate is not the same at all: at a skewness of 2.26 a Frank copula leaves 12.118% where a Gaussian leaves 21.539% and a t on four degrees of freedom leaves 23.640%. A factor of 2.0 between two copulas that are both symmetric, both matched on rank correlation, and both harmless on their own. So the copula matters to the marginal's leak without breaking any symmetry of its own, which is a milder version of the same finding and applies to every trial rather than to the asymmetric ones. Both halves of the dependence at once

A copula that halves a marginal

Three copulas break nothing on their own and put a factor of two between the same skewed covariate's leaks — 12.118% under a Frank against 23.640% under a t, at the same rank correlation.

A rate that does not know how large the trial is. The share of equal splits admitted by a tolerance of 1 coin-spreads on 3 functions, at six trial sizes. The first two are exact — 12,870 and 184,756 splits, walked, averaged over eight draws of the units — and the rest are sampled. From a hundred units on, the rate sits on (2Φ(1) − 1)^3 = 0.3182, which contains no n at all. The two small trials are 29.2% and 27.7% short of it, so the sixteen-unit measurement understates the rate rather than bracketing it. Meanwhile the admissible count — the rate times C(n, n/2) — goes from 2^11.5 to 2^393.7: the exhaustion a small trial runs into is a fact about small trials. When the set is too large to walk

A count that has to be estimated

At sixteen units the admissible assignments can be counted by walking all 12,870 of them. At four hundred there are about 2^393.70, and the share admitted is 0.31885 against a closed form of 0.31818 that has no trial size in it at all. The exhaustion a small trial runs into is a fact about small trials.

Thinness stays put and reachability does not. The same balancing rule and the same tolerance at five trial sizes. The admitted share barely moves — one admissible assignment in 27, 30, 30, 20, 18 — because the acceptance rate of a rerandomisation is a fact about the basis rather than about the number of units. What does move is the number of single swaps available: 36, 49, 64, 81, 100, growing like a quarter of the square of the trial size. Filled marks are sizes whose admissible set falls into more than one piece. The set is in 4 pieces at 12 units, 2 at 14, and one piece from 16 upwards. The fourteen-unit result is a statement about fourteen units. What a chain cannot report

A defect that is about size

The admitted share of a rerandomisation barely moves with the number of units. The number of admissible neighbours grows like the square of it, and that is what decides whether the walk can go everywhere.

What a longer list actually changes. How often the five candidates choose different tuning parameters, at a true null where every one of them contains the truth, so a disagreement is manufactured rather than discovered. The order's list is an interval of integers, and thinning it moves the rate smoothly from 0% at two values to 39% at thirteen. The window's is not an interval — it runs 0, 1, 2, 4, 8, 12, 20, 30 — so a thinned window list jumps depending on whether it happens to keep the width the criterion wants, between 0% and 42% with no order to it. So "the same length" was never quite the same thing for the two rules, and it is a smaller effect than the field it was invoked to explain. How long the list is

A list is not a rule

How often five candidates disagree about a tuning parameter runs from nothing at two values on the list to two draws in five at thirteen. What the disagreement costs does not move at all.

What each probe can see. How far apart the two components of the admissible set are on each probe, over the spread inside a component, on 100 designs whose set is enumerated and split. It is the population quantity a chain is trying to report. The separating direction itself reads 10.5646; the projected fourth power 5.0800, the design's own leverage 3.8362, the modelled active set 1.9529, the counted active set 2.0170 and a random direction in the same subspace 0.9422. The two active-set probes beat the random direction and lose to both of the earlier field's, which is the field's answer to the question that opened it. A probe from what the rule blocks

A quantity that loses to a heuristic

Leverage is a heuristic about which units a balancing rule has most to say about. The constraint's active set is the thing the rule actually does. As a probe, the heuristic wins by 4.4 paired standard errors.

The split depends on the order. How much two searches share, measured both ways round, over 300 draws. Pinning the first search at its own answer and searching the second gives one overlap; pinning the second and searching the first gives another. A break paired with an independent column reads -0.004395 one way and 0.002163 the other, at 8.70 paired standard errors and on opposite sides of zero. The excess the two components subtract to is the same in both orders by construction, so what changes is only how it is attributed. There is no order-free way to say which of two searches found ground both can reach, and the two orders bracket it. Overlap and complementarity, separated

A split that depends on the order

Run the second search first and pin that instead, and the same draw gives a different overlap and a different interaction — with the same difference. And one pair has no second order at all.

Flat along a row, apart between them. The probability that a per-candidate tuning list changes the winner, at three list lengths on three candidate tables, over 800 draws in each of the nine cells. Along a row — the reading the earlier field takes — it moves by a factor of at most 1.21, so that field's invariant survives on every table. Down a column it moves by up to 1.98. The list length is the dial that does not move this number and the table is one that does, and the earlier field varied only the first. What decides whether a tuning list decides

A table and a list

A nested ladder of candidates differing by one coefficient was predicted to turn over more often at every list length. It turns over less at every one, and its list changes the winner half as often.

Four cells change their answer. The four cells of the twenty whose excess changes sign as the dependence strengthens, over 7 recalibrations. Above the line the two failures compound — the cell leaks more than adding the copula's own leak and the marginal's — and below it they cancel. All four start above and end below, and all four are at the two most skewed covariates: skew 0.90 under heavy-tailed, skew 0.95 under heavy-tailed, skew 0.90 under upper tail, skew 0.95 under upper tail. Whether two failures of a dependence compound or cancel is therefore not a property of the pair. It is a property of the pair at a strength of dependence, and a fifth of the table changes its answer inside the range measured here. The same table at seven correlations

An answer that changes

Eleven of twenty cells cancel and nine compound, at one rank correlation. Sweep the correlation and four of the twenty change sides — all four from compounding to cancelling, all four at the most skewed covariates.

How often the split is taken, and by which rule. Over 400 draws on each of five laws. The first two rows have no break in them at all, the last two have one at row 60, and the middle one is a moving average. A criterion that counts a fitted two-regime model's parameters and nothing else takes the split on 99% of draws where there is no break. Counting the break point as one more parameter brings that to 67%. Charging what the search actually manufactures — 5.16 units, measured on a law with no break — brings it to 16%, and still takes the split on 61% of draws where there is one. Paying for a search

Choosing whether to break

Charging what the search manufactures takes a rule from splitting a stationary sample on 99% of draws to 16%. It also costs regret, because the two mistakes a rule can make are not the same size.

One likelihood, three answers. The concentrated Gaussian log-likelihood of one sample of 120 rows under AR(1) at 0.8, as a function of the correlation the errors are whitened at. Three rules put three different numbers on this curve. The two-step rule reads the least-squares residuals and lands at 0.7616, giving up 0.304 of log-likelihood. Iterating moves it to 0.8080 and gives up 0.002. The maximum is at 0.8044. The curve is not flat between them: what a fixed point of the residual update finds is a solution of a different equation, and the difference is the Jacobian term ½log(1 − ρ²), which grows as the correlation does. Fitted together, or fitted after

Iterating is not maximising

Re-reading a correlation from the generalised residuals and refitting converges in seven steps. What it converges to solves the first-order condition of a sum of squares, and the likelihood has one term more than that.

What each instrument costs to read. The number of draws each instrument needs to separate a rectangular block from a trapezoidal one at two standard errors, at a block length of 20 and 120 rows — measured from each instrument's own spread on the same draws. The implied variance needs 7.0 and the 95% point needs 20.2, a factor of 2.90 at this block length. There is a closed form beside it and it does not depend on either the scale or the size of the gap: the standard error of a p-quantile is √(p(1−p))/f(q) over √B where a standard deviation's is σ/√(2B), which at the 95% point of a nearly normal reference distribution is 3.30 times as many draws for the same statement. And the quantile route needs every one of those draws resampled, where the variance route needs none. Where a taper's case begins

Measuring a variance rather than a quantile

A resample's implied long-run variance can be computed from the sample with no resampling in it at all. A critical value cannot, and the difference is a factor of three in the draws before any of the resampling is counted.

Two factors that interact — where each design looks. The true response at the four corners is -10, 6, 4, 0. One factor at a time visits three of them, sees that raising either factor alone helps, and recommends raising both — a corner it never ran, and one that is worse than either single change. It picks the best corner 0.0% of the time against the factorial design's 99.4%, on the same number of runs. Decided before the data

One factor at a time

Changing one thing per experiment estimates each effect from two conditions; changing everything at once estimates each from every run. The ratio is (k+1)/2 and it is exact — and when two factors interact, the one-at-a-time design recommends a setting it never tried.

R² against the number of useless predictors, n = 30. The response is pure noise and so is every predictor, so the true relationship is nothing at all. R² rises from 0.000 to 0.648 anyway, following k/(n − 1) — which is what a criterion that rewards higher R² is actually rewarding. Regression, and what the summary hides

R² is not a measure of fit

Adding a predictor with no relationship to anything cannot reduce R², and in expectation raises it by 1/(n − 1). Twenty useless predictors on thirty points give an R² of 0.69 from pure noise.

Two measurements of the same thing, correlated 0.60. Pick the worst 15% on the first measurement and their average rises by 0.78 on the second. Pick the best and theirs falls by 0.48. No treatment was given to anybody. Reversals that are not errors

Regression to the mean

Select the worst performers, measure them again, and they improve. Select the best and they decline. No intervention is required for either, the size of the apparent effect is predictable from the correlation alone, and it is the reason so many things appear to work.

Which repair goes with which defect. The share of true nulls rejected at a nominal 5% by a reference distribution generated from the fitted benchmark, over 120 draws with 59 resamples each. Where the errors are well behaved every resampling is fine and all four are conservative. Where the variance is a function of the design, the two that detach a residual from its own row reject 5.8% and 8.3% — and the block bootstrap, which is the resampling three earlier fields on this site reach for, repairs nothing at all, because the dependence it is built for is between origins and the rolling scheme reproduces that on its own. Where the errors are skewed the symmetric multiplier is the one that is wrong, and Mammen's two-point version is the only one of the four that is right in both columns. A search with no fixed point

Residuals that keep their own variance

A reference distribution for a search has to be generated from a fitted model, and the generator draws residuals. Four ways of drawing them keep four different things — and the one this site has reached for three times repairs nothing at all here.

The weights may not read the block they weight. A weighted least squares decomposition needs weights that are constants, or at least independent of the differences they multiply. One λ̂ pooled across the trial is estimated on hundreds of degrees of freedom and is effectively a constant; a λ̂ estimated inside each block is estimated on that block's own two or three, and is correlated with the difference it weights. Coverage falls from 94.68% to 82.76% — and the interval gets wider while doing it, 0.5163 against 0.3024, which is the signature of weights that are noise. The weights the corner needs

The condition that cannot be dropped

The weights may not read the block they weight. Estimate the variance ratio inside each block rather than across the trial and the coverage falls to 83% — on an interval that is at the same time seventy per cent wider.

A cut at a quantile, and a cut at a value. Two rules that read identically in a protocol. One splits each covariate at its median; the other splits it at 1 on the covariate's own scale — a dose, a temperature, a clinical threshold. At a correlation of 0.5 the first removes exactly nothing of the interaction between its own two splits, under every marginal here, because a median split is a function of the sign of the latent normal whatever the marginal is. The second removes what the bars show, and it does so on a normal covariate too: the threshold sits at 1.000 on the latent scale rather than at zero, so it is 59.4% odd and 40.6% even. The exact zero was never about the cut; it was about the cut being at the median. A guarantee that needed a symmetry

The cut that is not a quantile

A protocol that says split the covariate at a threshold and one that says split it at the median read the same and are different rules. One has an exact guarantee under every marginal and the other has none under any.

Two arms leave one degree of freedom per block unaccounted for. Each point is one run. The one-mean field's identity is (b − 1) + (N − b) = N − 1, and every schedule moves along that line rather than off it. Two arms give the rule N − 2b and the interval b − 1, which come to N − b − 1 — short of the N − 2 two arms leave by exactly one per block, since a block's arm counts absorb one degree of freedom each and only one of the two directions carries the difference. The hollow points add what the block sums are worth, b − 1 more, and land on the total. The missing degrees of freedom are not lost; they are in a place the interval has to be shown it may read. A promise about two arms

The degrees of freedom in the sums

One arm partitions N − 1 exactly. Two arms give the rule N − 2b and the interval b − 1, which is short by one per block — and the missing ones are in the block sums, which are correlated with the differences at −0.79 and are usable anyway.

Four rules of four change sign. The margin between the two block windows in points of coverage, under each of four rules, on each of three intervals built from the same resamples, over 300 draws. Positive is the tapered window covering better. On the percentile interval the taper wins at all four rules, by 5.33, 1.67, 5.00 and 4.00 points, which is the earlier field's own reading. On the studentised interval the rectangle wins at all four, by 4.33, 7.00, 4.33 and 2.67. And a normal interval, which uses no resampling at all, puts the two within a third of a point at every rule — so the disagreement is manufactured entirely by what is done with the resamples. The interval, studentised

The ordering reverses again

One field found two of four rules changing sign between two readings of one resampling. Turn the same resamples into a studentised interval instead of a percentile one and all four change sign.

The row count entered twice, and a penalty is one place. Regret against the best available model as the errors are made persistent. Counting rows more than quadruples; both penalty repairs — the trace, and the scalar effective sample size — are worse than it at every persistence measured; and whitening the sample and keeping the ordinary penalty falls, recovering 86.9% of what counting rows gives up at ρ = 0.85. Doing it at an estimated ρ recovers 80.6%, so having to estimate the dependence from the rows being selected on costs 7.3% of what knowing it is worth. Mallows' forms are drawn beside the logarithmic ones and behave the same, which is what rules the linearisation out. Counting what is independent

The repair that was exact and made it worse

A penalty computed from the trace is exactly the optimism it estimates, and selecting with it gives up a fifth more than not correcting anything. The row count entered the criterion twice, and a penalty is the second place.

The reversal is a property of the instrument. The margin between the two block windows under each of four rules, on three readings of the same resampled means, signed so that a positive bar is the tapered window winning. On the implied long-run variance the taper wins at the best available block length and at one estimated from the sample and loses at a length written into a protocol and at the rule of thumb — which is the reversal the earlier field's whole argument turns on, at 1.48 and 4.52 points. On the 95% point a test actually reads, the taper wins at all four, by 6.13 to 7.08 points. On the coverage the interval actually delivers, the taper wins at all four again, by 2.50 to 5.75 percentage points. Two of the four rules change sign between the first reading and the other two, and the two that change are exactly the two the earlier field's recommendation is about. The block length read on a quantile

The reversal that was the instrument's

On an implied variance the rectangle wins at a protocol length and at the rule of thumb. On the 95% point a test reads, and on the coverage an interval delivers, the taper wins at all four rules.

Two constructions on one triangle, and a third that is not. Three resamplings that all keep runs of neighbours, on the same residuals at a block length of 5, with the lags running past ℓ so that the tapers separate. A blocked multiplier never moves a residual; a fixed-length moving block moves every one; and they attenuate identically, worst gap 1.4 standard errors, both sitting on γ_resid(k)(1 − k/ℓ)⁺ and both exactly zero past ℓ — so the attenuation is the block boundary rather than the multiplier. The third is the stationary bootstrap, whose runs are geometric rather than fixed: its taper is γ_resid(k)(1 − 1/ℓ)^k, it agrees with the other two at the first lag and at no other, and at lag 6 it still carries 0.0081 where they carry -0.0005. Estimating the dependence, not naming it

The triangle that was not the multiplier's

A resampling that leaves each residual on its own row can keep only what the residuals have, times a triangle. A construction that moves every one of them has the same triangle — and the one in this collection's own table has a different taper entirely.

The window a whitening wants is not the memory of the errors. Regret under a five-period moving average as the tapered estimate is given more lags, over 120 draws at n = 120. The best window is L = 30; the automatic bandwidth is 4 and the error model's own likelihood chooses 9.7 on average. Both land in the same place and both are short, and the reason is the taper: a Bartlett weight at lag k is 1 − k/(L + 1), so a window of 8 keeps 0.556 of whatever the fourth lag carries and a window of 30 keeps 0.871. A window has to be several times the memory before it stops removing the memory. The dashed line is the rule told the errors are a first-order autoregression, which needs no window at all. The shape a dependence has

The window a whitening wants

Every law here is best whitened by a window several times longer than its own memory, including the one whose memory ends at the fourth lag. The three ways of choosing it from the sample all land in the same place, and it is the wrong one.

12,000 studies of a real effect of 0.3, n = 16. Power is 21%. The studies that reached significance report a mean effect of 0.621 — 2.07 times the truth. Every one of them is honest; the selection did the inflating. Tests, and the second number

The winner's curse

Filter honest studies down to the ones that reached significance and the effects they report are systematically too large. At low power the inflation is a factor of two, nobody has done anything wrong, and the selection did all of it.

The guarantee that survives a correlation, and the one that does not. What a balancing rule handed both main effects removes of the pure interaction between them, as the covariates become dependent. For median splits it is exactly zero at every correlation, because sign(x)² = 1: the interaction sign(X)sign(Y) is orthogonal to sign(X) and to sign(Y) whatever ρ is. For the product of the raw covariates it is 4ρ²/(1+ρ²)² — 64.00% by ρ = 0.5, rising to all of it at perfect correlation. A cut away from the median sits between them and is not small: 23.01% at a cut of one. The zero is not a fact about interactions. It is a fact about a dictionary whose functions square to a constant, which a polynomial one does not. A cut point, at a correlation

The zero that survives a cut

A rule holding both main effects removes half of a pure interaction between correlated powers and exactly none between correlated median splits. The guarantee that a correlation destroyed was never about interactions.

What a rule reads, against what the outcome uses. The variance of the treatment estimate relative to a coin's, for four things a rule might balance against three shapes the outcome might have, over 350 trials of 200 units. The diagonal is the easy part — a rule that reads the function the outcome uses removes about half the variance. What the table is for is the off-diagonal: reading the covariate alone is worth nothing against a quadratic (0.755), and reading all three is worth nearly as much against every shape as the matching rule is against its own (0.532, 0.493, 0.514). The shape the covariate enters by

Three functions of one number

A rule that balances the covariate is exposed to every shape the outcome might have. A rule that balances three functions of it costs two points of variance against the shape the first was built for and takes the worst case from a coin's to about half of it.

The ladder is the same ladder under every law. Each pair's overlap under each of the four laws, over 400 draws apiece. The scale is fixed at both ends by construction: two searches over disjoint sets of independent columns read -0.001, 0.013, -0.002, 0.004, and a break search paired with a step column it already contains reads exactly one under every law. Between them the pair that reads one residual series twice runs 0.763, 0.805, 0.577, 0.752 — lowest under long memory, where a whitening has most to do and the break search has least left to find that the whitening has not taken. The rung that moves most is the pair of step dictionaries, from -0.148 under a moving average to 0.419 under a break; and the pair that is negative is negative under all four. Two searches over different features

Three quarters of the way to one search

The pair that started this reads 0.762 on a scale whose one is containment. And the pair that shares nothing but its response reads −0.306, so the sign the earlier field found does not transport at all.

The overshoot is the last block size and nothing else. A run stops at a multiple of its own block sizes and cannot land between them, so it ends past its own target by about half a block. Fixed sizes overshoot by 1.5, 2.2, 3.1, 4.7, 8.5 observations as the size goes 2, 3, 5, 8, 16. Every schedule here ends in blocks of two and every one of them lands where blocks of two land — 1.62, 1.32, 1.37 against 1.48 — while having spent most of the run inside blocks four and eight times larger. That is the one thing on this page a schedule genuinely takes from both ends. The block size as a schedule

What a schedule actually buys

Big blocks early and small blocks late is the right instinct and it does not take both ends of the trade, because there are not two ends to take. What it does take is the overshoot — about four per cent of the observations — and a steadier stopping point.

The argument is a third of the size of the thing it is inside. Three quantities on one scale, in points of the error in a block resample's implied long-run variance, at 120 rows. The gap between the two windows at the best available block length — the whole subject of the comparison this field inherited — is 2.12 points. What the best rule a practitioner could actually run gives up against that same best length is 7.26, a factor of 3.42. What the rule of thumb gives up is 26.01. So the ordering between windows is worth establishing and is not worth arguing about, and the sentence that follows from it is not use the taper but estimate the block length, because that is where the points are. A block length chosen from the data

What choosing the length costs

The gap between two block windows at the best available length is 2.12 points. What the best rule a practitioner could run gives up against that same length is 7.26. The argument is a third of the size of the thing it is inside.

How often "there is no spread between the groups" is reported about data that has one. Every dataset here was generated with a real population spread of 1. The moment estimator is the difference between the observed spread and what noise alone would produce, clamped at zero, and the difference comes out negative often: at eight groups it reports exactly zero on 32.6% of datasets, which is an instruction to pool completely and give all eight groups the same estimate. The rate falls to 4.2% at 48 groups. The spread, and its own uncertainty

When the spread estimates to zero

The usual estimate of a population spread is a difference of two positive quantities, clamped at zero. On a third of eight-group datasets with a real spread in them the difference comes out negative, the estimate is exactly zero, and every group is pooled completely on data that said no such thing.

A class that is a subspace has no guarantee below its own dimension. Each cell is the worst case over every unit-variance function in a class of dimension m, for a rule reading k functions: the smallest squared principal-angle cosine between the two subspaces. Wherever k is less than m the number is zero to machine precision, and that is not a weak guarantee but the absence of one — some direction of the class is orthogonal to the entire basis, and against an outcome in that direction the rule does exactly what a coin does. An experimenter who declines to name the shapes and asks instead to be protected against everything smooth is asking for the cells above the diagonal. Choosing what the rule reads

Where the guarantee is exactly zero

An experimenter who declines to name the shapes, and asks instead to be protected against anything in a class, is asking for a number that is not small but zero. Bounding the class is unavoidable, and the two ways of doing it choose different bases.

Which tail the threshold is in. What a rule balancing a threshold at 1 on each covariate's own scale removes of the interaction between the two thresholds, on five copulas matched at a Spearman rank correlation of 0.4 with a normal covariate throughout. This rule never had a zero to lose — the earlier field establishes that under every marginal — so what is left is a size, and the size depends on where the dependence lives. A Clayton copula, whose density piles up in the lower tail, leaves 5.33%; the same copula turned over, so that it piles up in the upper tail where the threshold is, leaves 33.36%. Same rank correlation, same Kendall tau, same marginal, same threshold: 6.26 times the leak, decided by which end of the distribution the dependence and the cut are both in. The other half of the dependence

Which tail the cut sits in

The same copula and its reflection have the same rank correlation, the same Kendall tau and the same marginals. A balancing rule holding a threshold at a dose leaves 5.33% under one and 33.36% under the other.

Two structures in three are made worse. What adjusting for every covariate measured does to the bias in the treatment's estimated effect, against adjusting for none, over 4000 randomly drawn structures of 6 covariates each. Each covariate is independently a common cause with probability 0.25, a cause of the treatment only, a cause of the outcome only, a cause of neither, a step on the causal path, or a common effect. The rule leaves a larger bias on 65.5% of structures, a smaller one on 33.8%, and the same on 0.7%. The share is a property of that population of structures rather than of adjustment, which is why the weights are stated; what does not depend on them is that the rule has no direction — it is not a conservative default that occasionally overcorrects, it is a rule whose error is whatever the structure happens to be. What conditioning on a variable does

Adjusting for everything

"Control for every covariate that was measured" leaves a larger bias than controlling for nothing on 65.5% of four thousand randomly drawn structures and a smaller one on 33.8%. Its squared error is 4.110 times that of using no covariate at all, and half of it sits in its worst tenth of structures.

What stratifying on a mediator estimates, and what it does not. The direct effect stays at -0.109 throughout, because it is defined with the mediator held fixed. The total effect runs from -0.218 to 0.027 and crosses zero at 1.5. On 15% of the sweep the two have opposite signs, and both are correct answers. When the stratified answer and the pooled one disagree

Conditioning on what the treatment caused

When the grouping variable lies on the path from treatment to outcome, the stratified answer is the direct effect and the aggregate is the total effect. Both are correct. Over 15% of a sweep of the indirect path they have opposite signs, and no arithmetic on the table says which question was being asked.

What the family-wise correction does to the effect it lets through. At two standard errors the estimate that clears an uncorrected 5% threshold averages 1.35 times the truth, and the one that clears the family-wise threshold averages 1.69 times it. The correction fixes the error rate by demanding a larger estimate, and a larger estimate is a more selected one. The analyses that were available and not run

The correction that makes the estimate worse

Correcting for twenty analyses repairs the p-value by demanding a larger statistic, and a larger statistic is a more selected one. At two standard errors the surviving estimate averages 1.35 times the truth before the correction and 1.69 times it after — so the honest error rate is bought with a more inflated effect.

How far apart four summaries put the quartet. Each bar is the largest value minus the smallest across the four datasets. Pearson spans 0.0018, which is the construction working. Distance correlation — the measure that is zero if and only if the variables are independent — spans 0.101, and Spearman spans 0.491. What a summary of a scatter is a property of

The summary that was meant to work

Distance correlation is zero if and only if two variables are independent, which is exactly the guarantee a correlation coefficient lacks. Run on the four datasets that share a correlation, it spreads them by 0.10 — and Spearman, which guarantees nothing, spreads them by 0.49.

What a design does to the residuals of a correct model. Every residual has standard deviation sigma times the square root of one minus its leverage. On this design the leverages run from 0.045 to 0.663, so the residual spreads differ by a factor of 1.68 — and the model is exactly right. The high-leverage point's residual averages 0.46 of the fitted spread where a typical point's averages 0.79. What a diagnostic plot is showing

Residuals are not the errors

A residual's standard deviation is σ√(1 − hᵢᵢ), so a design whose leverages run from 0.045 to 0.663 produces residuals whose spreads differ by a factor of 1.68 with the model exactly right. On the samples where the high-leverage point really did have the largest error, a raw residual plot shows it as the largest on 0.0% of them.

Clopper–Pearson, mid-p and the randomised interval: coverage across the proportion, 20 trials. Clopper–Pearson never falls below 95% and runs up to 99.80%. The mid-p interval, which is the randomised interval with its coin fixed at one half, runs from 92.94% to 99.80%. The randomised interval covers 95% at every proportion, to within the 0.043% of the numerical integration over the coin. A proportion's interval near the boundary, and the coin

The coin that makes it exact

Every interval for a proportion either covers less than 95% somewhere or more than 95% on average, because a count is discrete. One construction covers exactly 95% at every proportion: it adds a uniform random draw to the count. At thirty trials it is 0.9% wider than Wilson's interval and narrower than both exact ones — and two analysts with the same data report different intervals, and one study in forty that sees nothing reports an empty one.

Two studies of the same effect, z statistics with mean 1.96: where one is significant and the other is not. Five hundred pairs. With the true effect identical in both, exactly one of the two is significant in 50.0% of pairs; among those, the difference between the two is significant in 9.7%. The dashed lines mark 1.96 on each axis; the diagonal lines mark a significant difference. An interval read beside something else

Significant in one, not in the other

Two studies of exactly the same effect, each with 50% power, disagree about significance half the time — and when they do, the test of the difference between them is significant in 9.75% of cases. A p of 0.01 beside a p of 0.20 is a difference with p = 0.36. Among four subgroups sharing one effect, at least one significant and one not happens 87.5% of the time, and the test that would tell a real difference apart needs four times the sample the effect itself needed.

What a schedule is allowed to read, and what happens when it reads more. The construction allows the block sizes to be anything at all as long as they are functions of the within-block contrasts, which are independent of every block mean. A schedule that shrinks the block whenever the between-block spread is running above what the contrasts say is a direct attempt to hold down the quantity the interval will be built from, and it succeeds: the estimate lands at 0.8373σ² against the honest 0.9831, and the coverage goes with it. Reading the running mean instead pushes the other way and over-covers — which is not a repair, it is the same violation with the sign reversed, and the level is no longer a property of the procedure at all. The block size as a schedule

A schedule that reads the mean

The block sizes may be anything at all provided they are functions of the contrasts. Two natural schedules break that, in opposite directions — and the most natural mistake of the three is not a schedule at all but a stopping rule, at 86.87% coverage and fewer observations.

What a second break adds. Over 200 draws, the likelihood ratio a search over one break point reports, and how much more a search over an ordered pair adds on top of it. Under AR(1) at 0.8, which has no break at all, the first search manufactures 5.697 and the second adds 4.278. Under a law with exactly one break — where a second one is as absent as the first was in the row above — the first search reports 9.442 and the second still adds 5.800. Searching for something that is not there costs the same whether or not something else was there to find. Paying for a search

A second break on a flat profile

Searching a hundred and twenty rows for one change point manufactures five units of likelihood. Searching for a second manufactures four more, on a series that has at most one — and on a profile whose whole range is under seven.

One window for the table, or one each. The regret of the same fifteen-candidate table under AR(1) at 0.8 over 150 draws, with the window attached three ways. Chosen once from the fullest candidate's residuals it gives up 0.02518. Chosen from each candidate's own residuals, with the covariance estimate still shared, it gives up 0.02799 — a paired cost of 0.00281 at 2.0 standard errors for the tuning parameter alone. Estimating the covariance per candidate as well costs 0.01087, so the objection already on record is about 3.9 times the size of the one that was not. Fitted together, or fitted after

A window for every candidate

The window and the order a whitening needs are chosen once, from the fullest candidate, on an argument that was made about an estimated covariance. A tuning parameter is not a covariance, and the two cost different amounts.

A guarantee that stops being a number. The worst case of each dictionary over six outcome shapes, at a correlation of 0.5, against the skewness of the covariate. Under a symmetric marginal every rule made of odd functions has a worst case of exactly zero, and the rule holding a mean and a median split of each covariate — the two things every trial balances — is one of them. Under skew that zero becomes 0.74%, 1.83%, 2.24%, 2.49%: small numbers, each of which depends on a marginal nobody stated. The guarantee has not improved by becoming positive. It has stopped being a guarantee, because it can no longer be written down without the covariate's distribution in it. A guarantee that needed a symmetry

Balancing a skewed covariate

The worst case of the rule every trial runs goes from exactly zero to somewhere between a quarter of a per cent and two and a half. Which is small, and is a number that cannot be stated without the covariate's distribution in it.

One of them is mostly leverage. How much of the design's own leverage direction each active-set probe carries, once both are standardised and projected off the rule's span — which is what a probe is, so it is the comparison that matters. Over 189 designs the modelled active set agrees with leverage at |r| = 0.8359 ± 0.0114 and the counted one at 0.4239 ± 0.0216. So the modelled probe is largely leverage under another name and the counted one is genuinely a different direction — and the counted one is the worse probe, at 0.3854 of alignment against 0.4272. What the active set contains beyond leverage points away from where the set splits. A probe from what the rule blocks

Counting it exactly does not help

If a modelled active set lost because the model was crude, the exact one would win. It is computed at a cost no trial can pay, and it is worse — so the approximation was never what was costing the probe.

Where a walk is cheaper than a hunt. Both costs in the same unit. A rejection sampler evaluates 1/p assignments per independent draw and does not care how large the trial is; a walk evaluates one per step and yields an effective draw every τ steps, and τ is a property of the constraint and the statistic together. They cross at a tolerance of 0.194 standard deviations, where about one assignment in 396 is admissible — far tighter than any trial is designed at. And the walk does not remove the acceptance cost; it pays it once, hunting for somewhere to start. When the two are not independent

Draws that repeat each other

A hunt costs 1/p evaluations per independent draw. A walk costs one per step and yields an effective draw every τ steps. Both are counted in the same unit, and the walk is dearer at every tolerance a trial is designed at.

What each construction carries, against what there was. The autocorrelation of a resampled error series at five lags, averaged over 60 samples of 40 resamples each. Three facts are in the picture. The residuals lie below the errors at every lag, which is the ceiling a multiplier cannot exceed. The blocked multiplier and the fixed-length block lie on top of each other below it — they attenuate identically, because the attenuation is the join — while the stationary bootstrap, whose runs are geometric rather than fixed, sits above them both. And the sieve is the exception in kind rather than in degree: at lag six it carries 0.0638 where the residuals have 0.0300 and the multiplier has -0.0011, because a fitted model extrapolates past the lags it was told about and a truncated sample sequence cannot. Estimating the dependence, not naming it

Errors generated from a fitted model

The one construction that is not bounded by the residuals, because a model extrapolates past the lags it was told about and a truncated sample sequence cannot. It is nearly exact where the only defect is dependence, and it pays for it where there are two.

What a positive test means, sensitivity 90%, specificity 95%. At a prevalence of one in a thousand, 98 of every hundred positives are false. At one in 10, 33 are. The test has not changed. The prior, doing visible work

The base rate was always Bayes

The screening arithmetic everybody finds counter-intuitive is a posterior update with a prior of one in a thousand. Naming it that way turns a famous puzzle into an instance of a rule, and makes the sequential version obvious.

What the diagnostic says at two hundred units. The same test run 8 times on independent streams, at five tolerances of a two-hundred-unit trial, 40,000 steps each. At the loosest tolerance every run says the same thing — the walk reaches the whole set — and it keeps saying it as the set is thinned. Past a point the runs stop agreeing with each other: at the tightest tolerance here 6 of 8 report that the chains have not mixed and 2 report a split, which is a diagnostic disagreeing with itself rather than a property of the set. That disagreement is the honest answer at this size, and it is one nothing in this collection could give before: an enumeration stops at about twenty-four units. What a chain cannot report

The diagnostic at two hundred

Pointed at a trial size no enumeration reaches, the test gives three answers rather than one — and past a certain thinness it stops agreeing with itself, which is the honest reading and the one nothing could give before.

Where the taper's case begins, and it is not where the algebra says. The block length at which a trapezoidal block's implied variance stops being more biased than a rectangular one's, against the length of the sample. Computed exactly — from the law's own autocovariances, with no sampling in it — the answer is 19.2 and does not depend on the sample at all. What a sample of 120 rows reports is 13.3, and the reported crossing walks out towards the exact one as the sample grows: 13.3, 15.0, 16.4, 18.0. The mechanism is that the autocovariances the window is applied to are themselves attenuated, worst at the longest lags, and the window that discards those lags loses less of them. Where a taper's case begins

The error no window repairs

Every block window's best estimate of a long-run variance is wrong by about forty per cent at a hundred and twenty rows, and the largest part of that is not a bias at all. Choosing the window moves a twentieth of it.

Three estimates of the same effect, and the honest one is the worst. 8,000 trials in which every arm has an effect of exactly 0. The arm carried forward is reported by its first stage at +0.187 above the truth — it was chosen for being ahead — and by its second stage at -0.0032, which is unbiased by construction because the selection could not see it. The combination that gets published is at 0.123, exactly the share of the first stage's bias the first stage contributes. Root mean squared error: 0.240, 0.261, 0.182 — the unbiased estimate is the least accurate of the three. Designs that change while they run

The estimate after the choice

An arm chosen for being ahead is ahead by more than it should be, and the trial then publishes the average of the stage that chose it and the stage that did not. The unbiased estimate is the one built from a third of the data — and it is the least accurate of the three.

Where the maximum is, from 15 runs. One dataset, one fitted quadratic, and two answers to "where is the best setting". The delta method reports 0.80 ± 0.46, a finite interval it will report whatever the data does. Fieller's set is 0.49 to 1.76, because the curvature here has t = -4.04. The true optimum is at 0.75. The surface between the corners

The optimum is a ratio, and its interval is sometimes the whole line

The best setting is −b₁/2b₂: a ratio of two estimates whose denominator is a curvature the design can often barely see. The delta method reports a finite interval every time and covers 68.8% where the curvature is weak; Fieller's set covers 95% and says so by being unbounded.

Four sequences, and the rule only ever sees the last one. Under long memory at d = 4/9, four things that are all called the dependence. The law itself is the top line. What a sample of 120 rows reports on average is the second, computed exactly: subtracting a sample mean takes the first lag from 0.800 to 0.538. What a candidate's residuals report is the third, lower again at 0.472, because a fit removes dependence along with signal. The autoregressions are fitted to that third sequence and reproduce it exactly out to their own order — the Yule–Walker equations are solved to make it so — so everything they say past that is extrapolation. At the twentieth lag the law has 0.576, the residuals report 0.006, and an AR(8) extrapolates 0.028. The shape a dependence has

The order the tail is drawn at

A fitted autoregression reproduces the sample exactly at the lags it was fitted on, so everything it says past them is extrapolation — and the order is the dial that decides how much of it there is.

The correction does not arrive at the truth, it passes it. The average decay factor a forecast applies to the last observation, at φ = 0.85 and 50 observations, 3000 series per horizon. The middle curve is φʰ, what the model actually does. Below it is the uncorrected forecast, which uses φ̂ʰ and reverts too fast — 24.8% short at h = 4, 30.0% short at h = 6, 32.7% short at h = 8. Above it is the forecast built on the corrected estimate, which overshoots, and the reason is arithmetic rather than a bad correction: raising an unbiased estimate to a power does not give an unbiased estimate of the power, and the higher the power the more the spread of φ̂ is converted into overshoot. Comparing two forecasters

The repair that moves the wrong number

Correcting the bias in a persistence parameter is one line of arithmetic that works. Feeding the corrected estimate into a forecast repairs the number everybody looks at, makes the forecast worse by squared error at moderate persistence, and improves the interval for a reason that has nothing to do with bias.

A fit takes the low frequencies out of what it leaves behind. The autocorrelation of the errors, of the residuals of a fitted benchmark, and of those residuals rescaled by their own leverage. (I − H) removes the component of the errors lying in a column space that is itself slow-moving, so the residuals are less persistent at every lag — by 5.9% at the first and 26.6% by the fourth. The leverage correction is the standard repair for what a fit does to a residual's size; drawn here against what it does to a residual's dependence, it does nothing. Counting what is independent

The residuals are not the errors

A fit removes the part of the errors lying in its own column space, and a persistent design's column space is itself slow — so what is left behind is smoother than what went in, at every lag, by an amount that grows with the lag.

A thin enough set is not one set. Every admissible set of 14 units this table can enumerate, by how much of the assignment space it admits and how many pieces it falls into under single swaps. A walk is uniform on the piece it starts in and never leaves it. The pieces are not fragments: at 522 admissible assignments the set splits into 3 halves of exactly 520 each, and every assignment's complement is in the other half — no sequence of admissible single swaps takes an assignment to its own mirror image. Two-swap proposals reconnect four of the six disconnected sets here — the two they do not are the thinnest, where a two-unit move rarely lands anywhere admissible either — which makes a bigger proposal a correctness repair rather than the speed dial it was measured as. What a dictionary buys and what it costs

The walk that cannot cross

A thin enough admissible set is not one set. It splits into an assignment and its mirror image, no sequence of admissible single swaps joins them, and the walk that samples it is uniform on half the reference distribution for ever.

The false-positive rate against the number of analyses, on pure noise. The data has no effect in it. Each analysis is correct and each p-value is honest. With 20 correlated outcomes available, something reaches p below 0.05 58% of the time. Tests, and the second number

Twenty analyses of nothing

Twenty honest, correct analyses of data with no effect in it find something significant 57% of the time. Nobody p-hacked, every individual p-value is right, and the reported one is the smallest of twenty.

One group 6 population widths from the rest. Squared error for each group under partial pooling, with each group's own mean beside it. Seven of the eight are estimated better by pooling. The eighth, which was never from the population, is estimated 6.1 times worse — 13.9 against 2.3. Groups that borrow

When borrowing goes wrong

Partial pooling wins on the total and can lose badly on one group. Placed six population widths out, the group that was never from the population is estimated six times worse than by its own mean — and nothing in the output says so.

Where the constraints exhaust the randomisation. At 16 units there are 12,870 equal splits, so the ones meeting a stated tolerance can be counted rather than estimated. With each of the first k standardised imbalances required to be within 0.4 of a coin's own spread, the admissible count runs 3874 → 1006 → 314 → 0 → 0 → 0 — and at 4 functions there is no admissible assignment at all. The count is the number of distinct answers a randomisation test can give: at 3 functions its finest attainable p-value is 1 in 314. Balance improves with every constraint and the reference distribution shrinks with it, and the two run out at different rates. Choosing what the rule reads

When the constraints run out

Every function added to a basis is a constraint the assignment has to satisfy with the same units. At sixteen units and a stated tolerance the admissible assignments run 3,874, then 1,006, then 314, then none — and the count is exact, because the assignment space is finite.

The same dial, on a list that steps by one. The probability that a per-candidate tuning list changes which candidate the table selects, at each of 5 separations, for both tuning parameters at a matched list length of 8. The sieve order runs from 13.1% to 2.6%, a factor of 5.00; the whitening window, from 16.4% to 1.5%, a factor of 11.75. What is held is the number of options, the table, the law, the sample size and the seeds; what cannot be held is the size of a step, since an integer step and a geometric step are different amounts of change. The dial moves both, and it moves them by 2.35 times as much on one as on the other. What decides whether a tuning list decides

A step that is not a ratio

Run the separation sweep on a tuning list of integers rather than a geometric ladder and the two factors still point opposite ways. The invariant does not survive: along a row of integers the probability moves by 2.163 where along the geometric ladder it moves by 1.208.

Either model is enough; neither is not. The bias of three estimators of an average effect of 1.0000, over 600 samples of 600 units with the assignment rule at strength 1, in each of the four cells made by getting each nuisance model right or wrong. The wrong model in both cases is one that omits the second covariate, which the outcome and the assignment both depend on. The outcome model alone is off by 0.8064 whenever it is the wrong one; weighting alone is off by 0.8190 whenever the propensity model is. The augmented estimator built from both is off by -0.0085, -0.0083 and -0.0016 in the three cells where at least one of them is right, and by 0.8118 in the fourth — which is between its two components rather than better than either. Weighting one sample into another

Either model, but not neither

The augmented estimator's bias is −0.0085, −0.0083 and −0.0016 wherever one nuisance model is right, against components off by 0.8064 and 0.8190. One step past the overlap sweep it is the least biased estimator on the table at 0.0857 and the worst on it at 1.9265.

Two groups, a baseline and a follow-up, and nothing happening in between — baseline reliability 0.6. 600 units in two pre-existing groups whose true means are 1.00 apart, read once at baseline and once at follow-up, with no change for anybody. The two groups' mean changes are −0.075 and −0.032, so the change-score analysis reports a group difference of 0.043. The regression of follow-up on baseline and group reports 0.409, against a closed form of (1 − λ) × 1.00 = 0.400: at any one baseline reading the two groups' lines sit that far apart, because each group's units regress towards their own group's mean. The pooled slope in this sample is 0.614, the baseline's reliability. Reversals that are not errors

Two analyses of one baseline

Two groups read at baseline and again at follow-up, with no change for anybody. Subtracting the baseline reports a group difference of −0.0014 and adjusting for it reports 0.4008 — and each analysis is exactly right about one reason the groups started apart and wrong by 0.40 about the other.

The fixed-width trial's coverage when the outcomes are not normal, for both stopping rules. normal: stopping on the arms 94.05% after 18.1 blocks, on the report 89.95%; log-normal, skewness 0.95: stopping on the arms 94.70% after 18.5 blocks, on the report 90.80%; log-normal, skewness 2.26: stopping on the arms 94.15% after 19.3 blocks, on the report 90.25%; log-normal, skewness 4.75: stopping on the arms 94.45% after 18.7 blocks, on the report 90.50%; t, five degrees of freedom: stopping on the arms 94.35% after 18.3 blocks, on the report 90.30%; skewness 4.75, arm A only: stopping on the arms 93.80% after 26.0 blocks, on the report 89.90%; skewness 4.75, arm B only: stopping on the arms 93.60% after 14.2 blocks, on the report 89.95%; equal variances, normal: stopping on the arms 94.75% after 11.4 blocks, on the report 90.90%; equal variances, skewness 4.75: stopping on the arms 94.05% after 11.1 blocks, on the report 92.00%. When a fixed width is reached

A width rule on skewed outcomes

The blinded fixed-width rule rests on a within-arm spread being independent of the arm means, which only normal samples guarantee. On outcomes with a skewness of 4.75 the independence fails and the overall coverage barely notices — 93.60% to 94.70% across every shape counted, against 94.05% on normal outcomes. What skew moves is the runs that stop by twelve blocks, which cover about 90% with the skew in one arm, and the trial's length: a variance ratio corrected on normal theory lengthens it from 18.1 blocks to 26.0 with the skew in the first arm and shortens it to 14.2 with the skew in the second.

Where the bias lands. The drift in the log variance ratio, fitted across 12 blocks over 4000 trials. E[log λ̂_b] is log λ_b plus ψ(k_B/2) − log(k_B/2) − ψ(k_A/2) + log(k_A/2), which depends on nothing but the degrees of freedom — so the tempting sentence is that it goes into the intercept and leaves the slope alone. It does not, because the blocks alternate between allocations and the alternation is correlated with the covariate being fitted: the lopsided blocks carry 0.5383 of bias and the even ones carry none. Uncorrected the slope reads 1.5597 against a truth of 1.5, which is 8.0 standard errors. Subtracting the two digammas block by block leaves 1.4976. What a block may vary

The bias that lands in the slope

The bias in a log variance estimate depends on nothing but its degrees of freedom, so it goes into the intercept — unless the degrees of freedom alternate with the design, which is exactly what a block-randomised trial makes them do.

The ceiling a multiplier cannot reach past. A wild-type resampling forms e*_t = e_t·w_t with the multiplier independent of the residual, so what comes out has autocovariance γ_resid(k)·γ_w(k) — the residuals' own, multiplied by the multiplier's. Since |γ_w| ≤ 1 the reference distribution's dependence is bounded above by the residuals', and the residuals' is already below the errors'. The two shortfalls compose. For a block of ℓ the multiplier's autocorrelation is exactly the triangle (1 − k/ℓ)⁺, drawn here as the dashed prediction against the realised resamples at ℓ = 5; the bound is attained only at ℓ = n, where the reference distribution is built from one sign. Counting what is independent

What a multiplier cannot keep

Two reasons were named for the quarter a blocked resampling falls short, and taking either away makes the gap larger. What is left is a bound — a multiplier can only take dependence out, and the residuals' own is already below the errors'.

Where the bootstrap works and where it does not. Uniform data on [0, 1]. For the mean the percentile bootstrap covers 93.5%. For the maximum it covers 0.0%, because a resample can never contain a value larger than the largest one observed, so the interval cannot reach above it. Intervals, counted

Where the bootstrap lies

Resampling is the most generally useful trick in the subject and it has a failure mode that is easy to state: it cannot see past the data. For a statistic that lives at the edge of the sample, coverage collapses from 95% to almost nothing.

The pairing recovers most of it and passes nothing. How much of the separating direction each probe carries, over 100 designs of 14 units whose admissible set is enumerated and split in two. The four rows the pairing adds are the dominant direction of what each blocking matrix keeps past its degrees, and the cut that direction's signs induce. Counted, they read 0.5266 and 0.5258 against the counted per-unit share's 0.3854 — most of the gap between that share and the design's own leverage at 0.5395, closed. Modelled, they read 0.4274 and 0.4954 against 0.4272. Nothing built from the active set passes leverage, and the projected fourth power is still ahead of all of them at 0.6583. A probe from what the rule blocks

A set of pairs, not a vector

The active set is a graph on the units, and every probe built from it so far has been its degree. Read as a graph it recovers 0.1326 of the alignment the summary lost — and draws level with leverage rather than passing it.

The line is the sample, and the sample is the finding. 900 draws of two independent standard normal causes, with the 453 of them past a threshold of 0.00 marked and the 447 that fall short left pale. In the population the two are independent by construction. Inside the selected sample the correlation is -0.4669 in closed form and -0.5050 counted on these 453 rows, and the least-squares line through them has a slope of -0.545. The mechanism is visible in the picture rather than argued: the threshold removes one corner of the cloud, and a cloud with a corner missing is a cloud whose two coordinates carry information about each other. What conditioning on a variable does

The sample is a condition

Two independent standard normals, selected on their sum exceeding its median, read a correlation of exactly −1/(π − 1) = −0.4669 inside the sample. Nothing is measured badly and nothing is missing — and both halves of that split read it, in the same direction, while the population containing both reads zero.

A cohort screened once, the top tenth enrolled, and followed up with nothing given — correlation 0.6. 2000 people read once at screening and once at follow-up, with a test–retest correlation of 0.6 and no treatment. The 215 above a cut at the top ten per cent of one reading (1.282 standard deviations) are enrolled. Their mean screening reading is 1.744 and their mean follow-up reading 0.982, a fall of 0.762 ± 0.055 with nothing done to anyone. The closed form for the fall is (1 − ρ) times the truncated-normal mean, (1 − 0.6) × 1.755 = 0.702. Reversals that are not errors

The measurement that got them enrolled

Enrol the top tenth of one screening reading and give them nothing, and they fall by 0.702 standard deviations at follow-up. Measured from a fresh reading taken after enrolment they fall by nothing. Averaging ten screening readings still leaves 0.101, and it takes twenty-one to get under 0.05.

The risk of one cause, estimated two ways. Two causes of an ending event with constant hazards 0.2 (the one of interest) and 0.3 (the competitor), random dropout at 0.1 and follow-up to 6. The lower line is the cumulative incidence, (0.2/0.5)(1 − e^(−0.5t)), the chance of actually having had this event by t; the dots on it are the Aalen–Johansen estimate over 2000 studies of 300, 0.3670 at t = 5 against 0.3672. The upper line is 1 − e^(−0.2t), and the dots on it are one minus Kaplan–Meier with the competing event treated as censoring: 0.6318 at t = 5 against 0.6321. The second is larger by a factor of 1.722 at t = 5, and it is not an error of estimation. It estimates, correctly, the risk in a population where the competing cause does not exist. When the data stops early

One minus Kaplan–Meier is not a risk

With two ways for observation to end, one minus Kaplan–Meier for one cause reads 0.6318 at t = 5 where the chance of actually having had that event is 0.3670. Added across the two causes, the complements reach 1.4088 — more than the whole cohort. Nothing is estimated badly: the complement estimates, correctly, the risk in a world where the other cause does not exist.

Estimating a weight you already know is worth doing. The variance of an inverse-probability estimate weighted by a propensity fitted from the sample, over the variance of the same estimate weighted by the true propensity, paired on the same 500 samples of 600 units at each of five settings. Every reading is below one: the stabilised estimator keeps 27.8% of its true-weight variance where the assignment is nearly a coin toss and 72.0% where it is nearly decidable, and the unstabilised one 30.0% and 49.3%. Neither estimator is materially biased, so this is a variance rather than a trade. The true weights are right about the population and know nothing about the draw; the fitted weights are the value that sets this draw's own imbalance to zero, and that imbalance was what the variance was made of. Weighting one sample into another

The estimated weight is the better one

The propensity is known exactly here, so it can be weighted by — and estimating it from the same data and weighting by that gives a variance ratio of 0.4769 on paired draws. The reason is a projection: the draw's own imbalance explains 56.33% of the true-weight variance and 0.05% of the estimated-weight one.

What each estimator costs, against how many groups there are. Each point is 4,000 datasets, with the population spread estimated from the data rather than supplied. Partial pooling first beats BOTH of the estimators it sits between at 5 groups; below that, complete pooling — which estimates nothing at all — is the better answer. Its own cost falls from 1.751 at 2 groups to 0.795 at 40. Groups that borrow

The fewest groups that can borrow

At three groups the estimator that shrinks towards its own data's mean returns the group means untouched, on every dataset, because its constant is J − 3. At two it expands instead of shrinking. And the number of groups at which partial pooling starts to be worth doing is five, or two, or never — it depends on how far apart the groups are.

What a 8-run fraction of 4 factors confounds. The defining relation is I = ABCD, so the resolution is 4. A is estimated as A + BCD; B is estimated as B + ACD; C is estimated as C + ABD; D is estimated as D + ABC. Each of those is an identity about the design rather than an approximation about the data. Decided before the data

The word a fraction costs

A half fraction estimates each main effect as an exact sum of that effect and everything it is confounded with — no error term, no sample-size argument. With every interaction at 0.8 the design reports a true effect of −1 as −0.20, and the design cannot test the assumption that makes the number mean anything.

The average decay factor each route produces, φ = 0.85, 6 steps ahead. The truth is φ^6 = 0.3771. no correction averages 0.2616 with a spread of 0.1646 and a squared forecast error of 3.2516; the formula, on the persistence averages 0.4213 with a spread of 0.2528 and a squared forecast error of 3.4827; the bootstrap, on the persistence averages 0.4355 with a spread of 0.2655 and a squared forecast error of 3.5120; the bootstrap, on the decay factor averages 0.3375 with a spread of 0.2278 and a squared forecast error of 3.4132. 800 series, 100 bootstrap refits each. Comparing two forecasters

Correcting the forecast instead

The complaint against the usual repair is that a correction aimed at the persistence lands on the wrong quantity. Aiming it at the decay factor the forecast actually uses fixes exactly that — the error stops compounding with the horizon, 69.7% becomes 9.5% at twelve steps — and the forecast still gets worse.

What the fit calls the shape, against what it is. One eigenvalue held at −3 and the other swept from −2 to 2, so the truth is a maximum on the left and a saddle on the right and the change happens at exactly zero. At an eigenvalue of −0.25 — a genuine maximum — the fit reports a saddle on 26.4% of studies; at +0.25 — a genuine saddle — it reports a maximum on 25.1%. The standard error of a squared coefficient under this design is 0.3791, and the region of confusion is about that wide either side of zero. The surface between the corners

The sign the curvature has

A fitted surface reports a maximum, a minimum or a saddle, and the report is a comparison of two estimated eigenvalues against zero. At a true second eigenvalue of −0.25 the fit calls a genuine maximum a saddle on 26.4% of studies, and at +0.25 it calls a genuine saddle a maximum on 25.1%.

An exact test rejecting a true hypothesis a fifth of the time. How often each analysis reports an effect when the average treatment effect is exactly zero and the effect varies between units, at 150 units with 25% treated. The permutation test on the difference in means reads 4.20% where the effect is constant — where the two nulls coincide and its exactness applies — and 22.93% where the effect varies with a standard deviation of 3. The same test on the studentised difference reads 6.27% there, and the ordinary large-sample t, which makes no exactness claim at all, reads 6.60%. The reference distribution the design supplies

The null the exactness is for

A permutation test is exact under the hypothesis that the treatment changed nothing for anybody. Under the hypothesis it changed nothing on average, with a quarter of the units treated and the effect varying between them, it rejects a true null 22.93% of the time.

What the guess is worth, when it is worth anything. The variance cost of an even split relative to the variance-minimising one for a risk difference, against the first arm's proportion, with the second at 0.3. The cost is a pure number: it does not depend on the trial's size. It is exactly zero at 0.3 and at 0.70, where the two arms have the same p(1 − p); it is 0.19% at a half and 4.36% at a tenth. Across the whole range from a tenth to nine tenths it never exceeds 4.36%, which is what the variance-minimising rule is worth here — and what it is worth is the reason it is safe to use with a guess. Splitting the units

The arm whose variance is its answer

With a binary outcome the allocation rule is a function of the proportions the trial exists to estimate. It costs at most 4.36% of variance to ignore it anywhere between a tenth and nine tenths, because √(p(1−p)) stays within a factor of two of its peak across 98% of the unit interval.

Reading the draw changes what is charged, not what is tracked. The correlation between the band width each rule picks and the best band width on the same draw, over 400 draws. The three fixed charges read -0.069, -0.066, 0.012. The three that read the sample read -0.012, -0.019, -0.041. None of the six is distinguishable from nothing. The statistic the first plug-in reads does vary — the draw's own summed squared autocorrelation runs from 2.06 to 10.19 with a mean of 4.36 — so the failure is not that the charge stopped moving. It is that what it moves with carries no information about which width this draw wanted. A charge that is not a straight line

A charge that reads the draw

Three charges built to read the sample track the best band width on their own draw at −0.012, −0.019 and −0.041, deliver more error than the fixed rule they are calibrated to, and pick a width half again as variable. The statistic moves; the answer does not.

The covariate the treatment caused, and what it hides. A covariate on the causal path: the treatment causes it and it causes the outcome, so the treatment's total effect of 1.130 runs partly through it. Adjusting for it returns the direct edge alone, 0.500, which is what somebody wanting the total effect should not have asked for. The dashed variable is the second problem: an unmeasured cause of both the covariate and the outcome. It does not touch the treatment, so the unadjusted regression still recovers 1.130 exactly. It does touch the covariate, so once the covariate is conditioned on the treatment and the outcome are linked through it, and the adjusted coefficient lands on 0.050 — neither the total effect nor the direct one. What conditioning on a variable does

The variable the treatment caused

Adjusting for a covariate the treatment caused stops estimating the total effect and starts estimating the direct one. When that covariate shares an unmeasured cause with the outcome it estimates neither: the total effect is 1.1300, the direct effect is 0.5000, and the regression returns 0.0500.

One row at x = 9 on a wrong line, fitted three ways. Least squares gives slope −0.389, Huber 0.171 with the extra row at weight 0.115, and least trimmed squares 0.420, fitted to the 12 rows it keeps. The twenty clean rows alone give 0.495. Open circles are the rows the trimmed fit leaves out. Regression, and what the summary hides

A robust loss and a far x

One far row drags least squares to a slope of −0.389. Huber's loss, the standard robust line, reaches only 0.171, and carried further out the same row gets its full weight back. Least trimmed squares reads 0.420 at every distance, and at the normal model keeps 7.13% of least squares' efficiency to do it.

The top of a heavy-tailed population keeps its lead; the top of a light-tailed one gives it back. Select the top share on the first reading and read the group again: the share of its mean lead the second reading keeps, by integration over the true score (lines) and counted on 400,000 draws a parent in 20 batches (points, with two standard errors). The normal keeps exactly 0.6 at every selection. At the top half the Laplace keeps 0.541, the t 0.535 and the uniform 0.648 — the heavy tails keep LESS than the correlation. By the top one per cent the order has reversed: 0.761, 0.784 and 0.443. At one in ten thousand the t keeps 0.977 and the uniform 0.346. Reversals that are not errors

A lead that a heavy tail keeps

Four populations whose readings all correlate at exactly 0.6, and whose least-squares slopes all read 0.6. Select the top one per cent on one reading and measure them again: they keep 60% of their lead if the true scores are normal, 76.1% if they are Laplace, 78.4% if they are a t on four degrees of freedom — and 44.3% if they are uniform. The correlation predicts the regression of the extremes for one shape of population only.

Weights that balance a sample by construction. What three sets of weights leave of the standardised difference between the arms on each covariate, as a root mean square over 1200 samples of 600 units. The true propensity leaves 0.1317 and 0.1186 — a sampling error, since it is right about the population and knows nothing of the draw. A likelihood fit leaves 0.0770 and 0.0657, having absorbed part of the draw's imbalance as a side effect of fitting the treatment. Weights fitted so that each arm's weighted means are the sample's leave 1.4e-14 and 1.2e-14, which is the arithmetic's floor rather than a small number: the largest gap between a weighted arm mean and the sample mean in any draw is 9.8e-14. Weighting one sample into another

A weight fitted to balance

Weights fitted so that each arm's weighted covariate means equal the sample's leave a difference of 1.4×10⁻¹⁴ between the arms and give the estimate a third of the variance of weights fitted by likelihood — 0.011883, within a relative 5.8% of the bound no estimator can beat. In the world where the assignment carries a square nobody named, the same exact balance leaves the square further apart than no weighting at all, and where the outcome carries it too the estimate is wrong by 0.6973 with an interval that covers 1.5%.

Five treatments of an estimate above one, φ = 0.95, n = 25. The correction exceeds one on 31.1% of series at this setting. left where it lands: squared forecast error 12.828, average decay factor 0.7974 against a true 0.7351; capped at 0.995: squared forecast error 5.680, average decay factor 0.5950 against a true 0.7351; capped at 1 − 1/n: squared forecast error 5.535, average decay factor 0.5256 against a true 0.7351; correction scaled to fit: squared forecast error 5.535, average decay factor 0.5256 against a true 0.7351; correction refused where it leaves: squared forecast error 5.868, average decay factor 0.4423 against a true 0.7351. Comparing two forecasters

The correction that leaves the region

The bias correction adds (1 + 3φ̂)/n whatever φ̂ is, so it pushes the estimate above one whenever φ̂ exceeds (n − 1)/(n + 3) — on 31.1% of series at φ = 0.95 and twenty-five observations. Five obvious things to do about it differ by a factor of 2.3 in squared forecast error, and none of them is documented as a choice.

Six cells, and 5% is the right answer in all of them. How often a regression between two independently generated series is called significant at the 5% level, for two worlds and three treatments, at 200 observations. Every pair is independent by construction, so 5% is correct everywhere and every other reading is a failure. Untreated: 82.9% and 100.0%. With a fitted line removed: 74.2% and 33.5%. Differenced: 5.0% and 5.2%. The treatment that controls the rate in both worlds is the one that discards the level and the trend, which is the quantity a study of trending series was about. When the observations repeat each other

The repair that keeps the question

A regression between two independent trending series is significant 82.9% of the time on random walks and 100.0% on trend-stationary ones. Subtracting a fitted line leaves 74.2% and 33.5%; differencing leaves 5.0% and 5.2% and throws away the trend the study was about.

Three contrasts on one dataset, three different splits. The variance-minimising allocation for each of three ways of reporting the same two-arm comparison, against the first arm's proportion, with the second at 0.1. A risk difference wants the arm with the larger p(1 − p) to get more units; a log odds ratio wants it to get fewer, and the two curves are exact reflections of each other in the half line. A log risk ratio wants something else again. At a first-arm proportion of 0.6 they ask for 62.0%, 21.4% and 38.0% of the units. A trial reporting more than one of them cannot be optimal for either. Splitting the units

Two contrasts, one split

A risk difference wants 62.0% of the units in the first arm, a log risk ratio wants 21.4% and a log odds ratio wants 38.0% — on one dataset, with one pair of proportions. The difference's rule and the odds ratio's are exact reflections of each other, so no split can be near-optimal for both.

One row at x = 9 on a wrong line, fitted three ways. Least squares gives slope −0.389, Huber 0.171 with the extra row at weight 0.115, and least trimmed squares 0.420, fitted to the 12 rows it keeps. The MM-estimator carried on from the trimmed fit gives 0.479, with the extra row at weight 0.000 and the scale fixed at 0.319. The twenty clean rows alone give 0.495. Open circles are the rows the trimmed fit leaves out. Regression, and what the summary hides

The start an efficient robust line inherits

The MM-estimator carries a trimmed fit on through a redescending loss, and it does what it promises on one far row: slope 0.479 at every distance, the row at weight exactly zero, and 87.2% of least squares' efficiency at twenty rows. What it cannot do is choose. At eight far rows of twenty the exact trimmed fit picks the wrong half on 111 datasets; the efficient step repairs none of them, spoils none of the other 89, and ends nearer the wrong line than the start did.

One weighting told the means and one told the second moments, in five worlds. The bias of the fit to balance over 600 samples of 600 units in each world, fitted to the covariates' means and fitted to their means, squares and product. Told the means it is off by -0.0004, -0.0020, 0.0103, 0.6973, 0.2698 in the worlds with no square, a square in the assignment, a square in the outcome, a square in both and a cube in both; told the second moments, by -0.0009, -0.0000, 0.0009, -0.0045, 0.3099. Its interval covers 94.0%, 94.7%, 95.0%, 1.5%, 51.0% and 93.7%, 89.8%, 94.0%, 91.0%, 48.3%. Weighting one sample into another

The moments a balance is told

Weights fitted to balance the covariates' means were wrong by 0.6973 in the world where both the assignment and the outcome carry a square. Told the squares and the product as well, the same construction is off by −0.0045 there and its interval covers 91.0%. The failure moves up a moment rather than away: with a cube in both, the second-moment balance is off by 0.3099 and leaves the cube twice as far apart as no weighting. And where overlap is thin, 37.0% of samples have no such weights at all.

Where each criterion's optimum puts the information. the D-optimal design's smallest eigenvalue is 0.09927, attained once; the A-optimal design's smallest eigenvalue is 0.16516, attained once; the I-optimal design's smallest eigenvalue is 0.17541, attained once; the E-optimal design's smallest eigenvalue is 0.19999, attained 2 times. A criterion that reads the smallest eigenvalue has no derivative where that eigenvalue is repeated, and the E-optimal design is exactly there. A design chosen rather than looked up

The criterion with no derivative

E-optimality maximises the smallest eigenvalue of the information matrix, and at its own optimum that eigenvalue is attained twice — which is exactly where the function has a corner. The multiplicative search this field's other three criteria use assumes a derivative that is not there, and stops at 37.2% of the optimum.

What a confirmation run at the chosen setting would find. The true optimum is worth 62.348. At σ = 2 the fit predicts 62.679 at the setting it recommends and the truth there is 61.821 — a gap of 0.858, which is 0.72 of the prediction's own standard error. The setting itself gives up 0.527 against the best available. The surface between the corners

The run that confirms it

The setting a response-surface analysis recommends was chosen because the fitted surface was highest there, so the height the fit predicts at it is a maximum over a random field. At twice the noise the fit predicts 0.858 more than is there — 0.72 of the prediction's own standard error — and the gap is not noise, it is the selection.

One of these converges and the other does not. Two measurements on the same fits, against the sample length, for a system with 2 genuine relations. The distance from the fitted plane to the true plane falls from 0.1438 at 100 observations to 0.0075 at 1600 — halving with each doubling, which is the 1/n rate this field's estimates converge at. The angle between the leading fitted relation and the leading generating one reads 29.6° and 29.0° at those same lengths, and is flat in between. The plane is an estimate; the relation inside it is not. Three series, and a count

A space is not a relation

The fitted plane approaches the true one at rate 1/n — 0.1438 at a hundred observations and 0.0075 at sixteen hundred. The angle between the leading fitted relation and the leading generating one reads 29.6° and 29.0° at those same lengths, and never moves.

The law is the eigenvalues, and nothing else. The mean and the skewness of n(ĝ − g) at the stationary point, measured over 40,000 draws, against the closed forms ½ Σλ and 2√2 Σλ³ ⁄ (Σλ²)^(3⁄2). The worst disagreement anywhere is 0.028. In one variable the second-order law is a single χ² and its sign is the sign of g″; here it is a weighted sum with the Hessian's eigenvalues as weights, so a bowl and a valley differ in both moments and a saddle has both equal to zero. The distribution itself

A flat point with more than one direction

At a stationary point of a function of several means the second-order law is ½ Z′HZ, so the bias is half the Hessian's trace — 2.008 for a bowl, 5.028 for a valley, and −0.006 for a saddle, where the eigenvalues cancel. The saddle's coverage is the worst of the three.

All themes