Concept

Monte Carlo — where it appears

Answering a question by simulating the process many times and counting. Its error falls only as the square root of the number of runs, so a quantity computable in closed form should be computed rather than counted — and where both are available the disagreement between them is the check.

Named by 143 essays across 51 fields — each of them below, with the objects they name alongside it.

The six best bases of 2 functions, and what each protects. Every cell is R²(g | span B) — the share of the imbalance in that shape a rule balancing that basis removes — computed from exact inner products between Hermite functions and indicators, with nothing simulated. The rows are ordered by their worst cell, which is the number an experimenter who does not know the shape is exposed to. The best row here guarantees 26.8% against every shape in the list, and the worst of the six guarantees 15.1%: the difference between them is entirely which subspace was picked, at the same cost per arrival.

A basis is a subspace

A balancing rule cannot tell one basis from another with the same span, so choosing what to hand it is choosing a subspace — and then what it removes of any outcome shape is a projection, computable exactly, with no trial anywhere in it.

basis · Criterion
One experiment, with the blocks getting smaller as the target comes into range. A single run at a requirement of 0.25, with the block sizes 5, 5, 11, 25, 11, 8, 3, 2 and a total of 70 observations in 8 blocks. The rule stops when the observations in hand reach z²σ̂²/d², with σ̂² pooled from the within-block contrasts — an estimate that moves as the run goes on, so the target moves too. Early blocks are large because the target is far away and cannot be overshot; late ones are small because a block is the granularity of the answer. The interval afterwards is built from the 8 block means and from nothing the rule looked at, and it has 7 degrees of freedom against the rule's 62.

A block size that changes

The blinded rule's exactness never needed the blocks to be the same size. Letting the size be chosen from the contrasts as the run goes on leaves the coverage exactly where it was — and runs straight into an identity that says what a schedule can and cannot buy.

pace · Stopping
What a median split can see. A standard normal covariate with its median marked, and the mean of each category as a vertical rule: -0.7979, 0.7979. A rule that balances the categories is balancing those numbers and nothing else, so the part of the covariate it can act on is the variance between them — 0.6366 of the total, which at two categories is exactly 2/π because the two half-normal means are ±√(2/π). The rest, 0.3634, is variation inside the categories that the rule cannot see and does not touch: the assignment within a category is still a coin. Everything the next figure measures is a consequence of this one, and it is available before any unit has arrived.

A covariate with no levels

Every balancing rule on this site reads a level. Age and blood pressure have none, so somebody cuts them into categories — and a median split can see exactly 2/π of a normal covariate, whatever the rule does with the halves.

continuous · Assignment
What each rule gives up against an oracle that is arithmetic. Expected squared error of the candidate each rule selects, minus the expected squared error of the best candidate in the table, over 500 draws of 120 rows. Both quantities are closed forms — σ_S²(1 + q/(n − q − 1)) — so the only Monte Carlo here is over which candidate got picked. The hold-out spends half its rows measuring what the criterion computes, and pays 1.8 times as much for it. Schwarz's criterion is worst because it is answering a different question: which candidate contains the truth, rather than which one forecasts best.

A criterion is a prediction of the hold-out

A rolling hold-out spends half the sample measuring what a criterion computes from all of it. Against an oracle that is arithmetic rather than an estimate, the criterion gives up 0.01701 and the hold-out 0.03200 — and the number the hold-out reports for its own winner is optimistic by more than either.

proxy · Order-selection
A proposal that moves more, refused more often. The two halves of the trade, both exact, on the 410 admissible assignments of twelve units. The integrated autocorrelation time of an imbalance the rule was never handed falls from 7.30 at one swap to 3.97 at three, and the acceptance rate falls with it, from 58.8% to 40.8%. A rejected proposal costs one evaluation and leaves the chain where it was, so acceptance is not the price of anything and the ranking by acceptance is the reverse of the ranking by cost. Past three the family folds: exchanging k of six from each arm is the complement of exchanging six − k, so k = 5 has the same 36 proposals as k = 1 and k = 6 has 1.

A proposal that moves more than two units

The walk's autocorrelation is a fact about its step size and not about its acceptance rate. Exchanging three units from each arm mixes nearly twice as fast as exchanging one, and is refused a third more often.

blocks · Randomisation
One factor moves and the other does not. The two factors of the same average, each drawn against its own largest value so that they share an axis. The rate at which the five candidates disagree about the tuning parameter rises from 28.6% at 4 values on the list to 43.3% at 8, a factor of 1.52. What a disagreement costs, given that there was one, is 0.00975 ± 0.00224 and 0.00848 ± 0.00113 at the same two points — 0.5 standard errors apart, and the paired comparison on the draws that disagree under both lists puts it the other way. The guess this field was written to test was that a longer list makes disagreements commoner and each one smaller. The first half is right and there is no second half.

A rate times a size

A sweep reported what it costs to let every candidate choose its own tuning parameter and found it flat across the list. It was reporting a product, and the two things multiplied together do not behave the same way at all.

apiece · Order-selection
Three intervals, one shortfall. What each of three intervals actually covers, at four rules and two block windows, over 300 samples of 120 rows. All three are built from the same resamples on the same draws, so a difference between them is a difference in what is done with the resampled series. Not one of the twenty-four cells reaches the ninety-five per cent it promises. The studentised interval runs from 75.7% to 92.3%, the percentile interval — the earlier field's — from 80.0% to 89.7%, and a normal interval on the same scale from 81.7% to 89.0%. The standard repair for a percentile interval's shortfall does not repair it.

An interval that carries its scale

A percentile interval inherits the resampled distribution's skewness and its scale error together. The standard repair is one extra variance per resample. It was named and not run, so this runs it.

student · Bootstrap
The number the comparison was missing. What it costs to choose the tuning parameter for every candidate separately rather than once for the table, under AR(1) at 0.8, paired on the draw. The window's figure is the one the earlier field reported; the order's is the one it named and did not make. They are the same size — 0.00401 against 0.00360, at 2.30 and 1.72 paired standard errors — and matching the lists at eight values leaves them the same size again. The prediction that the longer list would make the order's cost the larger of the two is not what happens; what happens is that the two rules cost the same once they are scored by the same criterion, which took a missing term to arrange.

The comparison that was not made

Choosing a whitening's window separately for every candidate costs 0.00401 of regret. The same question about an order was named and left, because the two lists are different lengths. The order's answer is 0.00360, and matching the lists changes almost nothing.

lists · Order-selection
The eighth was not a constant. The probability that a per-candidate tuning list changes which candidate the table selects, at each of 5 separations between the candidates, over 800 draws apiece. The earlier field reports this flat at about an eighth across list length, on a table and a world it never varies. Vary how much the omitted coefficients are worth — one multiplier, with the table, the list, the law and the sample size all held — and it runs from 17.6% to 1.5%, a factor of 11.75. The world in which every candidate is true is the world in which the tuning list decides most; the world in which one candidate dominates is the world in which it decides nothing.

The eighth that was not a constant

How often a per-candidate tuning list changes which candidate wins is reported flat at about an eighth across list length. Vary how far apart the candidates are instead and it runs from 17.6% to 1.5%.

turnover · Order-selection
The fourth-order expectation, by two routes. Every inner product in an eight-term slice of the dictionary at ρ = 0.5, computed from the linearisation and Mehler's formula and counted from two hundred thousand draws of a correlated pair. The entries that matter are the ones off the main effects: ⟨f(X)u(Y), g(X)v(Y)⟩ is a fourth-order expectation, which the independent-covariate field could not write down. The worst departure is 1.99 standard errors over 36 pairs, measured in each pair's own error because the entries differ in size by two orders of magnitude.

The fourth moment that was missing

Mehler's formula makes the main effects exact at any correlation and stops there, because the interactions need an expectation of four Hermite functions rather than two. A linearisation turns the four into two, and the whole geometry becomes closed again.

joint · Blocking
What a sample shows, and what the algebra does. The difference between a rectangular block's implied long-run variance and a trapezoidal one's, as a share of the truth. Above the axis the rectangle is less biased and below it the trapezoid is. The heavy line is exact — computed from the law's own autocovariances — and it crosses at 19.2. The others are what samples of 120, 240, 480, 960 rows report, and every one of them exaggerates whichever window is ahead: at ℓ = 20, where the exact difference is 0.28 points, a sample of 120 rows shows 4.31 points — 15 times larger. That is the number the earlier reading of this comparison was missing: three tenths of a point is what the algebra says and not what a hundred and twenty rows report.

The gap a sample shows

The exact difference between two block windows at a block length of twenty is three tenths of a point. What a hundred and twenty rows report is four and a third, because the autocovariances the window is applied to are attenuated too.

crossing · Bootstrap
A quantile is the dearer reading, everywhere. The error each rule and window delivers on the two error readings, over 400 draws. The lower pair of lines is the implied long-run variance — the instrument the earlier field uses — and the upper pair is the 95% point of the standardised resampled mean, read against the finite-sample truth of 3.889 found by simulating the law directly. The quantile costs more at every one of the eight cells: at the plug-in rule it is 59.1% against 45.3% for the taper. That is not a defect in the bootstrap; a quantile is a statement about the shape of a distribution as well as its scale, and a fixed number of resamples estimates a tail worse than a variance. What matters for the comparison is that the two orderings between the windows are not the same, which the margins figure is about.

The instrument and the reading

Every comparison between two block windows in this collection is an error in an implied long-run variance. Nobody reads a long-run variance. Read on the 95% point a test uses, the same bootstrap costs half as much again.

readout · Bootstrap
Three rules and a target none of them is aimed at. Which block length each rule picks, over 400 samples of 120 rows, for the tapered window. Two of the rules are points: a length written into a protocol is 8.00 on every draw and the rule of thumb is 4.00, because n to the one third does not read the data at all. The plug-in reads the sample's own persistence and lands at 14.36 with a standard deviation of 2.93. The length that would actually have been best on that draw averages 24.57 with a standard deviation of 16.23 and runs from 10 to 48 between its tenth and ninetieth percentiles. The target moves five times as much as the best estimate of it does, which is why no rule can be close to it and why the two that do not try are not merely worse — they are somewhere else.

The length nobody has

Every comparison of block windows in this collection is made at each window's own best block length. That length has a standard deviation of sixteen across draws and averages twenty-five. No rule is aimed at it.

feasible · Bootstrap
The curvature is in the denominator. The optimism a Bartlett band of each width actually costs, divided by that width, on two ways of measuring the width, over 2000 draws at 120 rows. Measured in the weights the band spends — Σ w(k), which is what the earlier field levies its charges on — the reading falls from 0.9528 at two lags to 0.7486 at thirty, so a charge proportional to the summed weights is too dear at one end and too cheap at the other. Measured in the pairs the band uses — Σ w(k)(1 − k/n), because a lag of k is an average over n − k products — the same readings are flat from 4 lags up, at 0.0084 of χ² per width against 0.2359. The correction has no fitted parameter in it: it is a function of the window, the width and the sample size.

The width a band is measured in

A tapered covariance band spends 84% of its own weights at two lags and 74% at thirty. Every charge in the collection is a straight line through the origin in those weights, so it is too dear at one end and too cheap at the other.

curve · Criterion
Where one rule becomes three. Every arrival in 200 simulated trials is put to all three scores, and the picture is how often they would send that patient to different arms. The range and the pairwise sum are the same rule at two arms and at three — for sorted counts the pairwise sum is twice the range, so the arm that minimises one minimises the other — and they part company at four, where the pairwise sum is 3(d − a) + (c − b) and the range still sees only d − a. The variance disagrees with both from two arms onwards, on 5.4% of arrivals at two and 27.0% at five, because the scores are summed over 3 factors and a sum of squares does not order the candidates the way a sum of absolute values does. All three are called minimisation.

Three arms and three scores

Minimisation balances a trial by keeping the arms' counts even inside every prognostic factor. With two arms there is one way to measure how uneven two counts are. With three there are several, they are all called minimisation, and they send different patients to different arms.

multiarm · Assignment
What each rung is made of. Each pair of searches, over 300 draws, split into the two effects its excess is the difference of. The overlap is what the second search loses by having the first already run at its own answer; the interaction is what the joint search finds by moving the first off it. They subtract to the excess exactly, on every draw, because the pinned supremum cancels. Two disjoint dictionaries of independent columns read an excess of 0.000011 and are made of 0.000514 and 0.000503. A break paired with a dictionary of step columns has an interaction of exactly 0 and is all overlap. And a break paired with an independent column has an overlap of -0.004395 against an interaction of 0.002364, which is what puts its excess below zero.

Two effects in one number

How much two searches over one sample share is measured as the net of two things — ground both of them find, and configurations only the joint search reaches. One extra supremum per draw separates them exactly.

separate · Break point
Two searches find some of the same luck. What each search reports on a sample with no break in it, and what the two report together, on four dependences. The dashed line is the sum of the two — what a rule charging each search separately would levy — and the two together always come in below it: 19.30, 15.44, 14.24, 26.64 short, on 100%, 99%, 99%, 100% of draws. The shortfall is not a rounding. Under AR(1) at 0.8 it is 19.30 of the 34.70 the break search manufactures on its own, which is more than half of it. Two searches over one sample are looking at the same noise, and the second one has less left to find.

Two searches, one sample

A searched break in a regression manufactures 34.7 of likelihood ratio where a count of coefficients says 11.1. A searched window manufactures 84.0. The two together manufacture 99.4, not 118.7.

twice · Break point
How much of one search the other has already found. Five pairs of searches on one sample, on a scale whose zero and one are both fixed by construction. Zero is two searches over disjoint sets of independent columns: they remove shares of the residual sum that add, at 0.8 standard errors from exactly additive, and they read 0.004. One is a break search paired with a step column it contains, which reads exactly one on every draw because the step adds nothing at all. Between them: two dictionaries of step columns cut a few rows apart read 0.125, and the pair the earlier field measured — a break and a whitening window, both reading the same residual series — reads 0.762, three quarters of the way to one search containing the other. And below zero, a break paired with a search over independent columns reads -0.306: the joint search finds configurations neither half of it contains, so charging the two separately under-charges.

Two searches that share nothing

Two searches over independent columns remove shares of the residual sum that add exactly. On the scale a chi-square point is quoted on they look super-additive by a fifth of a unit, and none of it is overlap.

apart · Criterion
Coverage of four nominal 95% intervals, n = 20. Computed exactly by summing over all 21 possible counts, not simulated. The Wald interval drops to 18.2% and is jagged everywhere; Clopper–Pearson never falls below 95% and pays for it in width.

What the 95% refers to

An interval that claims 95% is making a checkable statement about a procedure, not about the interval in front of you. Build every possible sample and count, and the interval taught first turns out to cover 87.6% of the time.

intervals · Coverage
One forecast, and the band the arithmetic puts round it. An AR(1) with φ = 0.75, 60 observations, fitted by least squares and forecast 14 steps ahead. The point forecast decays towards the fitted mean at φ̂^h; the band is ±1.96 standard errors from σ̂²Σψ̂², which grows with the horizon and stops at the unconditional spread 1.72. The dashed pair is the same band computed at the true parameters, which nobody has. The marks past zero are what actually arrived: 12 of 14 inside the band this once, which is one draw and settles nothing.

What the model says next

The usual account of a time series stops at estimation. A forecast asks the other question — not what the parameter is but what the next observation will be — and the band round it is a closed form that grows with the horizon and then stops growing, at a value the series was going to reach anyway.

forecast · Forecast
Three quantities, and only one of them crosses zero. Two forecasts of an AR(1) — the last value carried forward and the mean of the last 60 observations — at 1 step ahead. The curve through zero is σ₁² − σ₂², the difference in expected squared error that a comparison of accuracy tests; it changes sign at φ = 0.4922. The two curves above it are σ₁² − σ₁₂ and σ₂² − σ₁₂, the quantities the two encompassing tests are about, and neither of them comes near zero anywhere: the smallest value either takes across the range is 0.008 times the variance of the series. All three are closed forms in φ, R and h with no simulation in them. Equal accuracy is one hypothesis about this picture and encompassing is another, and a set of numbers can satisfy either without the other.

What the other forecast adds

Two forecasters, one series, and two different questions about them. Which is more accurate has an answer that changes with the persistence of the series; whether either is redundant has an answer that never changes at all.

ranking · Forecast
One true null, one table, five readings. every subset of four, fifteen models, at a null where nothing any candidate holds is worth anything, over 500 draws. Each bar is the share of draws on which that reading declares a difference at a nominal 5%. The reading is the whole of the difference between the bars: the data is identical. An open search over all 210 ordered pairs rejects 76.2%; the table's own 5% point is 3.163 against the 1.671 a single comparison uses. Bonferroni takes the open reading to 0.6% — and on the nested ladder the same correction does not reach the nominal level at all, because there the excess is a shift in the mean rather than a maximum over many.

When the benchmark is a candidate

A specification search with a benchmark nailed down is the case with a closed form. Take the nail out — let the model that would have been reported be one of sixteen, chosen by the same data as its rivals — and the same true null is read three ways, at 2.0%, 7.8% and 76.2%.

select · Forecast
One comparison, and the two error bars it can be given. 60 rolling origins, a window of 60 observations, forecasts 4 steps ahead, at the persistence φ = 0.8256 where the two benchmarks have exactly equal population mean squared error. Each mark is one origin's difference in squared error; the horizontal line is their mean, 0.6522. The two vertical bars at the right are ±1.96 standard errors round that mean computed two ways — 0.5337 treating the differences as independent, 0.6880 allowing for the overlap between neighbouring forecasts. The null is true here by construction, so an interval that excludes zero is a mistake, and the narrow one does it far more often than the wide one.

Which forecast is better

Two forecasters, one series, and a difference in mean squared error. Whether that difference is real is a hypothesis test, its terms are not independent, and the standard error it needs is not the one a t-test computes.

evaluation · Forecast
One wrong model, four designs, four slopes. The slope a straight line converges to when the truth is a quadratic, under four covariate distributions, by two routes: the population projection in closed form, and the mean of 2500 fitted slopes at 200 rows apiece. The even spread over [0, 2] gives 1.6000 and the same spread moved to [1, 3] gives 2.6000, while widening it to [0, 4] gives 2.6000 — the same number as the shifted one, because a symmetric design's target is the truth's tangent slope at the design's own mean and does not read the spread at all. An exponential spread with the SAME mean as the first gives 2.6000. So two studies of one world, each fitting the same wrong model, honestly report slopes 1.0000 apart, and neither is making an error.

What a wrong model estimates

A straight line fitted to a curved truth converges on the tangent at its own design's mean. Two honest studies of one world, fitting the same wrong model, report 2.600000 and 1.600000, and neither is in error.

sandwich · Misspecification
The coverage is exact and it is not the nominal rate. ⌈(m+1)(1−α)⌉/(m+1) against m, the number of calibration points, at α = 0.05. It is a closed form and needs no data. It never falls below 95.0% and never reaches 1−α+1/(m+1), the two bounds the rank argument gives. It equals 95.0% exactly at 10 of the 182 sizes drawn — the sizes where (m+1)α is a whole number, which are 20 apart — and sits above it everywhere else, worst at 38 points where it is 97.4359%, or 2.4359% of coverage nobody asked for. Below 19 points there is no such order statistic and the interval is the whole line, which is where the curve starts.

Coverage from exchangeability alone

A conformal interval's coverage is a fact about the ranks of m+1 numbers, so it can be enumerated before any data arrive — all 40,320 orderings of eight values, agreeing with the closed form to machine precision. What that exactness delivers is not 95%.

conformal · Exchangeability
Three mechanisms leave the slope alone; one does not. The bias of the complete-case slope under each of four missingness rules, counted over 4000 studies of 200 rows at 35.0% missing, with the closed form printed beside each count. Missingness that depends on nothing, on the regressor, or on the second covariate leaves the slope exactly where it was — the closed forms are zero to machine precision and the counts are -0.0005, -0.0005 and -0.0011 against standard errors of about 0.0018. Missingness that depends on the outcome moves it by -0.1635, which is 27.3% of the slope being estimated. The same share of rows is lost in every case.

Three mechanisms and one dataset

Four rules for which outcomes go missing, each calibrated to lose the same 35% of the rows and each leaning on what it reads with the same coefficient. Three leave the fitted slope exactly where it was, and the one that reads the outcome moves it by 0.163531.

missing · Missingness
The curvature is in the denominator. The optimism a Bartlett band of each width actually costs, divided by that width, on two ways of measuring the width, over 2000 draws at 120 rows. Measured in the weights the band spends — Σ w(k), which is what the earlier field levies its charges on — the reading falls from 0.9528 at two lags to 0.7486 at thirty, so a charge proportional to the summed weights is too dear at one end and too cheap at the other. Measured in the pairs the band uses — Σ w(k)(1 − k/n), because a lag of k is an average over n − k products — the same readings are flat from 4 lags up, at 0.0084 of χ² per width against 0.2359. The correction has no fitted parameter in it: it is a function of the window, the width and the sample size.

A lag the sample has less of

A sample autocovariance at lag k is an average over n − k products, not n. Count a band's width in the pairs it actually has and the curvature in its charge goes away, on a correction with nothing fitted in it.

curve · Criterion
Two instruments, two block lengths. The block length that would actually have been best on each draw, for each of the two error readings, averaged over 400 samples of 120 rows. For the rectangular window the implied long-run variance wants 18.92 and the 95% point wants 16.05; for the tapered window, 21.82 against 17.74. The quantile wants a shorter block under both windows — a ratio of 0.848 and 0.813. That is the mechanism the whole field turns on: a rule for choosing a block length is a way of guessing a target, and the two instruments do not have the same target. A rule tuned to one is systematically long for the other, and the two windows do not pay the same price for being long.

A length for each instrument

The block length that is best for an implied variance is 18.92; the one best for the 95% point of the same resamples is 16.05. A rule is a way of guessing a target, and there are two targets.

readout · Bootstrap
Which window is better depends on who chose the block length. The margin between a rectangular block and a tapered one, on 400 samples of 120 rows, under four rules for choosing the block length. At the length that would actually have been best on each draw the taper is ahead by 2.12 points of a 35.8% error, at 22.1 paired standard errors; at a length estimated from the sample's own persistence it is ahead by 1.89. At the length this field's own figures use — eight — the rectangle is ahead by 2.04, and at the rule of thumb by 4.54. Every rule sees the same draws. What separates them is the length: the two rules that lose to the rectangle pick 4.00 and 8.00 where the best available is 24.57, and a tapered window at a quarter of the right length has thrown away most of what it was weighting.

An ordering that depends on the rule

The tapered block beats the rectangular one at the best available block length and at one estimated from the data. At a length written into a protocol, and at the rule of thumb, the rectangle wins — at every sample size measured.

feasible · Bootstrap
A trial designed 2:1:1, and what two scores deliver. 500 trials of 180 patients, three arms, a target of 2:1:1. The shaded bars are a minimisation score that divides each arm's count by the share that arm is supposed to receive before measuring the spread; it delivers 49.9% : 25.1% : 25.1%. The others are the same rule with the counts left raw, which delivers 33.4% : 33.3% : 33.3% — the balance it enforces inside every factor level is equality, and equality is what it gets. The marks are the shares that were asked for.

Balancing towards unequal targets

A three-arm trial allocating two to one to one is the ordinary case, and a balancing rule built from raw counts does not know it. It balances the arms towards equality inside every factor level, delivers a third to each arm, and reports that it minimised imbalance.

multiarm · Allocation
Eight candidates, one of them exactly as good as the benchmark. The candidate set: moving averages of the last 1, 2, 3, 5, 8, 13, 21 and 34 observations, each drawn as its expected squared error divided by the benchmark's — the mean of all 60. The persistence is not chosen, it is solved for: at φ = 0.4895 the best candidate in the set, the average of 2, has exactly the benchmark's expected squared error, and every other candidate is worse by between 0.5% and 5.1%. So the null that no candidate beats the benchmark is true, with one candidate on its boundary. Everything a set comparison claims about its own error rate has to be measured here, because anywhere further inside the null every procedure flatters itself.

Eight forecasters and one benchmark

A set of forecasters is a multiplicity problem on top of a dependence problem, and the two do not separate. Eight windows of one series carry the multiplicity of two and a half independent comparisons; eight separate problems carry eight.

ranking · Multiplicity
Stationary is not the same as convergent. How far each k-swap walk is from uniform after t steps, started at the least balanced admissible assignment of 410. Every one of these chains has a symmetric proposal and rejects by standing still, so every one of them is doubly stochastic and every one preserves the uniform distribution exactly. Only five of the six get there. Exchanging all six units of each arm is a single proposal — the complement — and the admissible set is closed under complement, so the walk takes it every time and oscillates between two assignments for ever: after 160 steps it has visited 1 state and sits 0.9976 from uniform. Its stationary distribution is a fact about the matrix; its limit does not exist.

Stationary is not convergent

A walk that exchanges every unit in each arm preserves the uniform distribution exactly and never gets near it. Every doubly stochastic matrix has the same stationary distribution; only some of them have a limit.

blocks · Randomisation
Sheppard's arcsine, by two routes. Corr(sign X, sign Y) as the covariates' correlation runs from zero to one, drawn twice. One route is a sixty-four-node quadrature of the orthant probability over the correlation — the general construction, which works at any pair of cut points; the other is (2/π) arcsin ρ, which is elementary and works only at the median. They agree to 3.3e-16 at every one of 81 correlations, which is what licenses the quadrature everywhere else. The curve is above the diagonal at small ρ and below it at large: two signs agree with probability ½ + arcsin(ρ)/π, so a correlation of 0.5 gives exactly ⅓ and a correlation of 0.8 gives 0.5903.

The arcsine that closes it, and the error that was overstated

Two median splits of a correlated pair agree with probability ½ + arcsin(ρ)/π, exactly. And the truncation the field was avoiding falls geometrically in the correlation, not algebraically in the order.

splits · Routes
Three charges, and only one of them is a test. What each of three thresholds does to the same decision, under AR(1) at 0.8, against the size of a genuine break in the mean at row 60. A chi-square on the 5 coefficients a split adds — 11.07 — declares a break on 73.6% of samples that have none: it is not a test at all. The break search's own 95% point, 69.6, carried into a rule that also chooses its window, fires on 0.0% of null samples and on 0.0% of samples with the largest break measured — the natural way of combining two published corrections does not lose a little power, it switches the test off. The calibrated charge, 27.2, holds 5.6% at no break and reaches 29.2% at the largest.

The charge that is not a sum

Charging two searches what each costs on its own is conservative, and conservative here means the test never fires. At the largest break measured it declares nothing, on every draw, while a calibrated threshold reaches 29%.

twice · Break point
Every candidate is behind by what its parameter count says. Each dot is one of the fifteen subsets of four predictors, fitted on a rolling window of 80 rows and scored against the benchmark out of sample over 60 origins, at a null where every one of them contains the truth. The line is σ²(q₀/(R − q₀ − 1) − q/(R − q − 1)), which is arithmetic on two integers and a window length. Most of these pairs are not nested — a subset of two predictors and a different subset of two share neither model — and the closed form does not care: the displacement is a statement about how many coefficients each side estimates. The candidates of the benchmark's own dimension sit at zero.

The displacement is a parameter count

A nested variant is behind its benchmark out of sample before anything is searched for. The closed form for how far turns out to have nothing about nesting in it — only two integers and a window length — and it prices a table where no candidate contains any other.

select · Multiplicity
What a 95% forecast interval covers, counted. 1200 series of 25 observations from an AR(1) with φ = 0.7, at each horizon, on one set of seeds. The upper line is the interval computed at the true parameters — it covers 95.3% on average, which is the check that σ²Σψ² is the right formula rather than a claim about anything a forecaster can do. The lower line is the same formula fed σ̂² and φ̂: 92.8% at one step and 87.3% at 6. The interval that would cover what it claims is 6.9% wider at one step.

The interval that forgets it estimated

The forecast band is derived for a model whose parameters are known, and then computed by putting estimates into it. Counted, the 95% interval covers 87.3% six steps ahead on twenty-five observations, and the point forecast inside it returns to the mean a third faster than the series does.

forecast · Forecast
Most of the rise is the optimiser's, and under one law it is not. The rise in log-likelihood from the tapered plug-in to the maximum over the same eight-lag band, beside what the same optimiser produces on a sample generated from the plug-in's own covariance — where the family is correctly specified by construction and there is nothing to find. Under AR(1) at 0.8 the raw rise is 5.72 and the manufactured baseline is 4.79, leaving 0.93 at 1.8 standard errors; under long memory the excess is 0.14, at 0.2. Under the moving average it is 11.87 at 19.4 standard errors, on every draw. The taper is a shrinkage, and it costs nothing where the sequence decays smoothly and a great deal where it stops dead.

The plug-in and the maximum

A tapered covariance estimate sits five and a half log-likelihood units below the maximum of the likelihood it is substituted into. Four fifths of that is what the optimiser would have found if nothing were missing.

family · Dependence
Two rates, not a factor. The standard deviation of the covariate imbalance under three rules, at five trial sizes, 260 trials each, on log axes. The upper line is a coin: its slope is -0.489, against a closed form of exactly −½. The middle line is minimisation on a median split; its slope is -0.519 — the same rate — because inside a category the assignment is still a coin, and what it buys is the constant, 0.654 of a coin's at n = 200. The lower line is the rule that reads x and maximises the information about the treatment effect: slope -0.987, nearly twice as steep. Its advantage is therefore not a number that can be quoted — it is 0.258 of a coin's at n = 50 and 0.065 at n = 800, and it keeps going.

The rule that reads the number

Stop categorising and let the rule read the covariate itself. What it should minimise is not an invented distance but the variance of the effect being estimated — and what comes back is not a better constant but a different rate.

continuous · Assignment
Where the residual test's statistic actually falls, at n = 200. Four thousand pairs of unrelated random walks, each regressed on the other and each residual tested for a unit root. The statistic is computed as a t and its distribution is not a t: five per cent of it falls below -3.38, where the ordinary one-sided 5% point of a t on 198 degrees of freedom is -1.65. Everything left of -1.65 — 70.2% of the whole distribution — is a pair of unrelated walks that a t table calls cointegrated.

The test with no table

The statistic that separates a real long-run relation from a spurious one is computed as a t and is not a t. At two hundred observations its 5% point is −3.38 where the t table says −1.65, and reading it against the table calls two unrelated random walks cointegrated 70.5% of the time.

cointegration · Spurious
No block size is best at both things the procedure claims. Two claims and one dial. The honest interval's half-width falls as the blocks get smaller, because the interval's degrees of freedom are the number of blocks: 0.2602 at blocks of two against 0.2933 at blocks of sixteen. The fixed-width claim — that the mean is within 0.25 of the truth — gets more reliable as they get larger, because the sample size is less variable: 92.40% against 95.00%. Both are computed from the same runs, and the second is reproduced to within a tenth of a point by E[2Φ(d√N/σ) − 1], which needs the sample-size distribution and nothing else. The schedules sit at the bottom left: as narrow as the smallest fixed block and as few observations, with the rule's spread estimate on half as many degrees of freedom again.

Two degrees of freedom, one total

The block size is a dial, and the two things a fixed-width procedure claims move in opposite directions along it. Divide the width by the square root of the sample size and one of them turns out to depend on the number of blocks and on nothing else.

pace · Width
Two factors, opposite directions. The two factors the cost of a per-candidate tuning parameter is a product of, as the candidates are pulled apart, over 800 draws at each of 5 separations. How often the candidates disagree about the tuning parameter rises from 31.8% to 88.8%; the share of those disagreements that change which candidate the table selects falls from 51.6% to 1.7%. So the setting where the candidates quarrel most about the tuning parameter is the setting where the quarrel matters least, and a sweep that reads the rate and stops has read the factor pointing the wrong way.

Two factors pointing opposite ways

As the candidates on a table are pulled apart, they quarrel about the tuning parameter three times as often and the quarrel decides the winner thirty times less often. A sweep that reads the first factor has read the one pointing the wrong way.

turnover · Order-selection
Coverage of four nominal 95% intervals, n = 30. Computed exactly by summing over all 31 possible counts, not simulated. The Wald interval drops to 26.0% and is jagged everywhere; Clopper–Pearson never falls below 95% and pays for it in width.

Two routes to every number

A site about probability that only simulates has one route to each answer and no way to tell a right one from a plausible one. Every important number here is computed twice, by arithmetic that shares nothing, and the two are required to agree.

method · Routes
Twenty walks up the same hill, σ = 2. Each walk fits a plane to the same four-corner factorial, takes its gradient as a direction, and steps along it until a run comes in below the one before. The true optimum is the cross. 80% of the walks stop before the best point on their own path — not because the direction was wrong, but because one noisy run is enough to stop them, and the direction error costs only 3.9% of the available gain.

Walking up the gradient

The fitted gradient is wrong by an angle with a closed form, σ/(|β|√N), and what that angle costs is its squared cosine — twelve per cent at twenty degrees. What costs a third of the gain is not the direction at all. It is deciding where to stop.

surface · Optimum
Both halves grow; the difference does not. The control pair's two components and their difference, against how much each of its two searches can find, over 1200 draws at each dictionary size. Two disjoint sets of independent columns are additive at every size — the excess stays inside a standard error or two of zero throughout — and it is not because there is nothing there. The overlap grows from 0.000112 at two columns to 0.000870 at ten, a factor of 7.76, and the interaction grows with it, staying within a factor of two of the overlap at every size. Two searches competing for one residual sum share ground and find configurations neither has alone, in almost equal measure, and their difference is what the earlier field's scale calls zero.

What a zero is made of

Two disjoint dictionaries of independent columns read an excess of 0.000116 and are made of an overlap of 0.000583 and an interaction of 0.000467. The control the whole scale is anchored on reads zero because two effects cancel.

separate · Break point
What studentising costs. How much wider the studentised interval is than the percentile one, cell by cell, over 300 draws, with what each cell gains in coverage beside it. Averaged over the eight cells the interval is 2.09 times as wide and covers 0.46 points better. At the two rules that choose short blocks the two intervals are within a fifth of each other; at the oracle's length, where a resample holds two or three whole blocks, the studentised interval is 4.37 and 5.04 times as wide. A repair that doubles the width and buys half a point is not one a reader could not have had by widening the interval it replaced.

What studentising costs

Averaged over eight cells the studentised interval is 2.09 times as wide as the percentile one and covers 0.46 points better. At the block lengths the rules choose, the scale it divides by rests on two or three numbers.

student · Bootstrap
A test between nested models, under a null that is true. 1000 comparisons: an AR(1) truth, forecast by a fitted AR(1) and by a fitted AR4 whose extra coefficients are zero. In population the two forecasts are identical, so every rejection is false. The larger model's mean squared error is 1.1663 against 1.0583 — worse, by exactly the noise in estimating coefficients that are not there — and the ordinary test therefore declares the smaller model significantly better 67.2% of the time. Read one-sided in the direction anybody asks about, it finds the larger model better 0.0% of the time. Adding the squared difference between the two forecasts back into the loss differential puts the level at 4.9%.

When one model contains the other

The comparison a forecaster most often wants is between a model and the same model with one more term. That is exactly the comparison the standard test cannot make — and it fails by declaring the smaller model significantly better, more confidently the more data it is given.

evaluation · Forecast
The crossing is in the dependence, not in the split. Regret of each rule as the design and the errors are made persistent at the same coefficient, scored on fresh rows because the closed form assumes exactly what is being taken away. An optimism theorem counts rows; when the rows repeat each other there are fewer of them than there are rows, the penalty is too small for the fit it is correcting, and the criterion starts buying coefficients it should not — its average winner grows from 3.31 coefficients to 3.90. The hold-out never used the theorem and overtakes at ρ ≈ 0.81. Schwarz's criterion, worst of the three on independent rows, is best on repeating ones — its heavier penalty is right for the wrong reason.

Where the two searches cross

The obvious dial between a criterion and a hold-out is how much of the sample to hold out, and moving it never changes the answer. The dial that does is one nobody chooses — how much each row repeats the one before it — and the two rules change places at about 0.81.

proxy · Forecast
A weak instrument gives back the problem it was hired for. The counted mean bias of two-stage least squares at 4 instruments and 200 rows, over 2000 draws a setting, against the standard approximation and against the least-squares inconsistency the instrument was brought in to remove. At π = 0.02 the counted bias is 0.3220 ± 0.0142 where least squares is out by 0.3594 — 89.6% of the way back. At π = 0.3 it is 0.0118 against 0.2647. The approximation, the inconsistency over the population first-stage F, tracks the count at the weak end and sits above it in the middle: 0.968, 0.971, 0.918, 0.810, 0.740, 0.722, 0.846 as the ratio of counted to approximated bias.

Weak, and back where it started

A consistent instrumental estimate at two hundred rows and a concentration parameter of 0.32 is biased by 0.3220 ± 0.0142 against a least-squares inconsistency of 0.3594 — 89.6% of the way back to the problem it was hired to solve. Just identified, it has no mean at all, and that is measured as a rate rather than assumed.

instrument · Exclusion
What each error is a claim about, and what the claim comes out as. Each variance estimate's average over 20000 draws, divided by the variance the slope actually has across those same draws, at 80 rows with the error variance leaning towards the edges of the design (γ = 0.8). One is a standard error that is right. The model-based estimate reads 0.6081 of the spread, so its standard error is 77.98% of the one it should report; the four robust corrections read 0.9576, 0.9821, 0.9961, 1.0362. Two further routes agree with the count and share none of its arithmetic: n times the counted variance is 4.8905 against a population sandwich of 4.9200, and the counted ratio of the two standard errors is 1.2799 against a closed form of 1.2806.

The bread and the filling

The robust standard error is not a safety margin. At one setting of the error variance it is 1.2806 times the model-based one and at another it is 0.8246 times it, and the sign of a single dial decides which.

sandwich · Misspecification
What a search costs is not a property of that search. The likelihood ratio a searched break in the regression reports, two ways, on every law. On its own — the whole rule being a split of the sample, at no whitening — it averages 34.70 under AR(1) at 0.8, against the 11.07 a chi-square on the five coefficients a split adds would use as a threshold. Inside a rule that also chooses a window from a list of eight, the same search adds only 15.40 — less than half. Most of what a break search finds under correlated errors is the correlation, and a whitening chosen from the same sample has taken it already. A charge measured for one search, carried into a rule that makes two, is not conservative in some harmless direction: it is measuring a different quantity.

A charge that depends on the rule

The break search's charge is 34.7 on its own and 15.4 once a window has been chosen from the same sample. Most of what a break search finds under correlated errors is the correlation, and a whitening has taken it already.

twice · Break point
A rate that does not know how large the trial is. The share of equal splits admitted by a tolerance of 1 coin-spreads on 3 functions, at six trial sizes. The first two are exact — 12,870 and 184,756 splits, walked, averaged over eight draws of the units — and the rest are sampled. From a hundred units on, the rate sits on (2Φ(1) − 1)^3 = 0.3182, which contains no n at all. The two small trials are 29.2% and 27.7% short of it, so the sixteen-unit measurement understates the rate rather than bracketing it. Meanwhile the admissible count — the rate times C(n, n/2) — goes from 2^11.5 to 2^393.7: the exhaustion a small trial runs into is a fact about small trials.

A count that has to be estimated

At sixteen units the admissible assignments can be counted by walking all 12,870 of them. At four hundred there are about 2^393.70, and the share admitted is 0.31885 against a closed form of 0.31818 that has no trial size in it at all. The exhaustion a small trial runs into is a fact about small trials.

product · Randomisation
A line in the right width beats two curves. How far each candidate charge sits from the measured optimism across the plateau, in units of each width's own standard error, over 2000 draws. The straight line through the origin in the band's summed weights — which is what the earlier field levies — misses by 0.2382 per width. The same straight line in the pairs the band actually uses, Σ w(k)(1 − k/n), misses by 0.0095. A fitted power law misses by 0.0293 and a fitted decaying rate by 0.0172, both on one fitted constant more. The deferral this field answers asked for a curve; the answer is a line, in a variable with nothing fitted in it.

A line that beats two curves

A deferral asked for a curve. Fitted against the same measurements, a straight line in a variable nobody had to fit describes the plateau better than either curve does with a constant more — and for three windows out of four it does not.

curve · Criterion
What a longer list actually changes. How often the five candidates choose different tuning parameters, at a true null where every one of them contains the truth, so a disagreement is manufactured rather than discovered. The order's list is an interval of integers, and thinning it moves the rate smoothly from 0% at two values to 39% at thirteen. The window's is not an interval — it runs 0, 1, 2, 4, 8, 12, 20, 30 — so a thinned window list jumps depending on whether it happens to keep the width the criterion wants, between 0% and 42% with no order to it. So "the same length" was never quite the same thing for the two rules, and it is a smaller effect than the field it was invoked to explain.

A list is not a rule

How often five candidates disagree about a tuning parameter runs from nothing at two values on the list to two draws in five at thirteen. What the disagreement costs does not move at all.

lists · Order-selection
The split depends on the order. How much two searches share, measured both ways round, over 300 draws. Pinning the first search at its own answer and searching the second gives one overlap; pinning the second and searching the first gives another. A break paired with an independent column reads -0.004395 one way and 0.002163 the other, at 8.70 paired standard errors and on opposite sides of zero. The excess the two components subtract to is the same in both orders by construction, so what changes is only how it is attributed. There is no order-free way to say which of two searches found ground both can reach, and the two orders bracket it.

A split that depends on the order

Run the second search first and pin that instead, and the same draw gives a different overlap and a different interaction — with the same difference. And one pair has no second order at all.

separate · Break point
Flat along a row, apart between them. The probability that a per-candidate tuning list changes the winner, at three list lengths on three candidate tables, over 800 draws in each of the nine cells. Along a row — the reading the earlier field takes — it moves by a factor of at most 1.21, so that field's invariant survives on every table. Down a column it moves by up to 1.98. The list length is the dial that does not move this number and the table is one that does, and the earlier field varied only the first.

A table and a list

A nested ladder of candidates differing by one coefficient was predicted to turn over more often at every list length. It turns over less at every one, and its list changes the winner half as often.

turnover · Order-selection
What each criterion selects, at 50 observations. 700 series from an AR(2) with coefficients 0.6 and -0.3, every order from 0 to 8 fitted to the same 42 responses so the log-likelihoods are comparable. AIC finds the true order 55.1% of the time and lands above it 25.7%; BIC finds it 54.1% and lands above it 4.0%. The closed form for one extra lag is P(χ²₁ > 2) = 15.73% for AIC, which does not depend on n at all, and P(χ²₁ > ln n) = 4.79% for BIC at this size, which falls to zero. Under the true order is the other failure and it is BIC's: 41.9% against 19.1%.

Choosing the order

One criterion is consistent and one is not, which is the whole of what gets said about them. At two hundred observations the consistent one is right 95% of the time and the other 70%; at fifty they are both right 54% of the time and wrong in opposite directions, and consistency has not started to mean anything yet.

forecast · Order-selection
Least squares estimates persistence low, by an amount with a formula. 3000 series of 50 observations at each persistence. The lower curve is the counted bias of the least-squares estimate of φ, and the open marks on it are −(1 + 3φ)/n, computed rather than fitted. The upper curve is the bias left after adding that quantity back, evaluated at the estimate rather than at the truth nobody has: -0.0020 at φ = 0.3, -0.0023 at φ = 0.5, -0.0039 at φ = 0.7, -0.0059 at φ = 0.8, -0.0108 at φ = 0.9, -0.0165 at φ = 0.95. The formula is a leading-order expression and it understates the bias where the persistence is nearest one — -0.0882 counted against -0.0770 predicted at φ = 0.95, which is the corner of the parameter space every one of these approximations is worst in.

Correcting the persistence

Least squares estimates how much a series remembers of itself as smaller than it is, at every value it can take, by an amount with a closed form. Subtracting that amount back is one line of arithmetic, and what the line costs is variance.

evaluation · Bias
The quantity that does not depend on the list. The probability that letting each candidate choose its own tuning parameter changes which candidate the table selects — the product of the two moving shares — against the length of the list, over 1200 draws apiece. It is 14.2%, 11.9%, 12.3%: a spread of 2.2% across a list length that moves the disagreement rate by a factor of 1.52. This is the invariant the whole field turns on. Everything downstream of the winner — the coefficients, the regret, whatever a reader is going to quote — is a function of whether the winner changed, and how often that happens is not something the list controls. A longer list changes how often the candidates quarrel and not how often the quarrel matters.

How often it matters

The disagreement rate rises by half across the list and the share of disagreements that decide anything falls by nearly the same factor. Their product — how often the tuning list changes which candidate wins — sits at an eighth and does not move.

apiece · Order-selection
What each instrument costs to read. The number of draws each instrument needs to separate a rectangular block from a trapezoidal one at two standard errors, at a block length of 20 and 120 rows — measured from each instrument's own spread on the same draws. The implied variance needs 7.0 and the 95% point needs 20.2, a factor of 2.90 at this block length. There is a closed form beside it and it does not depend on either the scale or the size of the gap: the standard error of a p-quantile is √(p(1−p))/f(q) over √B where a standard deviation's is σ/√(2B), which at the 95% point of a nearly normal reference distribution is 3.30 times as many draws for the same statement. And the quantile route needs every one of those draws resampled, where the variance route needs none.

Measuring a variance rather than a quantile

A resample's implied long-run variance can be computed from the sample with no resampling in it at all. A critical value cannot, and the difference is a factor of three in the draws before any of the resampling is counted.

crossing · Bootstrap
Which repair goes with which defect. The share of true nulls rejected at a nominal 5% by a reference distribution generated from the fitted benchmark, over 120 draws with 59 resamples each. Where the errors are well behaved every resampling is fine and all four are conservative. Where the variance is a function of the design, the two that detach a residual from its own row reject 5.8% and 8.3% — and the block bootstrap, which is the resampling three earlier fields on this site reach for, repairs nothing at all, because the dependence it is built for is between origins and the rolling scheme reproduces that on its own. Where the errors are skewed the symmetric multiplier is the one that is wrong, and Mammen's two-point version is the only one of the four that is right in both columns.

Residuals that keep their own variance

A reference distribution for a search has to be generated from a fitted model, and the generator draws residuals. Four ways of drawing them keep four different things — and the one this site has reached for three times repairs nothing at all here.

select · Bootstrap
One rule keeps its promise and the other keeps its budget. Both stopping rules at five requirements, 1,500 experiments each, with a first stage of 5. The upper curve is the two-stage rule: 97.1%, 96.2%, 95.9%, 96.1%, 96.0% — at or above 95% at every point, which is a theorem rather than a tendency, because its interval is built from a spread estimated before the stopping point was chosen. It pays 2.06×, 2.01×, 2.00×, 1.99×, 1.99× the observations that knowing σ would need. The lower curve is the rule that re-estimates after every observation: 94.5%, 89.3%, 91.0%, 91.5%, 94.3%, on 0.98×, 0.87×, 0.88×, 0.93×, 0.96×. The second rule is the one anybody would run and the first is the one whose claim is true.

Stopping when it is precise enough

An experiment that runs until its estimate is precise enough is the natural design and the one with a theorem against it. Its two-stage cousin keeps its promise exactly, for every unknown spread, and pays twice the observations for it.

guarantee · Stopping
Four analyses of the same 3-arm trials, under a true null. 250 trials of 150 patients, 3 arms, minimisation with p = 0.85, 99 re-randomisations for each exact test. Two statistics — an F on the arms alone and an F on the arms after the balanced factors — against two reference distributions: the table the statistic is named for, and the distribution the allocation rule itself generates when the outcomes are held fixed and the rule is re-run. Only the first cell is wrong, and it is wrong in the direction that costs power rather than the one that manufactures findings: 0.0% where 5% is claimed. Either repair works — adjusting for what the rule balanced, or asking the rule what it would have done.

The analysis after three arms

An unadjusted analysis after a two-arm balancing rule rejects 0.6% of true nulls where it claims 5%. With three arms and a deterministic rule it rejects none at all — and the repair is the same repair, which is a sentence and a column in the model.

multiarm · Assignment
One experiment finding out where to look. A single run of the fully sequential design: 40 runs, the first 8 placed at the guess K = 1, then the model refitted and the design revised after every 2. The marks are the settings the runs were made at. The horizontal lines are where a design built at the truth K = 3 would have put them — 1.875 and 10.00 — and the rule walks onto them without being told: its estimate of K after the first eight runs was 2.694, and by the end 2.765 against a truth of 3. The whole experiment is 96.5% as efficient as the design that knew the answer, where running all 40 at the guess would have been 81.1%.

The design that stops guessing

Every repair so far protects a guess. The alternative is to run part of the experiment, estimate the parameter from it, and design the rest at the estimate — which recovers most of what a threefold wrong guess costs, and has a best moment to stop guessing that is earlier than anyone expects.

robust · Local design
Sixteen candidates nobody would have run, and what they cost. At φ = 0.65 one candidate in the original set is genuinely better than the benchmark, and the question is how often each procedure finds it. The added candidates are stale copies of the last value — read two, four, six … steps late — every one of them worse than the benchmark by at least 43%, and not one of them is ever the best candidate in a sample. The reality check goes from 35.5% to 0.0% as they are added, because its reference distribution has to assume every candidate is exactly as good as the benchmark and sixteen such assumptions is a critical value nothing reaches. The recentred version, which drops from the recentring the candidates the data has already ruled out — 15.1 of 24 of them — goes from 25.5% to 24.5%. The top line never moves: reporting the winner's own p-value cannot notice a change to a set it never looks at.

The models that were never in the running

A reference distribution for a set has to assume something about every candidate in it. Assuming that all of them are as good as the benchmark is what makes the reality check honest, and it is what sixteen hopeless candidates use to destroy it.

ranking · Multiplicity
Four rules of four change sign. The margin between the two block windows in points of coverage, under each of four rules, on each of three intervals built from the same resamples, over 300 draws. Positive is the tapered window covering better. On the percentile interval the taper wins at all four rules, by 5.33, 1.67, 5.00 and 4.00 points, which is the earlier field's own reading. On the studentised interval the rectangle wins at all four, by 4.33, 7.00, 4.33 and 2.67. And a normal interval, which uses no resampling at all, puts the two within a third of a point at every rule — so the disagreement is manufactured entirely by what is done with the resamples.

The ordering reverses again

One field found two of four rules changing sign between two readings of one resampling. Turn the same resamples into a studentised interval instead of a percentile one and all four change sign.

student · Bootstrap
Where the +1 matters, and why nobody has noticed that it does. The true size of the two rules at every B, computed rather than simulated: under the null the count of re-randomisations reaching the observed statistic is uniform over {0 … B}, so both sizes are integer arithmetic. With the +1 the size is (⌊α(B+1)⌋)/(B+1), which never exceeds 5%. Without it the size is (⌊αB⌋+1)/(B+1), which is larger except at B = 19, 39, 59 — the values with B + 1 a multiple of 1/α, and the values everybody uses. At B = 19 the two rules are the same rule; at B = 20 the uncorrected one is a 9.5% test. The marks are simulated on 500 trials of 120 patients, as the second route to the same numbers.

The plus one and the round number

A sampled randomisation test counts the observed allocation as one of its own reference draws, and the correction is invisible at B = 19, 39, 59 and 999 — every value anybody uses. At B = 20 the version without it is an 8.00% test where the corrected one is 3.80%, and the convention protecting everybody is a preference for round numbers minus one.

exact · Reference
The reversal is a property of the instrument. The margin between the two block windows under each of four rules, on three readings of the same resampled means, signed so that a positive bar is the tapered window winning. On the implied long-run variance the taper wins at the best available block length and at one estimated from the sample and loses at a length written into a protocol and at the rule of thumb — which is the reversal the earlier field's whole argument turns on, at 1.48 and 4.52 points. On the 95% point a test actually reads, the taper wins at all four, by 6.13 to 7.08 points. On the coverage the interval actually delivers, the taper wins at all four again, by 2.50 to 5.75 percentage points. Two of the four rules change sign between the first reading and the other two, and the two that change are exactly the two the earlier field's recommendation is about.

The reversal that was the instrument's

On an implied variance the rectangle wins at a protocol length and at the rule of thumb. On the 95% point a test reads, and on the coverage an interval delivers, the taper wins at all four rules.

readout · Bootstrap
Each repair is for its own defect, and one is for both. The 95% point of the statistic's own distribution in each world, against the mean 95% point of five reference distributions built from one sample. Where the error variance is a function of the design, the two resamplings that detach a residual from its row fall short and the two multipliers that keep it there do not; where the rows repeat each other it is the other way round. With both defects at once the blocked multiplier — drawn once per run of 5 rows, so the residual never moves and its neighbours share a sign — is the closest of the five, at 2.999 against a truth of 3.803. It is still short by 0.804, and that shortfall is the next figure.

Two defects and one resampling

Four resamplings, each the repair for one defect and wrong about the other. Put both defects in the same world and the statistic's 5% point is 3.8028, where the best of the four reaches 2.8326 — until a multiplier that stays on its own row and shares a sign with its neighbours reaches 2.9988.

proxy · Bootstrap
How long a walk has to be given. Every equal split of twelve units is enumerated, the 410 admissible ones are found, the transition matrix is built, and the distance from uniform is computed exactly at each step — no simulation anywhere. The walk is started at the least balanced admissible assignment, which is the state a rejection sampler is least likely to have handed it and the one a burn-in has to cover. It is 0.0849 away after twenty steps and 0.00008 after ninety. A real cost, and a small one, and naming it is what stops it being assumed to be zero.

Walking the admissible set

A rerandomisation test hunts for admissible assignments and throws away the rest. A walk visits them instead — and it is exactly uniform only because it stands still when a proposal fails, which is the step that looks like waste.

joint · Randomisation
The overshoot is the last block size and nothing else. A run stops at a multiple of its own block sizes and cannot land between them, so it ends past its own target by about half a block. Fixed sizes overshoot by 1.5, 2.2, 3.1, 4.7, 8.5 observations as the size goes 2, 3, 5, 8, 16. Every schedule here ends in blocks of two and every one of them lands where blocks of two land — 1.62, 1.32, 1.37 against 1.48 — while having spent most of the run inside blocks four and eight times larger. That is the one thing on this page a schedule genuinely takes from both ends.

What a schedule actually buys

Big blocks early and small blocks late is the right instinct and it does not take both ends of the trade, because there are not two ends to take. What it does take is the overshoot — about four per cent of the observations — and a steadier stopping point.

pace · Nuisance
The argument is a third of the size of the thing it is inside. Three quantities on one scale, in points of the error in a block resample's implied long-run variance, at 120 rows. The gap between the two windows at the best available block length — the whole subject of the comparison this field inherited — is 2.12 points. What the best rule a practitioner could actually run gives up against that same best length is 7.26, a factor of 3.42. What the rule of thumb gives up is 26.01. So the ordering between windows is worth establishing and is not worth arguing about, and the sentence that follows from it is not use the taper but estimate the block length, because that is where the points are.

What choosing the length costs

The gap between two block windows at the best available length is 2.12 points. What the best rule a practitioner could run gives up against that same length is 7.26. The argument is a third of the size of the thing it is inside.

feasible · Bootstrap
Three analyses of the same trials, none of them wrong about the data. 320 trials at n = 60 with no treatment effect at all, so every rejection counted is a false one, and a covariate that drives the outcome with coefficient 1. The unadjusted comparison is at 5.94% after a coin — its level — and at 0.00% after the rule that reads the covariate: the design removed the imbalance and the analysis is still pricing it. Adjusting for the covariate gives 4.06%, and the rule's own reference distribution — hold the outcomes, re-run the rule 199 times, count — gives 3.13% against the 4.5% that 199 draws can deliver. The last of the three has to be told the assignment rule and nothing else, which is the one thing the experimenter certainly knows.

What the balanced trial is worth

A rule that reads the covariate removes three quarters of the imbalance. An analysis that does not know it happened prices the imbalance anyway, rejects one true null in two hundred instead of one in twenty, and finds a real effect less often than a coin-tossed trial does.

continuous · Randomisation
The crossing barely moves. Both methods' costs in one unit — assignments evaluated per usable draw — as the tolerance tightens. A hunt costs 1/p and rises without limit: from 2.22 at a tolerance of 1.2 to 357.14 at 0.18. A walk costs its autocorrelation time and barely moves. The two cross at a tolerance of 0.190 at one swap and 0.195 at eight — the whole family of proposal sizes crosses inside a band of about two hundredths, because where the crossing is, the large proposal has already lost its advantage. A multi-swap proposal is worth a factor of 5.65 in the regime where the walk should not be used at all.

Where the gain is, and where the decision is

A bigger proposal is worth a factor of six at a loose tolerance and nothing at a tight one. The tolerances where it helps are the ones where a hunt costs two evaluations a draw, and the crossing barely moves.

blocks · Assignment
Two structures in three are made worse. What adjusting for every covariate measured does to the bias in the treatment's estimated effect, against adjusting for none, over 4000 randomly drawn structures of 6 covariates each. Each covariate is independently a common cause with probability 0.25, a cause of the treatment only, a cause of the outcome only, a cause of neither, a step on the causal path, or a common effect. The rule leaves a larger bias on 65.5% of structures, a smaller one on 33.8%, and the same on 0.7%. The share is a property of that population of structures rather than of adjustment, which is why the weights are stated; what does not depend on them is that the rule has no direction — it is not a conservative default that occasionally overcorrects, it is a rule whose error is whatever the structure happens to be.

Adjusting for everything

"Control for every covariate that was measured" leaves a larger bias than controlling for nothing on 65.5% of four thousand randomly drawn structures and a smaller one on 33.8%. Its squared error is 4.110 times that of using no covariate at all, and half of it sits in its worst tenth of structures.

collider · Conditioning
Robust, at the sample sizes it is reached for. Counted coverage of five 95% intervals for a slope, at six sample sizes, under an error variance leaning towards the edges of the design (γ = 0.8), over 20000 draws at the small end. The model-based interval sits at about 87.06% everywhere and does not improve with the sample, because it is a claim about a variance it is not estimating. The robust ones do improve: HC0 covers 88.73% at 20 rows, 92.70% at 50 and 94.93% at 1,000. Its promise is asymptotic and its use is not, and the gap between those two facts is this picture. The leave-one-out correction read against a t on n − 2 is the only line that is near its promise at the small end: 94.55% at 20 rows.

Robust is not free

A robust standard error's promise is asymptotic and its use is not. Its 95% interval covers 88.73% at twenty rows, and under mild heteroskedasticity it is the worse of the two intervals until a hundred.

sandwich · Misspecification
One cohort of 40, and two intervals around the end of its curve. A single simulated study of 40 subjects with exponential survival at rate 0.35, dropout at rate 0.15 and follow-up to 6 — the first seed from 8811 upward whose plain band reaches below −0.05, chosen to show the failure rather than its frequency. The step curve is Kaplan–Meier and the smooth curve the truth. The plain band, the estimate plus and minus 1.96 Greenwood standard errors, first dips below zero at t = 3.78 and reaches −0.052; early on it also rises to 1.023, above one. At t = 5 the estimate is 0.069 with 1 subject still under observation, the plain interval runs from −0.052 to 0.191 and the log-log interval from 0.006 to 0.251, against a truth of 0.174. The log-log band is built on a scale that cannot leave [0, 1], and it bends away from the edge rather than through it.

The interval at the end of the curve

The interval most software prints around a survival curve covers 89.7% at five years, where 3.3 of forty subjects are still being watched and where the curve is actually read. The same variance carried on a log–log scale covers 94.8% there — and the failure was never the width.

survival · Censoring
Twenty cells of an interval that is exactly 95%, 1,000 replications each. The t interval covers exactly 95% in every cell. Estimated at 1,000 replications its cells read 93.9% to 96.5%, and 2 of the twenty are flagged by their own ±1.96 standard errors.

A coverage table with its own error

Twenty cells estimating the coverage of an interval that is exactly 95%, at a thousand replications each, read from 93.9% to 96.5% — and a table like that flags at least one of its correct cells on 69.9% of honest runs. Ten times the replications does not repair it: at ten thousand the same table still flags one 63.3% of the time.

method · Seeds
Forty O'Brien–Fleming trials at a true effect of 0.16, with the boundary written as an effect. The dashed line is the smallest effect a trial can report and still stop at each look: 0.510 at 80 observations, 0.255 at 160 observations, 0.170 at 240 observations, 0.128 at 320 observations, 0.102 at 400 observations. The true effect is 0.16, so at 3 of the five looks a trial cannot stop without reporting more than it. 29 of these forty trials stop before the last look, each marked where it stopped.

The effect a stopped trial reports

An O'Brien–Fleming trial at 88.45% power holds its error rate exactly and reports an effect 9.6% too large on average. The 11.39% of trials that stop at the second look report 1.83 times the truth, the ones that cross at the last look report 0.80 times it, and pooling every trial by its size gives the truth back to the last digit.

sequential · Stopping
What a fixed-width interval covers, by the number of blocks the trial ran before it stopped. Two thousand runs of each rule, the modelled weighting, a promise of 0.34. Reading its report: 4–8 blocks, 22.3% of runs, 78.2%; 9–12 blocks, 16.6% of runs, 90.4%; 13–16 blocks, 18.4% of runs, 96.2%; 17–20 blocks, 17.4% of runs, 96.0%; 21–28 blocks, 17.9% of runs, 96.4%; 29–36 blocks, 7.4% of runs, 99.3% — 91.45% overall. Reading the arms: 4–8 blocks, 0.0%, none; 9–12 blocks, 0.9%, 94.4%; 13–16 blocks, 30.4%, 95.6%; 17–20 blocks, 50.0%, 93.9%; 21–28 blocks, 18.0%, 94.4%; 29–36 blocks, 0.7%, 92.3% — 94.50% overall.

The trials that stopped early

A fixed-width trial that stops when its own interval is short enough covers 91.45% — an average of 78.2% among the 22.3% of runs that stop within eight blocks and 96% to 99% among those that run longer. Widening every interval by 17.1% brings the average to 95% and leaves the early stops at 85.6%, while 92.8% of runs now report an interval wider than the width they promised. Even doubling every interval leaves the early stops short.

stop · Width
The distribution the table does not have. 599 series simulated from the smaller model fitted to one comparison's own data, the whole rolling comparison re-run on each, and the ordinary statistic recorded. Under this null the two forecasts are the same forecast in population, so what is left in a sample is the larger model's estimation error and the statistic is centred at -1.134 rather than at zero. Its 95% point is 0.264; the standard normal drawn behind it puts that point at 1.645. Reading this statistic against that curve is not a poor approximation, it is a different distribution: the share of this one above 1.645 is 0.2%.

A distribution drawn from the null

Between nested models the ordinary comparison statistic has a null distribution centred at minus one and a 95% point of a quarter. A correction to its mean repairs the centre and leaves the shape; simulating the null repairs both.

ranking · Bootstrap
What a schedule is allowed to read, and what happens when it reads more. The construction allows the block sizes to be anything at all as long as they are functions of the within-block contrasts, which are independent of every block mean. A schedule that shrinks the block whenever the between-block spread is running above what the contrasts say is a direct attempt to hold down the quantity the interval will be built from, and it succeeds: the estimate lands at 0.8373σ² against the honest 0.9831, and the coverage goes with it. Reading the running mean instead pushes the other way and over-covers — which is not a repair, it is the same violation with the sign reversed, and the level is no longer a property of the procedure at all.

A schedule that reads the mean

The block sizes may be anything at all provided they are functions of the contrasts. Two natural schedules break that, in opposite directions — and the most natural mistake of the three is not a schedule at all but a stopping rule, at 86.87% coverage and fewer observations.

pace · Stopping
What balancing several numbers at once costs each of them. The criterion generalises without a word changing — the covariate imbalance becomes a vector and the correction a quadratic form — so the question is what it is worth rather than whether it can be done. At n = 200 with 200 trials per point, a rule balancing one covariate leaves 12.7% of a coin's imbalance in it; balancing eight leaves 23.2% in each. The assignment has a fixed amount of freedom and every covariate added takes a share of it. The rule degrades rather than failing: at eight covariates it is still four times better balanced than a coin, and the eight are being held simultaneously rather than in turn.

Balancing more than one number

The criterion generalises to several covariates without a word changing, which makes the question what it is worth rather than whether it can be done. Each one added takes a share of the assignment's freedom, and the imbalance left in every one of them rises.

continuous · Blocking
Where a walk is cheaper than a hunt. Both costs in the same unit. A rejection sampler evaluates 1/p assignments per independent draw and does not care how large the trial is; a walk evaluates one per step and yields an effective draw every τ steps, and τ is a property of the constraint and the statistic together. They cross at a tolerance of 0.194 standard deviations, where about one assignment in 396 is admissible — far tighter than any trial is designed at. And the walk does not remove the acceptance cost; it pays it once, hunting for somewhere to start.

Draws that repeat each other

A hunt costs 1/p evaluations per independent draw. A walk costs one per step and yields an effective draw every τ steps. Both are counted in the same unit, and the walk is dearer at every tolerance a trial is designed at.

joint · Reference
What a guesser gets, and what a guesser gets for nothing. 600 trials of 150 patients under a fully deterministic rule, with an investigator who knows the rule, the factors and every assignment so far. The guess rate falls with the number of arms — 87.8% at 2, 86.2% at 3, 81.0% at 4 — which reads like a trial getting safer and is not: what a guesser can trade on is the excess over the 50%, 33%, 25% they would get by naming an arm at random, and that goes the other way, from 1.76× chance at two arms to 3.24× at 4. The gap a guesser manufactures between the best and worst arm under a true null is 0.759, 0.733, 0.711 standard deviations — nearly unchanged.

Guessing one arm in three

A balancing rule is guessable because it is balancing. With three arms the next assignment is worked out less often than with two — and by more, relative to what a guesser gets for nothing, and the damage they can do is almost unchanged.

multiarm · Assignment
A bias against a variance, with the answer in between. How wrong one sample's reference distribution is, split into the two things it is wrong by. Sharing the multiplier over more rows keeps more of the dependence and closes the bias from 1.688 to 0.835; every row it is shared over also removes an independent sign from the 101 the sample started with, and the spread of the resulting quantile rises from 1.307 to 2.172. The distance a practitioner with one sample is actually exposed to is the two together, and it is smallest at ℓ = 5.

How long a block a multiplier shares

Sharing a sign over more rows keeps more of the dependence and leaves fewer independent signs to build a distribution from. The bias falls from 1.6885 to 0.8479 and the spread rises from 1.3073 to 2.1716, and the rejection rate walks straight through its nominal level on the way from 11.3% to 1.3%.

proxy · Reference
What the corner costs when the table is full of hopeless candidates. The benchmark holds two predictors, one of which is worth 1; a third predictor, worth the amount on the horizontal axis, is held only by candidates the benchmark does not contain. At the left the null is true and both procedures hold their level. To the right there is a genuinely better candidate, and the uncorrected reality check finds it 2.7% of the time while the same test with the clearly bad columns recentred finds it 51.0% of the time. The columns doing the damage are the ones nobody would have looked at twice: they are so far behind that they cannot win, and calibrating as though they might is what makes the test blind.

The corner the test is calibrated at

"No candidate is better than the benchmark" is not a null but a face of a region, and a reality check is calibrated at one corner of it. Fill the table with candidates that are hopeless rather than equal and the test finds a genuine improvement 0.0% of the time.

select · Reference
What the diagnostic says at two hundred units. The same test run 8 times on independent streams, at five tolerances of a two-hundred-unit trial, 40,000 steps each. At the loosest tolerance every run says the same thing — the walk reaches the whole set — and it keeps saying it as the set is thinned. Past a point the runs stop agreeing with each other: at the tightest tolerance here 6 of 8 report that the chains have not mixed and 2 report a split, which is a diagnostic disagreeing with itself rather than a property of the set. That disagreement is the honest answer at this size, and it is one nothing in this collection could give before: an enumeration stops at about twenty-four units.

The diagnostic at two hundred

Pointed at a trial size no enumeration reaches, the test gives three answers rather than one — and past a certain thinness it stops agreeing with itself, which is the honest reading and the one nothing could give before.

reach · Randomisation
Where the taper's case begins, and it is not where the algebra says. The block length at which a trapezoidal block's implied variance stops being more biased than a rectangular one's, against the length of the sample. Computed exactly — from the law's own autocovariances, with no sampling in it — the answer is 19.2 and does not depend on the sample at all. What a sample of 120 rows reports is 13.3, and the reported crossing walks out towards the exact one as the sample grows: 13.3, 15.0, 16.4, 18.0. The mechanism is that the autocovariances the window is applied to are themselves attenuated, worst at the longest lags, and the window that discards those lags loses less of them.

The error no window repairs

Every block window's best estimate of a long-run variance is wrong by about forty per cent at a hundred and twenty rows, and the largest part of that is not a bias at all. Choosing the window moves a twentieth of it.

crossing · Bootstrap
Four intervals at one stopping time. 3,500 experiments under the sequential rule with a first stage of 5 and a required half-width of 0.4, which is a demand that knowing σ would meet with 24.0 observations and which the rule meets with 20.8. The rule's own interval covers 90.3%. Replacing the fixed width by a t interval on the same data gives 92.0%. Keeping the rule's own random sample size and drawing a fresh sample of that size gives 89.8% at the fixed width and 95.5% for a t interval — so the sample size being random costs nothing, and the sample size being chosen by the data the interval is built from costs the rest. The spread estimated at the stopping moment is 17.1% below the truth, which is the same fact one level down.

The interval after a stop it chose

A rule that stops when the estimated precision is good enough stops on the samples whose estimate was small. Its interval covers 90% and claims 95%, and a fresh sample of the same random size covers 95.4%.

guarantee · Stopping
What the interval covers once the order is chosen as well. 1200 series of 40 observations from an AR(1) with φ = 0.7, at each horizon, on one set of seeds. The upper line is the interval computed at the true parameters — it covers 95.5% on average, which is the check that σ²Σψ² is the right formula rather than a claim about anything a forecaster can do. The lower line is the same formula fed σ̂² and φ̂: 93.6% at one step and 90.8% at 6. The third line chooses the order by AIC from the same data before computing the interval, which costs a further 0.8 points at h = 6.

The interval after the choice

Estimating the coefficients of a known model costs a 95% forecast interval about two points of coverage. Choosing which coefficients to estimate, from the same forty observations, costs another four and a half — so the step nobody records in the output is the more expensive of the two.

forecast · Order-selection
Where the maximum is, from 15 runs. One dataset, one fitted quadratic, and two answers to "where is the best setting". The delta method reports 0.80 ± 0.46, a finite interval it will report whatever the data does. Fieller's set is 0.49 to 1.76, because the curvature here has t = -4.04. The true optimum is at 0.75.

The optimum is a ratio, and its interval is sometimes the whole line

The best setting is −b₁/2b₂: a ratio of two estimates whose denominator is a curvature the design can often barely see. The delta method reports a finite interval every time and covers 68.8% where the curvature is weak; Fieller's set covers 95% and says so by being unbounded.

surface · Optimum
The correction does not arrive at the truth, it passes it. The average decay factor a forecast applies to the last observation, at φ = 0.85 and 50 observations, 3000 series per horizon. The middle curve is φʰ, what the model actually does. Below it is the uncorrected forecast, which uses φ̂ʰ and reverts too fast — 24.8% short at h = 4, 30.0% short at h = 6, 32.7% short at h = 8. Above it is the forecast built on the corrected estimate, which overshoots, and the reason is arithmetic rather than a bad correction: raising an unbiased estimate to a power does not give an unbiased estimate of the power, and the higher the power the more the spread of φ̂ is converted into overshoot.

The repair that moves the wrong number

Correcting the bias in a persistence parameter is one line of arithmetic that works. Feeding the corrected estimate into a forecast repairs the number everybody looks at, makes the forecast worse by squared error at moderate persistence, and improves the interval for a reason that has nothing to do with bias.

evaluation · Bias
The scale moves the width; the curve does not. The band width each charge picks, averaged over 150 draws of 120 rows. The two conventions — a unit a lag and half a log n a lag — pick 6.08 and 3.65 lags. The four charges derived from the measured optimism pick 15.05, 15.60, 15.47 and 15.44, against a best width on the draw of 14.13. So the scale a charge is levied on moves the width by a factor of 4.27 and the shape of the charge moves it by 3.6%. None of the six is an estimate of the draw's own best width: the correlations are -0.006, -0.003, 0.017, -0.001, 0.016, -0.004.

What a better charge buys

Four charges derived from the same measurements pick band widths within six per cent of each other and deliver errors within two per cent of the gap any of them leaves. The scale a charge is levied on decides the width; the shape of the charge decides nothing.

curve · Criterion
What the interval covers, after a design that read the data. 800 experiments of 12 runs, all at the same truth. The first pair is a design fixed in advance; the second is one whose settings were chosen from the first stage's own outcomes. If choosing the design from the data broke the inference, the second pair would sit below the first, and it does not — 91.3% against 92.6%. What does move the coverage is the shape of the interval rather than the design: the Wald interval assumes the estimate is normal around its own standard error and falls short under both designs, and the profile interval, computed from the same residual sums of squares with no derivative in it, covers 94.3% and 94.9%.

What a design chosen from the data costs

Two fields on this site measured what happens when a rule reads the data, and the error rate broke both times. A design that reads the data to decide where to put its runs breaks nothing — and the control that proves it also finds what the real shortfall is.

robust · Local design
The rate falls geometrically; the count does not fall at all. The share of equal splits of two hundred units that a tolerance of 1 coin-spreads admits, against the number of functions the tolerance is stated for, with the closed form (2Φ(1) − 1)^k drawn beside it. The rate falls by about two thirds with every constraint. The admissible count is that rate times C(200, 100), and it goes from 2^195 to 2^192 — it does not fall in any sense a trial cares about. What the falling rate costs is sampling: 9,878 draws to collect a thousand admissible ones at six constraints, against 1,465 at one.

What a reference distribution costs to sample

A randomisation test on a trial too large to enumerate has to sample its reference distribution, at 1/p attempts per draw and a p-value resolved to 1/(B + 1). Six constraints cost 9,878 attempts per thousand draws, and a thousand draws resolve p to 9.99·10⁻⁴ and not one digit finer.

product · Reference
Fitting them together is worth something under one law. Four fits of the same regression under four dependences: least squares, the two-step plug-in every whitened rule in this collection runs, the coefficients and the band maximised together, and a whitening at the law's own covariance that nobody has. Under the moving average — the one law the band family contains — the joint fit beats the two-step by 0.0077 at 3.3 paired standard errors. Under the autoregression, long memory and the break it is a tie: 0.4, 1.0, 0.3 standard errors. That is the same ordering the likelihood gap gave, arrived at through the coefficients rather than through the objective.

What fitting them together buys

Maximising over the coefficients and the covariance together beats the two-step under one of four dependences and ties under the other three. It is the one the band family contains, and the likelihood said so before any coefficient was compared.

family · Dependence
What the interval actually covers. The coverage of the two-sided interval each rule and window builds, over 400 samples of 120 rows, against the 95% it promises. Not one of the eight reaches it: the best is 91.0% and the worst is 80.8%, on a promise of 95%. So the first thing this instrument says is that the choice between the two windows is a choice inside a range that is already four to fourteen points short, which neither of the other two readings can express at all. The second is the ordering: the tapered window covers better under every one of the four rules, by 5.00, 2.50, 5.75 and 4.00 points — including at a length written into a protocol and at the rule of thumb, where the implied variance says the rectangle wins.

What the interval covers

Eight rules and windows, and not one of them reaches its promised 95%. The range is 80.8% to 91.0%, and the choice between two block windows is a choice inside a shortfall that is four times larger.

readout · Bootstrap
The same dial, on a list that steps by one. The probability that a per-candidate tuning list changes which candidate the table selects, at each of 5 separations, for both tuning parameters at a matched list length of 8. The sieve order runs from 13.1% to 2.6%, a factor of 5.00; the whitening window, from 16.4% to 1.5%, a factor of 11.75. What is held is the number of options, the table, the law, the sample size and the seeds; what cannot be held is the size of a step, since an integer step and a geometric step are different amounts of change. The dial moves both, and it moves them by 2.35 times as much on one as on the other.

A step that is not a ratio

Run the separation sweep on a tuning list of integers rather than a geometric ladder and the two factors still point opposite ways. The invariant does not survive: along a row of integers the probability moves by 2.163 where along the geometric ladder it moves by 1.208.

turnover · Order-selection
The extra 1/m, and the correction nobody quotes. What a pooled 95% interval covers against the number of imputations, counted over 2000 studies of 200 rows at 35.0% of outcomes missing. Rubin's rules — total variance W̄ + (1 + 1/m)B, read against a t distribution on (m − 1)(1 + W̄/((1 + 1/m)B))² degrees of freedom — cover 94.10% at two imputations and reach their promise by 5, at 95.25%. Dropping the (1 + 1/m) factor takes two imputations to 93.10%; using a normal quantile instead of the degrees-of-freedom correction takes it to 92.55%; dropping both takes it to 91.45%. The median degrees of freedom at two imputations is 12.95, which is why the second correction is the larger.

The variance between imputations

Pooling several filled datasets covers 94.10% at two imputations and reaches its promise at five, where a single fill covered 85.78%. The correction everybody quotes is the smaller of the two doing the work — 1.00 ± 0.22 points against 1.55 ± 0.28.

missing · Missingness
What each correction is worth, exactly. Each variance estimate's expectation under a constant error variance, divided by the variance the slope actually has, at 20 rows on an even design, by two routes: the closed form E[eᵢ²] = σ²(1 − hᵢᵢ) carried through each correction's own weight, and the mean of 20000 counted estimates. The maximum leverage here is 0.1857 and the design's fourth-moment share Σu⁴/(Σu²)² is 0.0897, which is the only thing the closed form reads. HC0 comes out at 0.8603 — short by construction, since its factor is exactly 1 − 1/n − Σu⁴/(Σu²)². HC1 reaches 0.9559, HC2 is exactly 1.0000 at every design and every sample size, and HC3 overshoots to 1.1647. On an even design the four are within a fifth of each other and the choice barely matters.

Three corrections and a leverage

On an even design of twenty rows the four robust corrections read 0.8603, 0.9559, 1.0000 and 1.1647 of the truth and the choice barely matters. Add one point at x = 8 and they read 0.3191, 0.3419, 1.0000 and 5.1127.

sandwich · Misspecification
Two groups, a baseline and a follow-up, and nothing happening in between — baseline reliability 0.6. 600 units in two pre-existing groups whose true means are 1.00 apart, read once at baseline and once at follow-up, with no change for anybody. The two groups' mean changes are −0.075 and −0.032, so the change-score analysis reports a group difference of 0.043. The regression of follow-up on baseline and group reports 0.409, against a closed form of (1 − λ) × 1.00 = 0.400: at any one baseline reading the two groups' lines sit that far apart, because each group's units regress towards their own group's mean. The pooled slope in this sample is 0.614, the baseline's reliability.

Two analyses of one baseline

Two groups read at baseline and again at follow-up, with no change for anybody. Subtracting the baseline reports a group difference of −0.0014 and adjusting for it reports 0.4008 — and each analysis is exactly right about one reason the groups started apart and wrong by 0.40 about the other.

paradox · Rtm
Which samples Wilson and Clopper–Pearson each cover, n = 50, p = 0.2. Each bar is the probability of one count, shaded by which interval built on that count contains 0.2. Both cover 95.1% of samples, only Wilson 0.0%, only Clopper–Pearson 1.6%, neither 3.3%. The correlation between their hits is 0.810, so on shared draws the variance of their difference is 4.891 times smaller than on independent ones.

The same draws for both methods

Two intervals computed on the same simulated datasets give a difference in coverage whose variance can be 4.891 times smaller than on separate datasets — or, for a pair that covers different samples, 1.164 times larger. Which one a comparison gets is an exact sum over the counts each interval covers, and a standard error that ignores the sharing covers 100.00% for one pair and 93.07% for the other.

method · Seeds
The false discovery rate of twenty correlated tests, against the correlation. BH, every null true: 5.08% at 0, 4.86% at 0.3, 3.70% at 0.6, 2.34% at 0.9. BH, 10 of 20 real: 2.55% at 0, 2.53% at 0.3, 2.26% at 0.6, 1.66% at 0.9. BY, every null true: 1.46% at 0, 1.31% at 0.3, 1.03% at 0.6, 0.69% at 0.9. BY, 10 of 20 real: 0.72% at 0, 0.75% at 0.3, 0.64% at 0.6, 0.50% at 0.9. 20,000 families at each correlation.

False discoveries that arrive together

Correlate twenty tests and Benjamini–Hochberg still holds its false discovery rate — 1.66% at a correlation of 0.9 with ten real effects, against 2.55% when the tests are independent. What changes is how the errors come. A family of true nulls reports anything 2.34% of the time instead of 5.08%, and when it does, it reports 16.56 false findings out of twenty.

multiplicity · Multiplicity
The fixed-width trial's coverage when the outcomes are not normal, for both stopping rules. normal: stopping on the arms 94.05% after 18.1 blocks, on the report 89.95%; log-normal, skewness 0.95: stopping on the arms 94.70% after 18.5 blocks, on the report 90.80%; log-normal, skewness 2.26: stopping on the arms 94.15% after 19.3 blocks, on the report 90.25%; log-normal, skewness 4.75: stopping on the arms 94.45% after 18.7 blocks, on the report 90.50%; t, five degrees of freedom: stopping on the arms 94.35% after 18.3 blocks, on the report 90.30%; skewness 4.75, arm A only: stopping on the arms 93.80% after 26.0 blocks, on the report 89.90%; skewness 4.75, arm B only: stopping on the arms 93.60% after 14.2 blocks, on the report 89.95%; equal variances, normal: stopping on the arms 94.75% after 11.4 blocks, on the report 90.90%; equal variances, skewness 4.75: stopping on the arms 94.05% after 11.1 blocks, on the report 92.00%.

A width rule on skewed outcomes

The blinded fixed-width rule rests on a within-arm spread being independent of the arm means, which only normal samples guarantee. On outcomes with a skewness of 4.75 the independence fails and the overall coverage barely notices — 93.60% to 94.70% across every shape counted, against 94.05% on normal outcomes. What skew moves is the runs that stop by twelve blocks, which cover about 90% with the skew in one arm, and the trial's length: a variance ratio corrected on normal theory lengthens it from 18.1 blocks to 26.0 with the skew in the first arm and shortens it to 14.2 with the skew in the second.

stop · Width
The ceiling a multiplier cannot reach past. A wild-type resampling forms e*_t = e_t·w_t with the multiplier independent of the residual, so what comes out has autocovariance γ_resid(k)·γ_w(k) — the residuals' own, multiplied by the multiplier's. Since |γ_w| ≤ 1 the reference distribution's dependence is bounded above by the residuals', and the residuals' is already below the errors'. The two shortfalls compose. For a block of ℓ the multiplier's autocorrelation is exactly the triangle (1 − k/ℓ)⁺, drawn here as the dashed prediction against the realised resamples at ℓ = 5; the bound is attained only at ℓ = n, where the reference distribution is built from one sign.

What a multiplier cannot keep

Two reasons were named for the quarter a blocked resampling falls short, and taking either away makes the gap larger. What is left is a bound — a multiplier can only take dependence out, and the residuals' own is already below the errors'.

effective · Reference
Largest where least is needed. What the pairs correction supplies against what each window's measured profile needs, across this field's plateau, over 2000 draws at 120 rows. Both are stated as the multiplicative rise the charge per unit of width has to take between four lags and thirty. What the correction supplies is arithmetic — (1 − μ(4)/n)/(1 − μ(30)/n), where μ is the mean lag of the weight the band adds — and it runs 1.0795, 1.0580, 1.0456, 1.0539 for the four windows. What the measurement needs runs 1.1076, 1.2928, 1.2296, 1.6550. The two orderings are opposite: the plain Bartlett window has the longest mean lag, so it gets the biggest correction, and the flattest profile, so it needs the smallest. They coincide to 0.9746 of each other, and nowhere else does the correction account for more than 85.0% of the fall.

What the correction assumes

A correction with nothing fitted in it repairs one window of four. The reason is that its size is set by where a window puts its weight and the curvature it must repair is set by something else — and for one window at one sample size the two happen to agree.

curve · Criterion
The line is the sample, and the sample is the finding. 900 draws of two independent standard normal causes, with the 453 of them past a threshold of 0.00 marked and the 447 that fall short left pale. In the population the two are independent by construction. Inside the selected sample the correlation is -0.4669 in closed form and -0.5050 counted on these 453 rows, and the least-squares line through them has a slope of -0.545. The mechanism is visible in the picture rather than argued: the threshold removes one corner of the cloud, and a cloud with a corner missing is a cloud whose two coordinates carry information about each other.

The sample is a condition

Two independent standard normals, selected on their sum exceeding its median, read a correlation of exactly −1/(π − 1) = −0.4669 inside the sample. Nothing is measured badly and nothing is missing — and both halves of that split read it, in the same direction, while the population containing both reads zero.

collider · Conditioning
The p-value of a study with 80% power, twenty thousand times. Twenty thousand two-sided z-tests, each on 25 observations whose true mean is 0.5603 standard deviations from the null, a noncentrality of 2.802. The bars are the counted share of p-values in bins a quarter of a power of ten wide, with the leftmost bin holding everything smaller; the line is the closed form. The middle eighty per cent of the p-values runs from 4.4×10⁻⁵ to 0.13, 3.46 orders of magnitude, the median is 0.0051, and 80.0% fall below 0.05, which is what the power means.

The p-value a replication gets

Under a true null a p-value is flat. Under a real effect its distribution is closed form and wide — a study with 80% power returns anything from 4.4×10⁻⁵ to 0.13 in eight runs of ten — and the chance that an exact replication of a p = 0.05 result is significant again is exactly one half, under both of the models people use without naming them.

testing · Uniformity
A cohort screened once, the top tenth enrolled, and followed up with nothing given — correlation 0.6. 2000 people read once at screening and once at follow-up, with a test–retest correlation of 0.6 and no treatment. The 215 above a cut at the top ten per cent of one reading (1.282 standard deviations) are enrolled. Their mean screening reading is 1.744 and their mean follow-up reading 0.982, a fall of 0.762 ± 0.055 with nothing done to anyone. The closed form for the fall is (1 − ρ) times the truncated-normal mean, (1 − 0.6) × 1.755 = 0.702.

The measurement that got them enrolled

Enrol the top tenth of one screening reading and give them nothing, and they fall by 0.702 standard deviations at follow-up. Measured from a fresh reading taken after enrolment they fall by nothing. Averaging ten screening readings still leaves 0.101, and it takes twenty-one to get under 0.05.

paradox · Rtm
Twenty runs simulating an exactly 95% interval, checked every 250 replications. Each line is one run's running estimate; the dashed band is where the Wilson interval of the running estimate still contains 95%, and a run stops, marked, the first time it leaves the band. 8 of these twenty stop before 10,000 replications. The exact probability of stopping, from the recursion over the count, is 29.54%.

A simulation that stops when it looks settled

A simulation of an interval that covers exactly 95%, checked every 250 replications for a significant departure and stopped when it finds one, flags that correct interval on 29.54% of runs. Stopped instead as soon as its estimate reaches 95%, it reports an interval that covers 94% as meeting its level on 37.21% of runs. Stopped when the estimate stops moving, it reports the right number — and has quietly chosen to run about fifteen hundred replications.

method · Seeds
Forty trials at a true effect of 0.16, under the rule "power at the trend < 10%". The upper line is the benefit boundary (4.56, 3.23, 2.63, 2.28, 2.04); the lower line is where the rule stops a trial for futility (0.40 at 80, 0.66 at 160, 0.95 at 240, 1.31 at 320). Of forty trials with a real effect, 29 cross for benefit and 11 are stopped for futility.

A boundary for giving up

Adding "stop if z is below zero" to an O'Brien–Fleming trial costs 5.20 points of power at the effect it was designed for and halves the observations a trial with no effect uses. Stopping when conditional power at the observed trend falls under 10% costs 13.23 points and stops 21.28% of trials with a real effect. Making that rule binding lowers the benefit boundary from 2.040 to 1.901, and a binding rule that is then ignored rejects a true null 3.523% of the time instead of 2.5%.

sequential · Stopping
Storey's estimate of the share of true nulls over twenty thousand families, independent and correlated at 0.6. The true share is 0.5. Independent tests: mean 0.610, spread 0.160, below half the truth in 0.92% of families. Correlated at 0.6: mean 0.609, spread 0.240, below half the truth in 9.33%.

Estimating how many nulls are true

Benjamini–Hochberg at 5% delivers 2.55% when half of twenty nulls are false, because it cannot tell how many are. Storey's estimate of that share, read off the p-values above one half, spends the rest and finds 81.93% of the real effects instead of 74.70% on independent tests. Correlated at 0.9, the same procedure reports a finding in 19.29% of families in which every null is true.

multiplicity · Multiplicity
The average decay factor each route produces, φ = 0.85, 6 steps ahead. The truth is φ^6 = 0.3771. no correction averages 0.2616 with a spread of 0.1646 and a squared forecast error of 3.2516; the formula, on the persistence averages 0.4213 with a spread of 0.2528 and a squared forecast error of 3.4827; the bootstrap, on the persistence averages 0.4355 with a spread of 0.2655 and a squared forecast error of 3.5120; the bootstrap, on the decay factor averages 0.3375 with a spread of 0.2278 and a squared forecast error of 3.4132. 800 series, 100 bootstrap refits each.

Correcting the forecast instead

The complaint against the usual repair is that a correction aimed at the persistence lands on the wrong quantity. Aiming it at the decay factor the forecast actually uses fixes exactly that — the error stops compounding with the horizon, 69.7% becomes 9.5% at twelve steps — and the forecast still gets worse.

evaluation · Bias
An exact test rejecting a true hypothesis a fifth of the time. How often each analysis reports an effect when the average treatment effect is exactly zero and the effect varies between units, at 150 units with 25% treated. The permutation test on the difference in means reads 4.20% where the effect is constant — where the two nulls coincide and its exactness applies — and 22.93% where the effect varies with a standard deviation of 3. The same test on the studentised difference reads 6.27% there, and the ordinary large-sample t, which makes no exactness claim at all, reads 6.60%.

The null the exactness is for

A permutation test is exact under the hypothesis that the treatment changed nothing for anybody. Under the hypothesis it changed nothing on average, with a quarter of the units treated and the effect varying between them, it rejects a true null 22.93% of the time.

exact · Nuisance
What the guess is worth, when it is worth anything. The variance cost of an even split relative to the variance-minimising one for a risk difference, against the first arm's proportion, with the second at 0.3. The cost is a pure number: it does not depend on the trial's size. It is exactly zero at 0.3 and at 0.70, where the two arms have the same p(1 − p); it is 0.19% at a half and 4.36% at a tenth. Across the whole range from a tenth to nine tenths it never exceeds 4.36%, which is what the variance-minimising rule is worth here — and what it is worth is the reason it is safe to use with a guess.

The arm whose variance is its answer

With a binary outcome the allocation rule is a function of the proportions the trial exists to estimate. It costs at most 4.36% of variance to ignore it anywhere between a tenth and nine tenths, because √(p(1−p)) stays within a factor of two of its peak across 98% of the unit interval.

allocation · Allocation
The damage and the warning, against the same dial. Two readings at each persistence. In the darker colour, how often a regression between two independent series of 200 steps is called significant at 5%: 4.9% at φ = 0, 34.2% at 0.8, 52.4% at 0.9, 83.4% at a unit root. In the lighter, how often the standard unit-root test refuses a unit root on one of those series — the chance the analyst is told the series is stationary and may be regressed: 87.2% at φ = 0.9 and 31.9% at 0.95. At φ = 0.9 both are high at once, which is a correct diagnostic licensing a regression that is wrong half the time.

The cliff that is a slope

A regression between two independent series is called significant 4.9% of the time at no persistence, 52.4% at a lag-one correlation of 0.9, and 83.4% at a unit root. The rule the field offers asks whether the last of those holds, and at 0.9 the unit-root test correctly refuses one 87.2% of the time.

timeseries · Spurious
The bounded error and the unbounded one. How the sequential trace procedure's answer is distributed, against the sample length, for a three-series system with 2 genuine relations. Over-counting — claiming a stationary combination that is a random walk — reads 4.9%, 7.2%, 5.7%, 6.2%, 5.9%, 4.2% across the six lengths, never far from the 5% of a single test. Under-counting reads 69.5%, 40.2%, 14.0%, 0.5%, 0.0%, 0.0%. The procedure is described as a 5% rule and the 5% applies to one of those columns.

The rank is a decision

The sequential procedure's 5% bounds one of its two errors. Over-counting reads between 4.2% and 7.2% at every sample length from fifty observations to three hundred; under-counting reads 69.5% at fifty and 0.0% at three hundred, and nothing in the procedure bounds it.

systems · Rank
Reading the draw changes what is charged, not what is tracked. The correlation between the band width each rule picks and the best band width on the same draw, over 400 draws. The three fixed charges read -0.069, -0.066, 0.012. The three that read the sample read -0.012, -0.019, -0.041. None of the six is distinguishable from nothing. The statistic the first plug-in reads does vary — the draw's own summed squared autocorrelation runs from 2.06 to 10.19 with a mean of 4.36 — so the failure is not that the charge stopped moving. It is that what it moves with carries no information about which width this draw wanted.

A charge that reads the draw

Three charges built to read the sample track the best band width on their own draw at −0.012, −0.019 and −0.041, deliver more error than the fixed rule they are calibrated to, and pick a width half again as variable. The statistic moves; the answer does not.

curve · Criterion
The top of a heavy-tailed population keeps its lead; the top of a light-tailed one gives it back. Select the top share on the first reading and read the group again: the share of its mean lead the second reading keeps, by integration over the true score (lines) and counted on 400,000 draws a parent in 20 batches (points, with two standard errors). The normal keeps exactly 0.6 at every selection. At the top half the Laplace keeps 0.541, the t 0.535 and the uniform 0.648 — the heavy tails keep LESS than the correlation. By the top one per cent the order has reversed: 0.761, 0.784 and 0.443. At one in ten thousand the t keeps 0.977 and the uniform 0.346.

A lead that a heavy tail keeps

Four populations whose readings all correlate at exactly 0.6, and whose least-squares slopes all read 0.6. Select the top one per cent on one reading and measure them again: they keep 60% of their lead if the true scores are normal, 76.1% if they are Laplace, 78.4% if they are a t on four degrees of freedom — and 44.3% if they are uniform. The correlation predicts the regression of the extremes for one shape of population only.

paradox · Rtm
Estimating P(Z > 5) = 2.8665×10⁻⁷ with plain draws and with four proposals. One seed each. Plain simulation draws nothing past 5 in 100,000 and estimates zero throughout. At 100,000 draws the proposal N(5, 1) reads 1.009 of the truth, N(9, 1) 1.041, N(4.5, 0.25²) 1.039 and N(5, 0.3²) 1.008. Values above 2.2 are drawn at the top edge.

The draws aimed at the tail

The chance a standard normal exceeds 5 is 2.8665×10⁻⁷, and a plain simulation needs 349 million draws to estimate it to within ten per cent. Draws aimed at the tail and weighted back need 565. Aimed slightly too narrowly, the same method has an infinite variance, an interval that covers 86.0% and gets worse with more draws, and an effective sample size that reads healthier than a proposal that works.

method · Seeds
Three intervals as one strength is spread thinner, at a concentration of 8. Coverage of three nominal 95.0% intervals on the same 1000 draws of 200 rows at each count, when a total concentration parameter of 8 is spread over 1 to 32 instruments. Two-stage least squares covers 97.2%, 96.4%, 94.0%, 86.7%, 73.2%, 51.5%. Building each row's fitted treatment from a first stage that never saw that row covers 97.1%, 97.2%, 98.3%, 97.8%, 97.9%, 98.7%. Limited-information maximum likelihood with its conventional standard error covers 97.2%, 96.7%, 95.7%, 90.9%, 85.2%, 79.0%. At one instrument the likelihood estimator is two-stage least squares exactly, which is why the first readings of those two agree to the last draw.

Leaving each row out of its own first stage

Spread a fixed first-stage strength over thirty-two instruments and two-stage least squares covers 51.5%. Build each row's fitted treatment from a first stage that never saw that row and the same draws cover 98.7% — through an interval 5.99 times as wide, around an estimate that misses by more than the whole effect on 34.7% of draws. At eight times the strength the same repair covers 95.3% and costs a width factor of 1.66.

instrument · Exclusion
Three ways to reject with two studies, drawn where the two z statistics live. Two one-sided studies, each summarised by its z statistic. Fisher's combination rejects outside a curve that runs parallel to both axes, so one study past z = 2.378 decides it alone; Stouffer's rejects above the straight line z₁ + z₂ = 2.326; Tippett's rejects when either z passes 1.955. Each region holds exactly 5% of the standard bivariate normal — Fisher's in closed form, e^(−c/2)(1 + c/2) at c = 9.488 — and of 100,000 counted null pairs they catch 4.95%, 4.88% and 5.04%. Two alternatives carry the same Stouffer evidence: one study at 2.326 and the other at nothing, where the powers are 62.7%, 50.0% and 65.4%; and both at 1.163, where they are 47.7%, 50.0% and 38.3%.

Two ways to combine p-values

Fisher's and Stouffer's combinations are both exactly right when every null is true, for the single reason that each p-value is flat. Under a real effect they disagree about which evidence counts: with Stouffer held at 50% power across ten studies, Fisher is the more powerful while the signal sits in six or fewer of them and the less powerful from seven.

testing · Uniformity
The chance of crossing later from each interim z, under six schedules with O'Brien–Fleming-type spending boundaries. Exact. At an interim |z| of 1.0: end only 3.64%, +0.6 3.41%, +0.75 3.20%, +0.9 3.41%, every 0.125 3.00%, every 0.05 2.88%. At 2.5: end only 38.30%, +0.6 45.80%, +0.75 45.36%, +0.9 41.70%, every 0.125 50.08%, every 0.05 53.24%. The heavy line is the largest of the six at each z.

A look the trend asked for

Under an O'Brien–Fleming-type spending function, every schedule of looks fixed in advance spends exactly 5.0000%. A committee that adds a look at three quarters of the trial whenever the interim z is 1.5 or more spends 5.2323% — 5.315% counted over a hundred thousand trials — and the most a committee choosing among six schedules could spend is 5.4390%.

sequential · Stopping
Twenty hypotheses tested in a declared order, the ten real effects listed first. Effects of three standard errors, ten real, familywise 5%. fixed sequence: 85.3% at position 1, 45.0% at 5, 20.4% at 10; overall power 46.10%; fallback: 49.1% at position 1, 56.4% at 5, 57.9% at 10; overall power 55.87%; Holm: 52.5% at position 1, 52.2% at 5, 52.5% at 10; overall power 52.53%.

An order that spends the error rate

Test twenty hypotheses in a declared order, each at the full 5% and each only if every one before it was rejected, and the first is found 85.3% of the time where Holm finds it 52.5%. The tenth is found 20.4% of the time, the product of the powers before it. Move one true null to the head of the list and every real effect behind it is found no more than 4.3% of the time.

multiplicity · Multiplicity
Five treatments of an estimate above one, φ = 0.95, n = 25. The correction exceeds one on 31.1% of series at this setting. left where it lands: squared forecast error 12.828, average decay factor 0.7974 against a true 0.7351; capped at 0.995: squared forecast error 5.680, average decay factor 0.5950 against a true 0.7351; capped at 1 − 1/n: squared forecast error 5.535, average decay factor 0.5256 against a true 0.7351; correction scaled to fit: squared forecast error 5.535, average decay factor 0.5256 against a true 0.7351; correction refused where it leaves: squared forecast error 5.868, average decay factor 0.4423 against a true 0.7351.

The correction that leaves the region

The bias correction adds (1 + 3φ̂)/n whatever φ̂ is, so it pushes the estimate above one whenever φ̂ exceeds (n − 1)/(n + 3) — on 31.1% of series at φ = 0.95 and twenty-five observations. Five obvious things to do about it differ by a factor of 2.3 in squared forecast error, and none of them is documented as a choice.

evaluation · Bias
What each wrong count costs, 4 steps ahead. Squared forecast error 4 steps ahead at each imposed rank, relative to the correctly specified fit, at 200 observations. With 1 genuine relations, imposing 0 costs 13.3% and imposing 2 costs 4.8%. With 2 genuine relations, imposing 1 costs 15.6% and imposing 3 costs 2.5%. Under-counting is the more expensive mistake in both systems, and it is the one the procedure's level does not bound.

Which mistake about the rank costs

On a system with two relations, imposing none costs 29.2% of squared forecast error and imposing three costs 2.5%. The expensive mistake is under-counting, which is the error the procedure's 5% does not bound — so the guarantee protects the cheap side.

systems · Rank
Too small breaks it and too large does not. Coverage of the weighted interval against the factor the true likelihood ratio is multiplied by, at a test population 80.0% drawn from the noisier group and 200 calibration points. The exact weight is the factor of 1 and covers 95.70%. Overstating it costs nothing: 96.13% at sixteen times too large. Understating it costs, and costs steeply below about a half — 94.93% at half, 88.37% at an eighth and 67.90% at a thirtieth. The question this answers was whether a wrong weight degrades smoothly or falls off a cliff, and the answer is that it does neither symmetrically: the curve is smooth and one-sided.

The weight that has to be estimated

A likelihood ratio sixteen times too large costs 5.5% of interval width and no coverage at all; one a thirtieth of the right size covers 67.90%. The estimate from a batch of five unlabelled covariates covers 95.10% against an exact repair's 95.30%, and the binomial says why.

conformal · Exchangeability
One statistic that is right under both hypotheses. Rejection rates for both statistics under both nulls, at 25% of 150 units treated, with the weak-null readings taken at an effect spread of 3. The difference in means is exact under the sharp null and rejects 22.93% of true weak nulls. The studentised difference is exact under the sharp null — 4.07% — and reads 6.27% under the weak one. The repair is a change of statistic inside the same construction: the same re-randomisations, the same fixed outcomes, a different number compared across them.

A statistic that is exact twice

Dividing the difference in means by its own separate-variance standard error before permuting takes the rejection rate under a true weak null from 20.47% to 6.07%, keeps the exactness under the sharp null at 4.07%, and costs 0.8 points of power against a real effect. At an even split it changes nothing at all, in every draw.

exact · Nuisance
Every finding Benjamini–Hochberg made in thirty families, with its interval, at real effects of 2. 85 findings, sorted by their estimate. 12 of their ordinary 95% intervals miss the true effect, every one of them on the far side; 6 of the wider false-coverage-rate intervals miss.

Intervals for the findings

Benjamini–Hochberg's findings usually go out each with its ordinary 95% interval. With ten real effects of two standard errors among twenty tests, 11.59% of those intervals miss their effect, every miss on the far side, and the interval around the most prominent finding covers 72.36% of the time — 2.38% when the effects are one standard error. Intervals widened for the number of findings hold the share that miss under 5%.

multiplicity · Multiplicity
How often each combination, and each union of them, rejects ten studies of nothing. Each combination alone rejects exactly 5% of null sets. Counted on 1,000,000 sets of ten null studies: Fisher or Stouffer 6.63%, Fisher or Tippett 8.05%, Stouffer or Tippett 8.96%, any of the three 9.66% — enclosed on a two-dimensional lattice between 9.18% and 10.09% — and all three together 1.05%. The three sizes add to 15%.

The smallest of three combinations

Reporting whichever of Fisher's, Stouffer's and Tippett's combinations is smallest is a test of its own, and on ten studies of nothing it rejects 9.66% of the time — not 5%, and nowhere near the 15% the three sizes add to, because the statistics are correlated at up to 0.903. Read at 2.448% each it is exact, and then it trails the best single combination by at most 7.45 points and leads the worst by at least 10.30.

testing · Uniformity
The honest curve, and the same forecaster in three coarse vocabularies. The ROC curve of the honest probability at full signal, area 0.8683, beside the same forecaster rounded to the nearest whole number, area 0.7684 — the two-valued liar — to the nearest half, area 0.8205, and to the nearest tenth, area 0.8650. A vocabulary of v values is v points on the square joined by straight segments, and tied reports count half.

A forecaster that rounds

An honest probability issued in tenths loses 0.0033 of ROC area and 0.000708 of resolution — the variance its bands average away, and 89.5% of the 0.000792 it adds to the Brier score. Two hundred records of two thousand forecasts show that loss on 189; it takes about 3,300 forecasts to put it two standard errors from zero. And 3.207 in every thousand forecasts in tenths are a 0% on an event that happened, which a logarithmic score charges without limit.

calibrate · Calibration
The difference in restricted mean survival at every horizon, in three worlds. Treatment minus control, in closed form, with dropout irrelevant to the truth. The proportional treatment's difference grows to 0.4766 at τ = 3 and the waning treatment's to 0.2675. The crossing treatment's rises to 0.1776 at τ = 2, near where the two survival curves cross, and falls back to 0.1366 at τ = 3. The ticks along the bottom are the eleven horizons, from 0.5 to 3 in quarters, at which a trial below reads its differences.

A horizon chosen after looking

A difference in restricted mean survival read at whichever of eleven horizons looks most convincing rejects 11.24% of trials in which the treatment does nothing, against 4.70% at a horizon fixed in advance. The correlation of the differences across horizons is closed, and the Gaussian process it defines prices the choice at a critical value of 2.317 — which brings the counted size back to 4.99% and keeps 96.92% of the power that a horizon nobody could have known to fix would have had.

survival · Censoring
Three intervals as one strength is spread thinner, at a concentration of 8. Coverage of four nominal 95.0% intervals on the same 1000 draws of 200 rows at each count, when a total concentration parameter of 8 is spread over 1 to 32 instruments. Two-stage least squares covers 97.2%, 96.4%, 94.0%, 86.7%, 73.2%, 51.5%. Building each row's fitted treatment from a first stage that never saw that row covers 97.1%, 97.2%, 98.3%, 97.8%, 97.9%, 98.7%. Limited-information maximum likelihood with its conventional standard error covers 97.2%, 96.7%, 95.7%, 90.9%, 85.2%, 79.0%. The same estimate with Bekker's many-instrument standard error covers 97.2%, 97.2%, 97.2%, 95.0%, 94.3%, 93.8%. At one instrument the likelihood estimator is two-stage least squares exactly, which is why the first readings of those two agree to the last draw.

A standard error that knows about the instruments

Limited-information maximum likelihood came out least biased when a concentration parameter of 8 was spread over thirty-two instruments, and its conventional interval covered 79.0%. Bekker's many-instrument standard error covers 93.8% on the same draws, at 63% of the jackknife's width — and it gets there with a median standard error of 0.561 against a true spread of 0.797, because it is large on the draws that need it. At eight times the strength it covers 94.9% at 91% of the jackknife's width, and nothing measured here beats it.

instrument · Exclusion
Twenty studies of a thousand readings estimate the regression of the extremes: the t, four degrees parent. Each thin line is one study of 1000 readings from the t, four degrees parent: Tweedie's formula with the log-density's slope estimated by a degree-5 log-spline, drawn up to that study's largest reading. The thick line is the share kept by integration over the true score, and the dashed line the correlation, 0.6. The top ten readings of the median study begin at 2.45. Over 400 studies the corrected share kept by the top one per cent averages 0.8190, with a spread of 0.0806, against 0.7842 by integration; a Gaussian kernel averages 0.7914 with a spread of 0.0853.

The slope of a density nobody can see

Tweedie's formula corrects a reading by the slope of the readings' own log-density, and a study has its readings. Estimated from a thousand of them, the correction for the top one per cent beats the correlation's linear rule on 84.0% to 98.0% of studies from heavy-tailed populations and on 75.0% to 81.5% from a bounded one — and costs an error of 0.09 to 0.12 where the population is normal and the rule was already exact. At 250 readings the log-spline loses to the rule it replaces, and at 16,000 the same log-spline gets worse on a power tail.

paradox · Rtm
What a confirmation run at the chosen setting would find. The true optimum is worth 62.348. At σ = 2 the fit predicts 62.679 at the setting it recommends and the truth there is 61.821 — a gap of 0.858, which is 0.72 of the prediction's own standard error. The setting itself gives up 0.527 against the best available.

The run that confirms it

The setting a response-surface analysis recommends was chosen because the fitted surface was highest there, so the height the fit predicts at it is a maximum over a random field. At twice the noise the fit predicts 0.858 more than is there — 0.72 of the prediction's own standard error — and the gap is not noise, it is the selection.

surface · Optimum
How fast a gap has to close before a sample can see it close. The power of the test against the half-life of a disagreement, at 100, 200, 400 observations, each read against its own simulated critical value. Every pair in every reading is genuinely tied together, so a non-rejection is a miss. At 200 observations a gap that halves in 3 steps is found 99.9% of the time and one that halves in 12 steps is found 15.3% of the time — and by 35 steps the reading is 6.1%, which is the test's own size. Beyond that the curves are flat because there is nothing left to detect with.

How slow a return a sample can see

At two hundred observations the test finds a gap that halves in five steps four times in five, one that halves in eight 37.3% of the time, and one that halves in fifty 4.95% of the time — which is the rate at which it finds pairs with no mechanism at all. The boundary moves with the sample, not with its square root.

timeseries · Spurious
One of these converges and the other does not. Two measurements on the same fits, against the sample length, for a system with 2 genuine relations. The distance from the fitted plane to the true plane falls from 0.1438 at 100 observations to 0.0075 at 1600 — halving with each doubling, which is the 1/n rate this field's estimates converge at. The angle between the leading fitted relation and the leading generating one reads 29.6° and 29.0° at those same lengths, and is flat in between. The plane is an estimate; the relation inside it is not.

A space is not a relation

The fitted plane approaches the true one at rate 1/n — 0.1438 at a hundred observations and 0.0075 at sixteen hundred. The angle between the leading fitted relation and the leading generating one reads 29.6° and 29.0° at those same lengths, and never moves.

systems · Rank
Three detectors for one departure, all at 5%. How often each of three checks on the calibration scores fires, against the size of the drift, with every critical value simulated under no drift so that all three sit at 5.0% exactly. The incumbent — a rank comparison of the first half of the scores against the second — reaches four-in-five power at a growth factor of 4.31. Reading each score's rank against its position reaches it at 2.65, and the largest running departure of the scores from their mean at 2.12. The ordering of the three is the ordering by how much of the sample's arrangement each one uses.

A detector built for the ordering

The best of three checks for a drifting scale fires at half the growth factor the standard one needs — 2.12 against 4.31 — and still leaves 6.50 points of coverage gone before it does, against 0.51 for serial correlation. The reversal was not a property of the test.

conformal · Exchangeability
One estimator, three answers, and only the reference changes. Coverage of the cluster-robust 95% interval for the slope against the number of clusters, at 30 rows in each. The estimator is identical in all three curves; what differs is the number it is compared against. At 5 clusters it covers 75.05% against a normal, 85.30% against a t on 4 degrees of freedom and 87.95% against a t on 3. At 80 clusters the three agree to within a point. The correction costs nothing: the same standard error, a different table.

The reference the sandwich is read against

The cluster-robust interval covers 75.05% at five clusters and 93.58% at eighty. The same estimate read against a t on G − 2 covers 87.95% at five, and the estimator is unchanged — three hundred rows grouped into five clusters cover 74.28% where the same three hundred grouped into seventy-five cover 94.63%.

sandwich · Misspecification
A proxy removes less than its reliability, always. The share of the confounding bias removed by adjusting for a proxy, against how well the proxy measures the confounder. The diagonal is the answer a reader would guess — a covariate that is 80% signal removes 80% of the problem. The curve is what the arithmetic gives: the reliability, times one minus the squared correlation between the treatment and the confounder, divided by one minus the product of those two. That squared correlation is 0.4475. A reliability of 0.8 removes 68.85% and one of 0.6 removes 45.32%. The two agree only at the ends, and the gap is widest where most applied covariates sit.

Adjusting for a shadow

A covariate that is 80% signal removes 68.85% of the confounding, not 80% — the share is λ(1 − ρ²)/(1 − λρ²) and it is below the reliability everywhere. The residual bias is 0.1084 against an effect of 0.5, and at 25,600 rows it is 17.6 standard errors wide.

collider · Conditioning
Two companions on one simulation, two hundredfold apart. How many times as many draws each companion is worth, on the same 4,000 simulated samples of 40 observations. The coverage of the interval is estimated with the observed count as its companion, whose expectation is 12 exactly; they correlate at 0.2665 and the companion is worth 1.08 times the draws. The expected width is estimated with p̂(1 − p̂) as its companion, whose expectation is 0.20475 exactly; they correlate at 0.9977 because the width is a monotone function of it, and the companion is worth 214 times the draws — 856 thousand simulated samples' worth of precision from four thousand.

The check worth more than the check

The same exactly known companion that verifies a simulation can sharpen it. On one set of four thousand draws, one companion is worth 1.08 times the draws and another is worth 214 times them, and the factor is 1 − ρ² with nothing else in it.

method · Routes
The law is the eigenvalues, and nothing else. The mean and the skewness of n(ĝ − g) at the stationary point, measured over 40,000 draws, against the closed forms ½ Σλ and 2√2 Σλ³ ⁄ (Σλ²)^(3⁄2). The worst disagreement anywhere is 0.028. In one variable the second-order law is a single χ² and its sign is the sign of g″; here it is a weighted sum with the Hessian's eigenvalues as weights, so a bowl and a valley differ in both moments and a saddle has both equal to zero.

A flat point with more than one direction

At a stationary point of a function of several means the second-order law is ½ Z′HZ, so the bias is half the Hessian's trace — 2.008 for a bowl, 5.028 for a valley, and −0.006 for a saddle, where the eigenvalues cancel. The saddle's coverage is the worst of the three.

normal · Clt
Each interval covers one question and not the other. Coverage of each interval for the overall mean, scored against both estimands, over 20,000 two-site studies of 10 observations apiece. The fixed-effect interval covers the mean of the two sites in hand 96.37% of the time and the population mean 54.77%. The random-effects interval covers the population mean 94.96% — exactly its level, from one degree of freedom — and over-covers the two sites in hand at 98.25%. Both are correct; they are answers to different questions printed in the same place.

What a two-unit study should report

The fixed-effect interval covers the mean of the two sites in hand 96.37% of the time and the population mean 54.77%. The random-effects interval covers the population mean 94.96% — exactly its level, from one degree of freedom — and is 11.6 times wider.

multilevel · Levels

Named alongside it

The objects these essays reach for when they reach for this one.

Closed formDependenceCoverageReference distributionError rateModel selectionSelection effectStatistical powerConfidence intervalCritical valueDegrees of freedomLong-run variance

All concepts