Concept

Degrees of freedom — where it appears

The number of independent contrasts behind an estimate, which sets how far a t quantile sits from a normal one. They are conserved: an observation after the first supplies one to a spread estimate or to an interval and never to both.

Named by 61 essays across 33 fields — each of them below, with the objects they name alongside it.

One experiment, with the blocks getting smaller as the target comes into range. A single run at a requirement of 0.25, with the block sizes 5, 5, 11, 25, 11, 8, 3, 2 and a total of 70 observations in 8 blocks. The rule stops when the observations in hand reach z²σ̂²/d², with σ̂² pooled from the within-block contrasts — an estimate that moves as the run goes on, so the target moves too. Early blocks are large because the target is far away and cannot be overshot; late ones are small because a block is the granularity of the answer. The interval afterwards is built from the 8 block means and from nothing the rule looked at, and it has 7 degrees of freedom against the rule's 62.

A block size that changes

The blinded rule's exactness never needed the blocks to be the same size. Letting the size be chosen from the contrasts as the run goes on leaves the coverage exactly where it was — and runs straight into an identity that says what a schedule can and cannot buy.

pace · Stopping
The profile a break point is chosen from. One sample of 120 rows under a break in the persistence, fitted as two first-order regimes at every admissible break point. The maximum is at row 78, where the true break is at 60. The shaded band is every break point within two log-likelihood units of the best one — 8 of the 73 positions searched, which is 11% of the range. The horizontal line is the one-regime fit the search is compared against; the whole profile is above it, at every position, which is the point: a maximum over 73 candidates is above the null by construction and not by evidence.

A break that was looked for

A two-regime whitening finds its change point by maximising a profile, and then reads a criterion that counts parameters. Under no break there is no parameter to count, because every position describes the same model.

charged · Break point
The correction is not a property of the sample. tr(HΩ)/q for each of fifteen candidates, at ρ = 0.7. Two candidates that fit the same number of coefficients need corrections that differ by as much as 1.49, because one of them is fitting the persistent predictors and the other is not — so no single number can be right for both, and the scalar n/n_eff = 5.537 is above every one of them. The four predictors carry persistences 0.9, 0.6, 0.3, 0; at one persistence for every column the whole spread collapses and a scalar looks exactly as good as the trace.

A penalty is a trace

Akaike's 2q is not a count of coefficients. It is the answer a trace collapses to when the rows are independent — and once they are not, the trace is still the right object and is no longer the count.

effective · Order-selection
What each variant loses before anything has been searched for. The mean loss differential of each of the eight variants against the benchmark, over 600 tables of 60 origins, with every fit given 71 rows. The series is an AR(1) and every variant adds a lag whose coefficient is zero, so in population the two forecasts are the same forecast and the difference drawn here is estimation noise and nothing else. The marked line is σ²(q₁ − q₀)/n = -0.01408, which is an expression in how many coefficients each model has and how many rows it was fitted on — it knows nothing about the series, the persistence or which lag the variant added, and every bar is within a fifth of it. This is the amount a reference distribution recentred at each column's own sample mean believes the candidates are already behind by.

A table of nested models

A benchmark and eight variants of it, each adding one thing. Every variant is behind before the search begins, by an amount that can be written down before the data exists — and the two most natural ways of reading the table are wrong in opposite directions.

search · Forecast
The construction survives a difference of two weighted means. Coverage of δ̂ ± t√(S_D²/H) on b − 1 degrees of freedom, over 900 runs at a requirement of 0.3, where δ̂ is the block differences weighted by h_b = (1/m_A + 1/m_B)⁻¹ and H is their total. The theorem the one-mean field rests on goes through with h_b in place of the block size, and the reason is that the weights a weighted least squares decomposition needs are the inverse variances — which is exactly what h_b is. The stopping rule reads only within-arm within-block contrasts, so it is a function of nothing the interval reports, whatever it does with the block sizes. Each bar is within 2.9% of the level it claims.

A width promised for a difference

The exact fixed-width interval was built for one mean. Two arms make the target 42.7 units of effective size and each unit costs four observations, so the same promise about a difference costs 169.4 rather than 42.7 — and the theorem survives untouched with the harmonic size in place of the block size.

contrast · Width
The area under the window is what the band actually costs. The three windows' weight sequences at a width of 30 lags, drawn against the lag as a share of the window. A truncated window applies a weight of one to every lag inside it and zero outside, which is why its sum is the width and why every conventional charge is right for it — and it is a covariance matrix on almost no sample, so it cannot be used. The Bartlett window falls linearly to zero and its weights sum to exactly 15.000000000000004, which is half the width, at every width: Σ(1 − k/(L+1)) over k = 1 … L is L − L/2. The Parzen window sums to 11.13 here, three eighths of the width, and it gets there by holding a weight near one over the first few lags and then falling faster. A plug-in estimate multiplied by a weight below one is a shrunk estimate, and a shrunk estimate is worth less than a free one — which is the whole of why a charge levied per lag is a charge for parameters the window has already spent.

The charge nobody derived

A band of lags is charged one log-likelihood unit apiece, because that is what a regression coefficient costs. A band's numbers are not regression coefficients, and measuring what they actually cost puts the convention out by a factor of nearly three.

dimension · Criterion
The number the comparison was missing. What it costs to choose the tuning parameter for every candidate separately rather than once for the table, under AR(1) at 0.8, paired on the draw. The window's figure is the one the earlier field reported; the order's is the one it named and did not make. They are the same size — 0.00401 against 0.00360, at 2.30 and 1.72 paired standard errors — and matching the lists at eight values leaves them the same size again. The prediction that the longer list would make the order's cost the larger of the two is not what happens; what happens is that the two rules cost the same once they are scored by the same criterion, which took a missing term to arrange.

The comparison that was not made

Choosing a whitening's window separately for every candidate costs 0.00401 of regret. The same question about an order was named and left, because the two lists are different lengths. The order's answer is 0.00360, and matching the lists changes almost nothing.

lists · Order-selection
Four runs, and the term they cannot reach. Every run sits at a corner, so x₁² and x₂² are 1 at every run and both columns are copies of the intercept. The normal matrix is singular: the design has no information about curvature at all, and no analysis can recover it.

The design that cannot see a curve

A two-level factorial has every run at a corner, where every squared term equals one — so the column that would estimate curvature is a copy of the intercept, and the design has no information about it at all. A few runs at the centre buy one number back, and only one.

surface · Factorial
The curvature is in the denominator. The optimism a Bartlett band of each width actually costs, divided by that width, on two ways of measuring the width, over 2000 draws at 120 rows. Measured in the weights the band spends — Σ w(k), which is what the earlier field levies its charges on — the reading falls from 0.9528 at two lags to 0.7486 at thirty, so a charge proportional to the summed weights is too dear at one end and too cheap at the other. Measured in the pairs the band uses — Σ w(k)(1 − k/n), because a lag of k is an average over n − k products — the same readings are flat from 4 lags up, at 0.0084 of χ² per width against 0.2359. The correction has no fitted parameter in it: it is a function of the window, the width and the sample size.

The width a band is measured in

A tapered covariance band spends 84% of its own weights at two lags and 74% at thirty. Every charge in the collection is a straight line through the origin in those weights, so it is too dear at one end and too cheap at the other.

curve · Criterion
Two searches find some of the same luck. What each search reports on a sample with no break in it, and what the two report together, on four dependences. The dashed line is the sum of the two — what a rule charging each search separately would levy — and the two together always come in below it: 19.30, 15.44, 14.24, 26.64 short, on 100%, 99%, 99%, 100% of draws. The shortfall is not a rounding. Under AR(1) at 0.8 it is 19.30 of the 34.70 the break search manufactures on its own, which is more than half of it. Two searches over one sample are looking at the same noise, and the second one has less left to find.

Two searches, one sample

A searched break in a regression manufactures 34.7 of likelihood ratio where a count of coefficients says 11.1. A searched window manufactures 84.0. The two together manufacture 99.4, not 118.7.

twice · Break point
How much of one search the other has already found. Five pairs of searches on one sample, on a scale whose zero and one are both fixed by construction. Zero is two searches over disjoint sets of independent columns: they remove shares of the residual sum that add, at 0.8 standard errors from exactly additive, and they read 0.004. One is a break search paired with a step column it contains, which reads exactly one on every draw because the step adds nothing at all. Between them: two dictionaries of step columns cut a few rows apart read 0.125, and the pair the earlier field measured — a break and a whitening window, both reading the same residual series — reads 0.762, three quarters of the way to one search containing the other. And below zero, a break paired with a search over independent columns reads -0.306: the joint search finds configurations neither half of it contains, so charging the two separately under-charges.

Two searches that share nothing

Two searches over independent columns remove shares of the residual sum that add exactly. On the scale a chi-square point is quoted on they look super-additive by a fifth of a unit, and none of it is overlap.

apart · Criterion
Exact in the corner, where nothing was. Coverage of a nominal 95% interval on five designs, at a required half-width of 0.3. The first four are the two-arm field's own and the fifth is its corner — two variances, block sizes that swing by eight, and an allocation that alternates between five to one and one to five — where neither of that field's two conditions holds. The effective-size weights over-cover there at 98.40%; the weights h_b(λ) = (1/m_A + λ/m_B)⁻¹ cover at 94.84%, and at 94.84% when λ is estimated from the within-arm contrasts rather than known. Nothing here is supposed to move.

Weights that need only a ratio

A fixed-width interval about a difference is exact under either of two conditions and under neither in the corner. It is exact there too, and the only thing it needs is how much larger one arm's variance is than the other's.

corner · Nuisance
20,000 p-values from a true null, n = 12. Flat, as it must be: under the null a p-value is uniform on (0,1). The Kolmogorov–Smirnov distance from uniform is 0.0090 (p = 0.81). That flatness is the check that catches an error a single rejection rate would miss.

A p-value that is not flat is not a p-value

Under a true null, p-values are uniform. That is stronger than saying the test rejects 5% of the time, it constrains the whole distribution rather than one point of it, and it catches implementation errors that a rejection rate sails past.

testing · Uniformity
The curvature is in the denominator. The optimism a Bartlett band of each width actually costs, divided by that width, on two ways of measuring the width, over 2000 draws at 120 rows. Measured in the weights the band spends — Σ w(k), which is what the earlier field levies its charges on — the reading falls from 0.9528 at two lags to 0.7486 at thirty, so a charge proportional to the summed weights is too dear at one end and too cheap at the other. Measured in the pairs the band uses — Σ w(k)(1 − k/n), because a lag of k is an average over n − k products — the same readings are flat from 4 lags up, at 0.0084 of χ² per width against 0.2359. The correction has no fitted parameter in it: it is a function of the window, the width and the sample size.

A lag the sample has less of

A sample autocovariance at lag k is an average over n − k products, not n. Count a band's width in the pairs it actually has and the curvature in its charge goes away, on a correction with nothing fitted in it.

curve · Criterion
What the second search finds, alone and afterwards. For four of the pairs, what the second search removes on its own and what it removes once the first has already run. The gap between the two is the overlap in absolute terms. Where the searches share nothing the two readings are the same: an independent column removes 0.0261 alone and 0.0260 afterwards. Where one contains the other they are 0.1387 and exactly zero. The pair the earlier field measured sits between: a whitening window removes 0.5033 alone and 0.3120 after a break search has run. This is the earlier field's own reading of its pair, on the share scale rather than in log-likelihood units, and it is the number a rule that runs both searches actually has to charge for.

A search that is already the other

A break search shifts every coefficient after a row, so a step column is one of the directions it can move in. Paired with a dictionary of them it reads exactly one, on every draw, and that fixes the top of the scale.

apart · Criterion
How wrong the ratio is allowed to be. λ enters only through the weights, so misstating it leaves the estimate unbiased and moves two things — the interval's calibration and its efficiency — both of which are closed forms of the design. Coverage stays at its level over a factor of two in either direction (94.93% at half the truth, 94.27% at twice it) and starts to go at a factor of five. An estimate on hundreds of within-arm degrees of freedom is never wrong by anything like that, which is what makes the feasible rule usable rather than merely definable.

Blinded, and still exact

The one number the exact interval needs is a ratio of within-arm spreads, which is a contrast and contains no mean — so a rule forbidden to look at the effect may compute it, on more degrees of freedom than the interval itself has.

corner · Width
The one candidate an effective sample size is right about. n/n_eff with the finite-sample inflation Σ(1 − |k|/n)ρ^|k| is not an approximation to tr(HΩ) for a fit with only an intercept — it is that trace, to machine precision, because the hat matrix of a constant column is 1/n everywhere and its trace against Ω is the mean of Ω. The quoted limit form n(1 − ρ)/(1 + ρ) is not even right about that one. And the average correction the table's fifteen candidates actually need is 3.318 per parameter, well below the scalar, so applying it to all of them over-charges every one.

One number for a table of candidates

An effective sample size is a real quantity, it is exactly right about one thing, and that thing is a mean. Substituted into Akaike's criterion it changes nothing at all, because the penalty it is meant to fix has no sample size in it.

effective · Dependence
Every candidate is behind by what its parameter count says. Each dot is one of the fifteen subsets of four predictors, fitted on a rolling window of 80 rows and scored against the benchmark out of sample over 60 origins, at a null where every one of them contains the truth. The line is σ²(q₀/(R − q₀ − 1) − q/(R − q − 1)), which is arithmetic on two integers and a window length. Most of these pairs are not nested — a subset of two predictors and a different subset of two share neither model — and the closed form does not care: the displacement is a statement about how many coefficients each side estimates. The candidates of the benchmark's own dimension sit at zero.

The displacement is a parameter count

A nested variant is behind its benchmark out of sample before anything is searched for. The closed form for how far turns out to have nothing about nesting in it — only two integers and a window length — and it prices a table where no candidate contains any other.

select · Multiplicity
Most of the rise is the optimiser's, and under one law it is not. The rise in log-likelihood from the tapered plug-in to the maximum over the same eight-lag band, beside what the same optimiser produces on a sample generated from the plug-in's own covariance — where the family is correctly specified by construction and there is nothing to find. Under AR(1) at 0.8 the raw rise is 5.72 and the manufactured baseline is 4.79, leaving 0.93 at 1.8 standard errors; under long memory the excess is 0.14, at 0.2. Under the moving average it is 11.87 at 19.4 standard errors, on every draw. The taper is a shrinkage, and it costs nothing where the sequence decays smoothly and a great deal where it stops dead.

The plug-in and the maximum

A tapered covariance estimate sits five and a half log-likelihood units below the maximum of the likelihood it is substituted into. Four fifths of that is what the optimiser would have found if nothing were missing.

family · Dependence
The term that cancels, and the term that does not. The volume each candidate's whitening moves — log|Ω̂| — for a sieve of order 4 on one sample of 120 rows. Estimated once from the fullest candidate and used for the whole table, it is the same number for every candidate, so it drops out of every difference the criterion reads: that is why nothing in this collection has ever needed to carry it. Estimated from each candidate's own residuals it ranges over 23.87, which is more than a parameter is worth, and the criteria being compared are then fits made under different error models with no term saying so. The window's rule has carried this term since the estimated-covariance field and the sieve's never had it.

The volume a whitening moves

A sieve's whitening has a determinant and this collection's criterion for it never carried one. Shared across a table the term cancels exactly, which is why nothing ever noticed; used per candidate it is worth more than a parameter and the whole comparison turns on it.

lists · Criterion
A window with two wrong ends. The regret of a rule whitened by a Bartlett-tapered Ω̂, as the window widens, against three rules that need no window at all. At L = 0 the estimate is the identity and the rule is exactly least squares — 0.08556, the same number to five places. It falls to 0.01754 at L = 20 and rises again by L = 30, because a quarter of the sample's lags are then being estimated from it. The automatic bandwidth a practitioner would reach for, 4(n/100) to the power 2/9, which at this sample size is 4, gives 0.02963 — 69% above the best available window. The rule told the dependence is an AR(1) sits at 0.01440 throughout, which is the price of not knowing the form.

The window that has to be chosen, and the term that was dropped

An estimated covariance has a bandwidth in it, and both ends of the dial are wrong for different reasons. The rule a practitioner would reach for is two thirds worse than the best window there is.

banded · Order-selection
No block size is best at both things the procedure claims. Two claims and one dial. The honest interval's half-width falls as the blocks get smaller, because the interval's degrees of freedom are the number of blocks: 0.2602 at blocks of two against 0.2933 at blocks of sixteen. The fixed-width claim — that the mean is within 0.25 of the truth — gets more reliable as they get larger, because the sample size is less variable: 92.40% against 95.00%. Both are computed from the same runs, and the second is reproduced to within a tenth of a point by E[2Φ(d√N/σ) − 1], which needs the sample-size distribution and nothing else. The schedules sit at the bottom left: as narrow as the smallest fixed block and as few observations, with the rule's spread estimate on half as many degrees of freedom again.

Two degrees of freedom, one total

The block size is a dial, and the two things a fixed-width procedure claims move in opposite directions along it. Divide the width by the square root of the sample size and one of them turns out to depend on the number of blocks and on nothing else.

pace · Width
Coverage of four nominal 95% intervals, n = 30. Computed exactly by summing over all 31 possible counts, not simulated. The Wald interval drops to 26.0% and is jagged everywhere; Clopper–Pearson never falls below 95% and pays for it in width.

Two routes to every number

A site about probability that only simulates has one route to each answer and no way to tell a right one from a plausible one. Every important number here is computed twice, by arithmetic that shares nothing, and the two are required to agree.

method · Routes
What a search manufactures, law by law. The average likelihood ratio a search over 120 rows reports, on each of the four laws, over 400 draws. Three of them have no break at all and report 5.080, 4.839 and 4.748; the fourth has one and reports 8.757, so the real break is worth only 3.677 beyond what the search would have found anyway. The dashed line is 2, which is what an information criterion charges for one extra parameter. A search costs about two and a half of them, and the number is a measurement rather than a count.

What a search costs in parameters

An information criterion's penalty is an estimate of the optimism a fit carries. For a break point the optimism can be measured and cannot be counted, and it comes to about two and a half parameters.

charged · Break point
Four windows, one line, and one that is off it. The optimism measured for each window at a band of 30 lags, against what that window's weights sum to, on 2000 pairs of independent samples of 120 rows. The diagonal is where a window that spent exactly its summed weights would sit. Three of the points are one shape at three levels — the Bartlett window, its square and its cube, whose sums stand in the ratio 6 : 4 : 3 — and they lie on a line through the origin at 0.767 of the diagonal, with 0.033 between the highest and the lowest. Scaling the weights scales the charge by the factor the weights predict, which is what makes the weights the mechanism. The Parzen window has a comparable sum and a different shape, and it sits at 0.871: its weights stay near one over the first few lags, and the first few lags are where the information is. A weight sum treats every lag as equally informative and no sample does.

What a window leaves free

A Bartlett window's weights sum to exactly half its width, which is a candidate for what the band costs. Varying the weights without varying anything else says the weights are the mechanism; varying the shape at the same weight says they are not the arithmetic.

dimension · Criterion
Three sets of weights, five designs, and no estimator that is exact everywhere. Coverage of the same interval under three weightings. h_b is the inverse variance when the arms share a variance or the allocation is constant; equal weights are right when every block has the same two counts; the estimated precision weights are right in the limit and exact nowhere, because the decomposition needs the weights to be the constants they are only estimating. In the corner — two variances, changing sizes, changing allocation — the two exact estimators are the ones that miss, at 98.45% and 95.65%, and the one with no theorem behind it is at 95.05%. That is the whole statement: there is an exact estimator under either condition, and none under both.

Which weights are the inverse variances

There is an exact estimator when the two arms share a variance and another when every block has the same two counts, and between them they cover every trial anybody designs on purpose. In the corner where neither holds, both cover 98.45% instead of 95%, and the only estimator at its level is the one with no theorem behind it.

contrast · Nuisance
The pooled two-sample test's size, with a true null everywhere. Forty units split between two groups, with the second group's variance a stated multiple of the first's, and the two population means equal. A 5% test should reject 5% of the time. The pooled test runs from 0.55% to 18.91% across this region; Welch's runs from 4.63% to 5.51%.

A degrees of freedom that is not a count

The pooled two-sample test's size runs from 0.55% to 18.91% across forty units split five ways against five variance ratios, with a true null in every cell. Welch's runs from 4.63% to 5.51% — bought with a degrees of freedom that is a function of the data, not an integer, and not a count of anything.

tails · Student
A line in the right width beats two curves. How far each candidate charge sits from the measured optimism across the plateau, in units of each width's own standard error, over 2000 draws. The straight line through the origin in the band's summed weights — which is what the earlier field levies — misses by 0.2382 per width. The same straight line in the pairs the band actually uses, Σ w(k)(1 − k/n), misses by 0.0095. A fitted power law misses by 0.0293 and a fitted decaying rate by 0.0172, both on one fitted constant more. The deferral this field answers asked for a curve; the answer is a line, in a variable with nothing fitted in it.

A line that beats two curves

A deferral asked for a curve. Fitted against the same measurements, a straight line in a variable nobody had to fit describes the plateau better than either curve does with a constant more — and for three windows out of four it does not.

curve · Criterion
Cut the charge and the width follows it. The band width each charge picks, averaged over 400 draws of 120 rows under AR(1) at 0.8, with the standard deviation across draws beside it. Schwarz's charge — half a log n a lag, which is 2.39 here — picks 3.67. Akaike's picks 6.02. Charging the numbers the window actually leaves free, which is half the width, picks 10.12; charging what the optimism measures, 0.767 of that, picks 14.15. A charge and the width it buys are very nearly reciprocal, which is what a likelihood rising at a fixed rate a lag implies and is why the four answers span a factor of 3.86. The width that was actually best on the draw averages 13.90 and moves by 10.30 from draw to draw — three times as much as any rule's answer does.

A width that moves and an error that does not

Four charges give four widths a factor of four apart and four errors half a per cent apart. The derived charge wins, significantly, by a quarter of what was on offer — and none of the four is an estimate of anything.

dimension · Criterion
What each criterion selects, at 50 observations. 700 series from an AR(2) with coefficients 0.6 and -0.3, every order from 0 to 8 fitted to the same 42 responses so the log-likelihoods are comparable. AIC finds the true order 55.1% of the time and lands above it 25.7%; BIC finds it 54.1% and lands above it 4.0%. The closed form for one extra lag is P(χ²₁ > 2) = 15.73% for AIC, which does not depend on n at all, and P(χ²₁ > ln n) = 4.79% for BIC at this size, which falls to zero. Under the true order is the other failure and it is BIC's: 41.9% against 19.1%.

Choosing the order

One criterion is consistent and one is not, which is the whole of what gets said about them. At two hundred observations the consistent one is right 95% of the time and the other 70%; at fifty they are both right 54% of the time and wrong in opposite directions, and consistency has not started to mean anything yet.

forecast · Order-selection
A wider band is always a better fit. The likelihood maximised over the band, at five widths, averaged over 30 samples. A band at L lags is a band at L + 1 with the last entry held at zero, so the families are nested and the maximised likelihood cannot fall — it does not, on any draw. What it does is rise at 0.984 of log-likelihood a lag. A parameter that is doing nothing buys half a unit in expectation and Akaike's criterion charges one, so this is a criterion very nearly indifferent between every width on offer. The dashed line is what a charge of one unit a lag would exactly cancel. Nothing in the fit chooses a width, and what does choose one is a charge somebody has to pick.

Nothing in the fit picks the width

A wider band is always a better fit, and it is better by about one unit of log-likelihood a lag — which is the order of what a criterion charges for a parameter. Three defensible rules choose widths a factor of three apart.

family · Order-selection
Four analyses of the same 3-arm trials, under a true null. 250 trials of 150 patients, 3 arms, minimisation with p = 0.85, 99 re-randomisations for each exact test. Two statistics — an F on the arms alone and an F on the arms after the balanced factors — against two reference distributions: the table the statistic is named for, and the distribution the allocation rule itself generates when the outcomes are held fixed and the rule is re-run. Only the first cell is wrong, and it is wrong in the direction that costs power rather than the one that manufactures findings: 0.0% where 5% is claimed. Either repair works — adjusting for what the rule balanced, or asking the rule what it would have done.

The analysis after three arms

An unadjusted analysis after a two-arm balancing rule rejects 0.6% of true nulls where it claims 5%. With three arms and a deterministic rule it rejects none at all — and the repair is the same repair, which is a sentence and a column in the model.

multiarm · Assignment
The weights may not read the block they weight. A weighted least squares decomposition needs weights that are constants, or at least independent of the differences they multiply. One λ̂ pooled across the trial is estimated on hundreds of degrees of freedom and is effectively a constant; a λ̂ estimated inside each block is estimated on that block's own two or three, and is correlated with the difference it weights. Coverage falls from 94.68% to 82.76% — and the interval gets wider while doing it, 0.5163 against 0.3024, which is the signature of weights that are noise.

The condition that cannot be dropped

The weights may not read the block they weight. Estimate the variance ratio inside each block rather than across the trial and the coverage falls to 83% — on an interval that is at the same time seventy per cent wider.

corner · Allocation
Two arms leave one degree of freedom per block unaccounted for. Each point is one run. The one-mean field's identity is (b − 1) + (N − b) = N − 1, and every schedule moves along that line rather than off it. Two arms give the rule N − 2b and the interval b − 1, which come to N − b − 1 — short of the N − 2 two arms leave by exactly one per block, since a block's arm counts absorb one degree of freedom each and only one of the two directions carries the difference. The hollow points add what the block sums are worth, b − 1 more, and land on the total. The missing degrees of freedom are not lost; they are in a place the interval has to be shown it may read.

The degrees of freedom in the sums

One arm partitions N − 1 exactly. Two arms give the rule N − 2b and the interval b − 1, which is short by one per block — and the missing ones are in the block sums, which are correlated with the differences at −0.79 and are usable anyway.

contrast · Blocking
Exact coverage, at every block size. Coverage of the interval each rule reports, at a nominal 95%, over 2,500 runs each with a standard error of 0.44 points. The blinded rule stops on the within-block contrasts and reports an interval built from the block means, and those two are independent whatever the rule does — so the interval is an ordinary t interval on b − 1 degrees of freedom and its coverage is exact. It is exact at every block size drawn. The interval a practitioner writes at the purely sequential rule's stopping time covers 91.72%, and Stein's two-stage rule is exact for the same reason as the blinded rule and spends 2.10 times the observations to be so. The bars are truncated at 86% so the differences can be seen.

The rule that cannot see the mean

A sequential rule stops when its own estimate of the spread is small, which is more often on the samples whose spread came out low — so the interval afterwards is short. There is a way to keep updating the estimate and stop being able to see the mean at all.

blind · Stopping
The ladder is the same ladder under every law. Each pair's overlap under each of the four laws, over 400 draws apiece. The scale is fixed at both ends by construction: two searches over disjoint sets of independent columns read -0.001, 0.013, -0.002, 0.004, and a break search paired with a step column it already contains reads exactly one under every law. Between them the pair that reads one residual series twice runs 0.763, 0.805, 0.577, 0.752 — lowest under long memory, where a whitening has most to do and the break search has least left to find that the whitening has not taken. The rung that moves most is the pair of step dictionaries, from -0.148 under a moving average to 0.419 under a break; and the pair that is negative is negative under all four.

Three quarters of the way to one search

The pair that started this reads 0.762 on a scale whose one is containment. And the pair that shares nothing but its response reads −0.306, so the sign the earlier field found does not transport at all.

apart · Criterion
The overshoot is the last block size and nothing else. A run stops at a multiple of its own block sizes and cannot land between them, so it ends past its own target by about half a block. Fixed sizes overshoot by 1.5, 2.2, 3.1, 4.7, 8.5 observations as the size goes 2, 3, 5, 8, 16. Every schedule here ends in blocks of two and every one of them lands where blocks of two land — 1.62, 1.32, 1.37 against 1.48 — while having spent most of the run inside blocks four and eight times larger. That is the one thing on this page a schedule genuinely takes from both ends.

What a schedule actually buys

Big blocks early and small blocks late is the right instinct and it does not take both ends of the trade, because there are not two ends to take. What it does take is the overshoot — about four per cent of the observations — and a steadier stopping point.

pace · Nuisance
A filled value is not an observation. What a 95% interval for the slope actually covers after each way of handling 35.0% missing outcomes, counted over 4000 studies of 200 rows. Dropping the incomplete rows covers 95.93%. Filling with the observed mean covers 13.85%, because the estimate itself has moved. Filling with a fitted value covers 80.85% against a closed prediction of 79.73%: the estimate is right and the reported standard error is short by a factor of 0.6567 against a predicted 0.6500, because the residual sum of squares is divided by the whole sample's degrees of freedom. Adding residual noise recovers the spread and covers 85.78% against a predicted 84.62%, since the interval still ignores the variance of having imputed at all.

One imputation is not an observation

Three ways of filling a missing outcome, under a mechanism that makes dropping the rows beyond reproach. Filling with the observed mean covers 13.85%, filling with a fitted value covers 80.85%, adding noise covers 85.78%, and the thing all three were meant to improve on covers 95.93%.

missing · Missingness
A wrong weight costs width; a random weight costs level. Five weightings on a trial whose variance ratio drifts by a factor of 20.1 between the first block and the last, over 4000 runs. The rule that knows every λ_b covers at 95.1% and sets the width. One ratio for the whole trial is wrong for every block and costs nothing in level — 94.8% — while being 20% wider; equal weights are calibrated by an identity and 22% wider. The ratio estimated inside each block is the only rule aimed at the quantity that actually varies, and it is the only one that misses the level, at 92.0%: a weight computed from a handful of degrees of freedom is mostly noise, and noise in a weight is not a wrong weight. Modelling the drift across blocks recovers the oracle's width at 94.8%.

A ratio that changes between blocks

A wrong weight costs width and a random weight costs level. The rule aimed at the quantity that actually varies is the only one that misses its own coverage, and the rule that models it across blocks recovers the whole of what knowing it is worth.

blocks · Nuisance
What a schedule is allowed to read, and what happens when it reads more. The construction allows the block sizes to be anything at all as long as they are functions of the within-block contrasts, which are independent of every block mean. A schedule that shrinks the block whenever the between-block spread is running above what the contrasts say is a direct attempt to hold down the quantity the interval will be built from, and it succeeds: the estimate lands at 0.8373σ² against the honest 0.9831, and the coverage goes with it. Reading the running mean instead pushes the other way and over-covers — which is not a repair, it is the same violation with the sign reversed, and the level is no longer a property of the procedure at all.

A schedule that reads the mean

The block sizes may be anything at all provided they are functions of the contrasts. Two natural schedules break that, in opposite directions — and the most natural mistake of the three is not a schedule at all but a stopping rule, at 86.87% coverage and fewer observations.

pace · Stopping
What a second break adds. Over 200 draws, the likelihood ratio a search over one break point reports, and how much more a search over an ordered pair adds on top of it. Under AR(1) at 0.8, which has no break at all, the first search manufactures 5.697 and the second adds 4.278. Under a law with exactly one break — where a second one is as absent as the first was in the row above — the first search reports 9.442 and the second still adds 5.800. Searching for something that is not there costs the same whether or not something else was there to find.

A second break on a flat profile

Searching a hundred and twenty rows for one change point manufactures five units of likelihood. Searching for a second manufactures four more, on a series that has at most one — and on a profile whose whole range is under seven.

charged · Break point
What balancing several numbers at once costs each of them. The criterion generalises without a word changing — the covariate imbalance becomes a vector and the correction a quadratic form — so the question is what it is worth rather than whether it can be done. At n = 200 with 200 trials per point, a rule balancing one covariate leaves 12.7% of a coin's imbalance in it; balancing eight leaves 23.2% in each. The assignment has a fixed amount of freedom and every covariate added takes a share of it. The rule degrades rather than failing: at eight covariates it is still four times better balanced than a coin, and the eight are being held simultaneously rather than in turn.

Balancing more than one number

The criterion generalises to several covariates without a word changing, which makes the question what it is worth rather than whether it can be done. Each one added takes a share of the assignment's freedom, and the imbalance left in every one of them rises.

continuous · Blocking
The normal density at sigma = 1.00. The bands hold 68.27%, 95.45%, 99.73% of the mass. Those figures are integrals of the curve drawn, not the memorised 68-95-99.7.

The shape, and where its mass is

68, 95, 99.7 is recited more often than any other set of numbers in the subject. They are integrals of a specific curve, they are worth computing rather than remembering, and the third one is the one people misuse.

normal · Bands
The scale moves the width; the curve does not. The band width each charge picks, averaged over 150 draws of 120 rows. The two conventions — a unit a lag and half a log n a lag — pick 6.08 and 3.65 lags. The four charges derived from the measured optimism pick 15.05, 15.60, 15.47 and 15.44, against a best width on the draw of 14.13. So the scale a charge is levied on moves the width by a factor of 4.27 and the shape of the charge moves it by 3.6%. None of the six is an estimate of the draw's own best width: the correlations are -0.006, -0.003, 0.017, -0.001, 0.016, -0.004.

What a better charge buys

Four charges derived from the same measurements pick band widths within six per cent of each other and deliver errors within two per cent of the gap any of them leaves. The scale a charge is levied on decides the width; the shape of the charge decides nothing.

curve · Criterion
What the interval covers, after a design that read the data. 800 experiments of 12 runs, all at the same truth. The first pair is a design fixed in advance; the second is one whose settings were chosen from the first stage's own outcomes. If choosing the design from the data broke the inference, the second pair would sit below the first, and it does not — 91.3% against 92.6%. What does move the coverage is the shape of the interval rather than the design: the Wald interval assumes the estimate is normal around its own standard error and falls short under both designs, and the profile interval, computed from the same residual sums of squares with no derivative in it, covers 94.3% and 94.9%.

What a design chosen from the data costs

Two fields on this site measured what happens when a rule reads the data, and the error rate broke both times. A design that reads the data to decide where to put its runs breaks nothing — and the control that proves it also finds what the real shortfall is.

robust · Local design
Free until the sums stop seeing what the differences see. Coverage with and without the block sums pooled into the interval's variance estimate. With one effect and one level they are free. With an effect that varies between blocks they are still free, because a block's sum picks that variation up exactly as its difference does. With a level that varies they make the interval 37% wider and conservative. And where the effect falls as the level rises — a ceiling, and not an exotic thing to suppose — the sums carry none of the between-block variation while the differences carry all of it, the pooled estimate is short, and the interval that uses it covers 88.75% on a width 20% narrower than the honest one.

What a two-arm rule may not pool

A spread computed "within the block" without the arm label carries a share of the effect, so the trial runs 173 observations at a null and 282 at an effect of 1.5. The stopping rule is reading the thing it exists to measure, and the phrase that produced it is one word long.

contrast · Allocation
What the exactness costs, and the dial it is bought with. The median half-width of the interval each rule reports, at a requirement of 0.4 and a first look after 5 observations. The flat line is the interval a practitioner writes at the purely sequential rule's stopping time, which covers 91.72% rather than 95%. The curve is the blinded rule, which covers its nominal level at every block size: it reads b − 1 degrees of freedom where the other reads n − 1, and pays for the exactness in width. The best block size is 3, at 0.4712. Larger blocks give the stopping rule a better estimate and the interval a worse one, and the two costs go opposite ways, which is what puts the minimum in the middle.

What the blindfold costs

The exactly-covering rule pays for it in the width of the interval, and the block size is a dial between two costs that run in opposite directions. And on an interval whose width was fixed in advance, the same repair buys nothing at all.

blind · Nuisance
One penalty, read along two dials. How much wider the studentised interval is than the percentile one, at every sample size and every block length on the grid, with the number of whole blocks each cell leaves written beneath. Read across a row and the block length changes; read down a column and the sample size does. The penalty is nearly a function of the block count alone: the cells at 15 blocks read 1.16, 1.20, 1.17, 1.15, 1.13, 1.10 across three sample sizes and three block lengths, while the cells at one block length read anything from 1.10 to 3.95. The largest penalty on the grid is 3.95, at the cell with 3 whole blocks in it.

The count or the length

A block length and a block count are one number read two ways at one sample size. Read at three, the studentised interval's width penalty tracks the count — with an R² of 0.9911 against a closed form that has no length in it — and its coverage tracks the length.

student · Bootstrap
The extra 1/m, and the correction nobody quotes. What a pooled 95% interval covers against the number of imputations, counted over 2000 studies of 200 rows at 35.0% of outcomes missing. Rubin's rules — total variance W̄ + (1 + 1/m)B, read against a t distribution on (m − 1)(1 + W̄/((1 + 1/m)B))² degrees of freedom — cover 94.10% at two imputations and reach their promise by 5, at 95.25%. Dropping the (1 + 1/m) factor takes two imputations to 93.10%; using a normal quantile instead of the degrees-of-freedom correction takes it to 92.55%; dropping both takes it to 91.45%. The median degrees of freedom at two imputations is 12.95, which is why the second correction is the larger.

The variance between imputations

Pooling several filled datasets covers 94.10% at two imputations and reaches its promise at five, where a single fill covered 85.78%. The correction everybody quotes is the smaller of the two doing the work — 1.00 ± 0.22 points against 1.55 ± 0.28.

missing · Missingness
What each correction is worth, exactly. Each variance estimate's expectation under a constant error variance, divided by the variance the slope actually has, at 20 rows on an even design, by two routes: the closed form E[eᵢ²] = σ²(1 − hᵢᵢ) carried through each correction's own weight, and the mean of 20000 counted estimates. The maximum leverage here is 0.1857 and the design's fourth-moment share Σu⁴/(Σu²)² is 0.0897, which is the only thing the closed form reads. HC0 comes out at 0.8603 — short by construction, since its factor is exactly 1 − 1/n − Σu⁴/(Σu²)². HC1 reaches 0.9559, HC2 is exactly 1.0000 at every design and every sample size, and HC3 overshoots to 1.1647. On an even design the four are within a fifth of each other and the choice barely matters.

Three corrections and a leverage

On an even design of twenty rows the four robust corrections read 0.8603, 0.9559, 1.0000 and 1.1647 of the truth and the choice barely matters. Add one point at x = 8 and they read 0.3191, 0.3419, 1.0000 and 5.1127.

sandwich · Misspecification
Where the bias lands. The drift in the log variance ratio, fitted across 12 blocks over 4000 trials. E[log λ̂_b] is log λ_b plus ψ(k_B/2) − log(k_B/2) − ψ(k_A/2) + log(k_A/2), which depends on nothing but the degrees of freedom — so the tempting sentence is that it goes into the intercept and leaves the slope alone. It does not, because the blocks alternate between allocations and the alternation is correlated with the covariate being fitted: the lopsided blocks carry 0.5383 of bias and the even ones carry none. Uncorrected the slope reads 1.5597 against a truth of 1.5, which is 8.0 standard errors. Subtracting the two digammas block by block leaves 1.4976.

The bias that lands in the slope

The bias in a log variance estimate depends on nothing but its degrees of freedom, so it goes into the intercept — unless the degrees of freedom alternate with the design, which is exactly what a block-randomised trial makes them do.

blocks · Width
Largest where least is needed. What the pairs correction supplies against what each window's measured profile needs, across this field's plateau, over 2000 draws at 120 rows. Both are stated as the multiplicative rise the charge per unit of width has to take between four lags and thirty. What the correction supplies is arithmetic — (1 − μ(4)/n)/(1 − μ(30)/n), where μ is the mean lag of the weight the band adds — and it runs 1.0795, 1.0580, 1.0456, 1.0539 for the four windows. What the measurement needs runs 1.1076, 1.2928, 1.2296, 1.6550. The two orderings are opposite: the plain Bartlett window has the longest mean lag, so it gets the biggest correction, and the flattest profile, so it needs the smallest. They coincide to 0.9746 of each other, and nowhere else does the correction account for more than 85.0% of the fall.

What the correction assumes

A correction with nothing fitted in it repairs one window of four. The reason is that its size is set by where a window puts its weight and the curvature it must repair is set by something else — and for one window at one sample size the two happen to agree.

curve · Criterion
Four intervals at 3 blocks of 32 rows. What four 95% intervals for the mean of a first-order autoregression at 0.7 cover, and how wide they are on average, at 120 rows cut into 3 whole blocks of 32, over 240 draws with 200 resamples each under the rectangle. The normal interval, the block-means variance with 1.96, covers 82.1% at a width of 0.708. The percentile interval covers 82.1% at 0.632 and the studentised one 90.4% at 2.495. The fourth resamples nothing: it is the normal interval with 1.96 replaced by Student's t on 2 degrees of freedom, and it covers 94.2% at 1.555, 0.62 times the studentised interval's width.

The interval with no resampling in it

Replace 1.96 in a normal interval on the block-means variance with Student's t on one fewer degrees of freedom than there are whole blocks, and resample nothing. Across twenty-four cells it covers at least as often as the studentised bootstrap interval at every one, by 0.42 to 10.42 points; it is narrower wherever seven blocks or fewer are left; and at fifteen blocks of 32 it covers 95.0%, which no resampled interval on the grid reaches.

student · Bootstrap
What each estimator costs, against how many groups there are. Each point is 4,000 datasets, with the population spread estimated from the data rather than supplied. Partial pooling first beats BOTH of the estimators it sits between at 5 groups; below that, complete pooling — which estimates nothing at all — is the better answer. Its own cost falls from 1.751 at 2 groups to 0.795 at 40.

The fewest groups that can borrow

At three groups the estimator that shrinks towards its own data's mean returns the group means untouched, on every dataset, because its constant is J − 3. At two it expands instead of shrinking. And the number of groups at which partial pooling starts to be worth doing is five, or two, or never — it depends on how far apart the groups are.

hierarchical · Pooling
What visiting fewer settings costs, 6 parameters. Carathéodory's bound puts the support of an optimal measure between 6 and 21. 6 settings: D-efficiency 88.90%, G-efficiency 57.18%, 0 degrees of freedom for lack of fit; 7 settings: D-efficiency 94.54%, G-efficiency 61.22%, 1 degrees of freedom for lack of fit; 8 settings: D-efficiency 95.99%, G-efficiency 64.60%, 2 degrees of freedom for lack of fit; 9 settings: D-efficiency 97.40%, G-efficiency 82.76%, 3 degrees of freedom for lack of fit. The saturated design has none, and buying the first one costs about five points of efficiency to get back.

How many places a design goes

Carathéodory's bound puts an optimal design's support between six and twenty-one settings, and every design in this field that can fit the model visits exactly nine. The count is not a choice anybody makes, it decides how many degrees of freedom are left to check the model with, and the first spare setting costs six points of efficiency to get back.

optimality · Equivalence
Student's t on 5 degrees of freedom, against the normal. The two-sided 95% critical value is 2.571 for t(5) and 1.960 for the normal — 31% wider. Using the normal at this sample size makes every interval too short by that much.

The correction for not knowing the spread

The t distribution exists because the standard deviation is estimated rather than known. At eight observations, using the normal instead makes every interval 12% too short — and the coverage that follows can be measured rather than argued about.

intervals · Student
The rows are held fixed; only the clusters move. Counted coverage of four 95% intervals for a slope, at five cluster counts with the row count held at 300 throughout and a within-cluster correlation of 0.1, over 6000 draws apiece, with the sizes equal. The interval that counts rows covers 53.42% at 5 clusters — a second closed form says 2Φ(z/√D) − 1 = 54.44% for a design effect of 6.900, and reads nothing about clusters at all. The cluster-robust interval read against a normal covers 74.43% there and 94.20% at 100 clusters; read against a t on G − 1 it covers 85.08% and 94.47%. The number of independent things is the cluster count, and every quantity here is blind to how many rows were typed.

The count that is not the rows

Three hundred rows in five clusters of sixty carry 6.9000 times the variance an independent-rows calculation reports, and the interval that counts rows covers 53.42%. The same five unequal sizes laid out two ways give design effects of 9.3158 and 5.4652.

sandwich · Misspecification
What a variance estimated from K units is worth. The between-unit mean square is a scaled chi-square on K − 1 degrees of freedom, so the estimator's whole distribution is decided by the number of units. At two units its interquartile range spans a factor of 13.03 and its ten-to-ninety range a factor of 171.3, and it comes out exactly zero on 26.7% of studies. The closed form and 3,000 simulated studies agree to 0.051 at every quantile.

A level with two units

A variance estimated from two units is a scaled chi-square on one degree of freedom. Its interquartile range spans a factor of thirteen, its ten-to-ninety range a factor of a hundred and seventy-one, and it comes out exactly zero on 26.7% of studies — so the design effect it decides runs from 1.00 to 7.01 against a truth of 4.69.

multilevel · Levels
What one lost run costs a 16-run factorial fitting 11 coefficients. Every run is worth the same: dropping any one multiplies every coefficient's variance by 1.2000, which is 1 + 1/(N − p) with N = 16 and p = 11, and gives every pair of coefficients a correlation of 0.1667 where the complete design had exactly zero.

The run that did not happen

Lose one run from any orthogonal design and every coefficient's variance is multiplied by exactly 1 + 1/(N − p), and every pair of coefficients acquires a correlation of exactly 1/(N − p + 1) where there was none. The price is set by the design's spare capacity and by nothing else, and a saturated design cannot survive it at all.

design · Factorial
One estimator, three answers, and only the reference changes. Coverage of the cluster-robust 95% interval for the slope against the number of clusters, at 30 rows in each. The estimator is identical in all three curves; what differs is the number it is compared against. At 5 clusters it covers 75.05% against a normal, 85.30% against a t on 4 degrees of freedom and 87.95% against a t on 3. At 80 clusters the three agree to within a point. The correction costs nothing: the same standard error, a different table.

The reference the sandwich is read against

The cluster-robust interval covers 75.05% at five clusters and 93.58% at eighty. The same estimate read against a t on G − 2 covers 87.95% at five, and the estimator is unchanged — three hundred rows grouped into five clusters cover 74.28% where the same three hundred grouped into seventy-five cover 94.63%.

sandwich · Misspecification
Each interval covers one question and not the other. Coverage of each interval for the overall mean, scored against both estimands, over 20,000 two-site studies of 10 observations apiece. The fixed-effect interval covers the mean of the two sites in hand 96.37% of the time and the population mean 54.77%. The random-effects interval covers the population mean 94.96% — exactly its level, from one degree of freedom — and over-covers the two sites in hand at 98.25%. Both are correct; they are answers to different questions printed in the same place.

What a two-unit study should report

The fixed-effect interval covers the mean of the two sites in hand 96.37% of the time and the population mean 54.77%. The random-effects interval covers the population mean 94.96% — exactly its level, from one degree of freedom — and is 11.6 times wider.

multilevel · Levels

Named alongside it

The objects these essays reach for when they reach for this one.

CoverageMonte CarloInformation criterionModel selectionClosed formDependenceFixed-width intervalNuisance parameterConfidence intervalBlindingTaperingLeast squares

All concepts