Concept

Least squares — where it appears

Choosing coefficients to minimise the sum of squared residuals, which for a linear model has a closed form. Its in-sample fit is optimistic by a factor of (n − q)/n and its risk on a fresh row inflated by 1 + q/(n − q − 1), both exactly.

Named by 41 essays across 21 fields — each of them below, with the objects they name alongside it.

The price of each thing the rule is not told. What each rule gives up against the best model available, at a persistence of 0.85 on a fifteen-candidate table, over 400 draws. Reading down: least squares with the ordinary penalty; the whitening at the true ρ; the same at a ρ̂ estimated per candidate; that rule with the term the Gaussian likelihood carries and it omits; a Bartlett-tapered Ω̂ estimated once from the fullest candidate at L = 8; the same estimated per candidate; and the truncated Ω̂, which exists on only 45.0% of draws and is averaged over those. Knowing ρ recovers 89.9% of what counting rows gives up, estimating it 83.4%, and estimating a whole covariance 75.6%.

A covariance with no parameter in it

The whitening that repairs a criterion is told the dependence is a first-order autoregression and left to find one number. A real dependence is not one number, and the obvious estimate of it is not a covariance matrix.

banded · Dependence
Where a D-optimal design puts its runs. The D-optimal measure over 121 candidate settings on a square region. It keeps 9 of them and discards the rest, and the 9 it keeps are the settings a catalogue would have offered without any of this arithmetic. What the search adds is the weights: 0.1458, 0.0962, 0.0802, which nine equal runs cannot express.

A design is a number

A standard design is taken from a catalogue and then measured. Turn the arithmetic round and a design becomes the answer to an optimisation — and over 121 candidate settings the search keeps nine of them, which are exactly the nine a catalogue would have offered, at weights nine equal runs cannot express.

optimality · Criterion
The correction is not a property of the sample. tr(HΩ)/q for each of fifteen candidates, at ρ = 0.7. Two candidates that fit the same number of coefficients need corrections that differ by as much as 1.49, because one of them is fitting the persistent predictors and the other is not — so no single number can be right for both, and the scalar n/n_eff = 5.537 is above every one of them. The four predictors carry persistences 0.9, 0.6, 0.3, 0; at one persistence for every column the whole spread collapses and a scalar looks exactly as good as the trace.

A penalty is a trace

Akaike's 2q is not a count of coefficients. It is the answer a trace collapses to when the rows are independent — and once they are not, the trace is still the right object and is no longer the count.

effective · Order-selection
What is left of a probe after the rule has had it. The share of each dictionary function a rule balancing x, x2, x3, cut0 has already taken, on trials of 14 units, averaged over 100 designs. Four of the eight functions are the basis, so their share is exactly one: a randomisation test run on one of them is asking about a quantity the rule forced to zero, and one of them is the default probe of the field this measurement comes from. The four that are not still read 0.919, 0.873, 0.903, 0.832 — between 0.832 and 0.919 of them is inside the span — against closed-form removed shares of 0.000, 0.692, 0.590, 0.692. At 14 units a rule with four functions in it takes most of anything it is shown.

The part the rule already took

A diagnostic that reports on what a balancing rule was not handed is run through a column that is 92% inside the span the rule balanced — because orthogonality in the population is not orthogonality on fourteen units.

aimed · Randomisation
A pair pulled back at 20% of the gap per step. Above, the two series. Below, the difference between them. The gap is pulled back towards zero by 20% of itself each step, so it stays inside a band of 14.3 while the series themselves travel much further. Nothing here is stationary except the difference. The faint line below is the gap for two free walks from the same seed, drawn for comparison.

The regression that is not spurious

Two random walks regressed on each other are called significantly related three times in four, so the time-series field ends in a warning. The exception it names and does not measure is here — and when the pair is genuinely tied, the fitted relation converges at rate 1/n rather than the usual 1/√n.

cointegration · Spurious
Six groups of 10, each fitting its own slope, then borrowing. Each faint line is one group's own least-squares slope through its own centre; each solid line is that slope after pooling towards the population slope of 0.79. Every group has the same 10 observations. The group whose x values span 0.4 has a slope standard error of 1.86 and moves 91% of the way in; the group spanning 2.0 has a standard error of 0.37 and moves 28%.

The slope that borrows

Pooling a mean makes it look as though how much a group borrows depends on how much data it has. Pool a slope instead and the illusion breaks — ten groups with ten observations each can borrow anything from 28% to 91%, decided entirely by where those ten observations were placed.

multilevel · Levels
How much of one search the other has already found. Five pairs of searches on one sample, on a scale whose zero and one are both fixed by construction. Zero is two searches over disjoint sets of independent columns: they remove shares of the residual sum that add, at 0.8 standard errors from exactly additive, and they read 0.004. One is a break search paired with a step column it contains, which reads exactly one on every draw because the step adds nothing at all. Between them: two dictionaries of step columns cut a few rows apart read 0.125, and the pair the earlier field measured — a break and a whitening window, both reading the same residual series — reads 0.762, three quarters of the way to one search containing the other. And below zero, a break paired with a search over independent columns reads -0.306: the joint search finds configurations neither half of it contains, so charging the two separately under-charges.

Two searches that share nothing

Two searches over independent columns remove shares of the residual sum that add exactly. On the scale a chi-square point is quoted on they look super-additive by a fifth of a unit, and none of it is overlap.

apart · Criterion
One forecast, and the band the arithmetic puts round it. An AR(1) with φ = 0.75, 60 observations, fitted by least squares and forecast 14 steps ahead. The point forecast decays towards the fitted mean at φ̂^h; the band is ±1.96 standard errors from σ̂²Σψ̂², which grows with the horizon and stops at the unconditional spread 1.72. The dashed pair is the same band computed at the true parameters, which nobody has. The marks past zero are what actually arrived: 12 of 14 inside the band this once, which is one draw and settles nothing.

What the model says next

The usual account of a time series stops at estimation. A forecast asks the other question — not what the parameter is but what the next observation will be — and the band round it is a closed form that grows with the horizon and then stops growing, at a value the series was going to reach anyway.

forecast · Forecast
The same covariate, three ways round. Three worlds over a treatment, an outcome and a covariate, joined by the same three edges at the same three strengths — 0.90, 0.50 and 0.70 — differing only in which way the two edges touching the covariate point. In the first the covariate causes both and adjusting for it recovers the effect of 0.50 exactly. In the second the treatment causes the covariate, the effect is 1.13, and adjusting returns 0.50 — the direct edge alone, with the part that travels through the covariate deleted. In the third the treatment and the outcome both cause the covariate, the effect is 0.50, and adjusting returns -0.087. The regression that produces those three numbers is one formula, and nothing in the data says which panel it is being run in.

One arithmetic, three decisions

A covariate beside a treatment and an outcome can be a common cause of both, a step on the path between them, or an effect of both. The regression that includes it is the same arithmetic in all three, and it is right in one — returning 0.5000, deleting 0.6300 of the effect, and turning 0.5000 into −0.0872.

collider · Conditioning
The first stage an instrument needs is set by the violation nobody can see. The error each estimator converges on when the instrument has a direct effect of 0.05 on the outcome — a path the exclusion restriction asserts is zero and no sample can check. The instrument's error is δ/π exactly, so it is the reciprocal of the very quantity that made the method work: 1.0000 at a first stage of 0.05 and 0.0833 at 0.60. Least squares carries the confounding instead, at 0.3440 at a first stage of 0.30. The two cross at π = 0.1389, and the crossing is exactly δ times 2.7778 — the first stage an instrument needs is proportional to the violation it is assumed not to have, and below that line the method being corrected is the better estimator.

The assumption nothing tests

An instrument buys a causal effect with an assumption no sample can check, and the price is set by the same quantity that made the method work. The first stage it needs is 2.7778 times the violation it is assumed not to have, so a direct effect of 0.05 demands a first stage of 0.1389 and least squares wins below it.

instrument · Exclusion
One wrong model, four designs, four slopes. The slope a straight line converges to when the truth is a quadratic, under four covariate distributions, by two routes: the population projection in closed form, and the mean of 2500 fitted slopes at 200 rows apiece. The even spread over [0, 2] gives 1.6000 and the same spread moved to [1, 3] gives 2.6000, while widening it to [0, 4] gives 2.6000 — the same number as the shifted one, because a symmetric design's target is the truth's tangent slope at the design's own mean and does not read the spread at all. An exponential spread with the SAME mean as the first gives 2.6000. So two studies of one world, each fitting the same wrong model, honestly report slopes 1.0000 apart, and neither is making an error.

What a wrong model estimates

A straight line fitted to a curved truth converges on the tangent at its own design's mean. Two honest studies of one world, fitting the same wrong model, report 2.600000 and 1.600000, and neither is in error.

sandwich · Misspecification
Three mechanisms leave the slope alone; one does not. The bias of the complete-case slope under each of four missingness rules, counted over 4000 studies of 200 rows at 35.0% missing, with the closed form printed beside each count. Missingness that depends on nothing, on the regressor, or on the second covariate leaves the slope exactly where it was — the closed forms are zero to machine precision and the counts are -0.0005, -0.0005 and -0.0011 against standard errors of about 0.0018. Missingness that depends on the outcome moves it by -0.1635, which is 27.3% of the slope being estimated. The same share of rows is lost in every case.

Three mechanisms and one dataset

Four rules for which outcomes go missing, each calibrated to lose the same 35% of the rows and each leaning on what it reads with the same coefficient. Three leave the fitted slope exactly where it was, and the one that reads the outcome moves it by 0.163531.

missing · Missingness
Twenty points and one more, at leverage 0.74. Without the distant point the slope is 0.495; with it the slope is -0.389. Its leverage is 0.737 and its Cook's distance is 24.1, against a conventional threshold of 1.

The line that one point drew

A single observation among twenty-one reverses the sign of a fitted relationship. Its leverage is known from its x value before the outcome is looked at, so this is a property of the design rather than a surprise in the data.

regression · Leverage
What the second search finds, alone and afterwards. For four of the pairs, what the second search removes on its own and what it removes once the first has already run. The gap between the two is the overlap in absolute terms. Where the searches share nothing the two readings are the same: an independent column removes 0.0261 alone and 0.0260 afterwards. Where one contains the other they are 0.1387 and exactly zero. The pair the earlier field measured sits between: a whitening window removes 0.5033 alone and 0.3120 after a break search has run. This is the earlier field's own reading of its pair, on the share scale rather than in log-likelihood units, and it is the number a rule that runs both searches actually has to charge for.

A search that is already the other

A break search shifts every coefficient after a row, so a step column is one of the directions it can move in. Paired with a dictionary of them it reads exactly one, on every draw, and that fixes the top of the scale.

apart · Criterion
The one candidate an effective sample size is right about. n/n_eff with the finite-sample inflation Σ(1 − |k|/n)ρ^|k| is not an approximation to tr(HΩ) for a fit with only an intercept — it is that trace, to machine precision, because the hat matrix of a constant column is 1/n everywhere and its trace against Ω is the mean of Ω. The quoted limit form n(1 − ρ)/(1 + ρ) is not even right about that one. And the average correction the table's fifteen candidates actually need is 3.318 per parameter, well below the scalar, so applying it to all of them over-charges every one.

One number for a table of candidates

An effective sample size is a real quantity, it is exactly right about one thing, and that thing is a mean. Substituted into Akaike's criterion it changes nothing at all, because the penalty it is meant to fix has no sample size in it.

effective · Dependence
Every candidate is behind by what its parameter count says. Each dot is one of the fifteen subsets of four predictors, fitted on a rolling window of 80 rows and scored against the benchmark out of sample over 60 origins, at a null where every one of them contains the truth. The line is σ²(q₀/(R − q₀ − 1) − q/(R − q − 1)), which is arithmetic on two integers and a window length. Most of these pairs are not nested — a subset of two predictors and a different subset of two share neither model — and the closed form does not care: the displacement is a statement about how many coefficients each side estimates. The candidates of the benchmark's own dimension sit at zero.

The displacement is a parameter count

A nested variant is behind its benchmark out of sample before anything is searched for. The closed form for how far turns out to have nothing about nesting in it — only two integers and a window length — and it prices a table where no candidate contains any other.

select · Multiplicity
What a 95% forecast interval covers, counted. 1200 series of 25 observations from an AR(1) with φ = 0.7, at each horizon, on one set of seeds. The upper line is the interval computed at the true parameters — it covers 95.3% on average, which is the check that σ²Σψ² is the right formula rather than a claim about anything a forecaster can do. The lower line is the same formula fed σ̂² and φ̂: 92.8% at one step and 87.3% at 6. The interval that would cover what it claims is 6.9% wider at one step.

The interval that forgets it estimated

The forecast band is derived for a model whose parameters are known, and then computed by putting estimates into it. Counted, the 95% interval covers 87.3% six steps ahead on twenty-five observations, and the point forecast inside it returns to the mean a third faster than the series does.

forecast · Forecast
A window with two wrong ends. The regret of a rule whitened by a Bartlett-tapered Ω̂, as the window widens, against three rules that need no window at all. At L = 0 the estimate is the identity and the rule is exactly least squares — 0.08556, the same number to five places. It falls to 0.01754 at L = 20 and rises again by L = 30, because a quarter of the sample's lags are then being estimated from it. The automatic bandwidth a practitioner would reach for, 4(n/100) to the power 2/9, which at this sample size is 4, gives 0.02963 — 69% above the best available window. The rule told the dependence is an AR(1) sits at 0.01440 throughout, which is the price of not knowing the form.

The window that has to be chosen, and the term that was dropped

An estimated covariance has a bandwidth in it, and both ends of the dial are wrong for different reasons. The rule a practitioner would reach for is two thirds worse than the best window there is.

banded · Order-selection
A test between nested models, under a null that is true. 1000 comparisons: an AR(1) truth, forecast by a fitted AR(1) and by a fitted AR4 whose extra coefficients are zero. In population the two forecasts are identical, so every rejection is false. The larger model's mean squared error is 1.1663 against 1.0583 — worse, by exactly the noise in estimating coefficients that are not there — and the ordinary test therefore declares the smaller model significantly better 67.2% of the time. Read one-sided in the direction anybody asks about, it finds the larger model better 0.0% of the time. Adding the squared difference between the two forecasts back into the loss differential puts the level at 4.9%.

When one model contains the other

The comparison a forecaster most often wants is between a model and the same model with one more term. That is exactly the comparison the standard test cannot make — and it fails by declaring the smaller model significantly better, more confidently the more data it is given.

evaluation · Forecast
Unrepresentative in every respect but the one that matters. Three properties of the complete cases as the chance of being observed leans harder on the regressor, in closed form, at 35.0% of outcomes missing throughout. The mean of the regressor among the rows kept climbs from 0.0000 to 0.5528 against a population mean of zero, and the mean of the outcome from 0.0000 to 0.3980 above its own. The bias in the fitted slope is exactly zero at every one of the ten settings, because selection acting on the regressor alone leaves the conditional law of the outcome given the regressor untouched and least squares conditions on exactly that. The sample is wrong about almost everything and right about the one quantity being estimated.

Dropping the incomplete rows

Push the missingness until the rows that survive have a covariate mean of 0.543905 against a population zero and a variance of 0.5041 against one, and the fitted slope is still exactly right. Where the rule reads the outcome instead, the same sweep takes coverage to 2.42% at eight hundred rows.

missing · Missingness
A weak instrument gives back the problem it was hired for. The counted mean bias of two-stage least squares at 4 instruments and 200 rows, over 2000 draws a setting, against the standard approximation and against the least-squares inconsistency the instrument was brought in to remove. At π = 0.02 the counted bias is 0.3220 ± 0.0142 where least squares is out by 0.3594 — 89.6% of the way back. At π = 0.3 it is 0.0118 against 0.2647. The approximation, the inconsistency over the population first-stage F, tracks the count at the weak end and sits above it in the middle: 0.968, 0.971, 0.918, 0.810, 0.740, 0.722, 0.846 as the ratio of counted to approximated bias.

Weak, and back where it started

A consistent instrumental estimate at two hundred rows and a concentration parameter of 0.32 is biased by 0.3220 ± 0.0142 against a least-squares inconsistency of 0.3594 — 89.6% of the way back to the problem it was hired to solve. Just identified, it has no mean at all, and that is measured as a rate rather than assumed.

instrument · Exclusion
The split decides the width. The width of the interval against the share of 200 observations spent on fitting rather than on calibrating, over 3000 draws. Spending more on the fit shrinks the residuals; spending more on calibration builds the interval at a less extreme order statistic. The two meet at 0.5, where the width is 4.0416 against 4.1603 at 0.1 and 4.3820 at 0.9. Full conformal, which spends the same 200 points on both jobs, is 3.9865 — so the whole cost of splitting is 1.38%.

What the split costs

Splitting a sample between fitting and calibrating looks like a trade against the guarantee, and it is not: coverage moves 0.63 points across nine splits and every reading sits on its own promise. The whole cost is 1.38% of width — and at sixty observations the width falls, rises and falls again.

conformal · Exchangeability
What each error is a claim about, and what the claim comes out as. Each variance estimate's average over 20000 draws, divided by the variance the slope actually has across those same draws, at 80 rows with the error variance leaning towards the edges of the design (γ = 0.8). One is a standard error that is right. The model-based estimate reads 0.6081 of the spread, so its standard error is 77.98% of the one it should report; the four robust corrections read 0.9576, 0.9821, 0.9961, 1.0362. Two further routes agree with the count and share none of its arithmetic: n times the counted variance is 4.8905 against a population sandwich of 4.9200, and the counted ratio of the two standard errors is 1.2799 against a closed form of 1.2806.

The bread and the filling

The robust standard error is not a safety margin. At one setting of the error variance it is 1.2806 times the model-based one and at another it is 0.8246 times it, and the sign of a single dial decides which.

sandwich · Misspecification
A line in the right width beats two curves. How far each candidate charge sits from the measured optimism across the plateau, in units of each width's own standard error, over 2000 draws. The straight line through the origin in the band's summed weights — which is what the earlier field levies — misses by 0.2382 per width. The same straight line in the pairs the band actually uses, Σ w(k)(1 − k/n), misses by 0.0095. A fitted power law misses by 0.0293 and a fitted decaying rate by 0.0172, both on one fitted constant more. The deferral this field answers asked for a curve; the answer is a line, in a variable with nothing fitted in it.

A line that beats two curves

A deferral asked for a curve. Fitted against the same measurements, a straight line in a variable nobody had to fit describes the plateau better than either curve does with a constant more — and for three windows out of four it does not.

curve · Criterion
16 groups shrunk towards a fitted line, at γ = 1.2. Hollow circles are the groups' own values, filled ones the estimates after pooling, and the diagonal is the line fitted through them with each group weighted by how well it is measured — slope 1.21, intercept 0.20. The horizontal rule is where the same 16 groups would have been shrunk to with no covariate. The spread left to borrow against is 0.64 with the covariate against 1.28 without, so every group is pulled further in than it would otherwise have been.

Borrowing towards a line

A group shrunk towards the average of all groups is being compared with groups it has nothing in common with. Fit a group-level predictor and it is shrunk towards what the predictor says a group like it should be — which halves the spread left to borrow against and takes a quarter off the squared error.

multilevel · Levels
What each criterion selects, at 50 observations. 700 series from an AR(2) with coefficients 0.6 and -0.3, every order from 0 to 8 fitted to the same 42 responses so the log-likelihoods are comparable. AIC finds the true order 55.1% of the time and lands above it 25.7%; BIC finds it 54.1% and lands above it 4.0%. The closed form for one extra lag is P(χ²₁ > 2) = 15.73% for AIC, which does not depend on n at all, and P(χ²₁ > ln n) = 4.79% for BIC at this size, which falls to zero. Under the true order is the other failure and it is BIC's: 41.9% against 19.1%.

Choosing the order

One criterion is consistent and one is not, which is the whole of what gets said about them. At two hundred observations the consistent one is right 95% of the time and the other 70%; at fifty they are both right 54% of the time and wrong in opposite directions, and consistency has not started to mean anything yet.

forecast · Order-selection
Least squares estimates persistence low, by an amount with a formula. 3000 series of 50 observations at each persistence. The lower curve is the counted bias of the least-squares estimate of φ, and the open marks on it are −(1 + 3φ)/n, computed rather than fitted. The upper curve is the bias left after adding that quantity back, evaluated at the estimate rather than at the truth nobody has: -0.0020 at φ = 0.3, -0.0023 at φ = 0.5, -0.0039 at φ = 0.7, -0.0059 at φ = 0.8, -0.0108 at φ = 0.9, -0.0165 at φ = 0.95. The formula is a leading-order expression and it understates the bias where the persistence is nearest one — -0.0882 counted against -0.0770 predicted at φ = 0.95, which is the corner of the parameter space every one of these approximations is worst in.

Correcting the persistence

Least squares estimates how much a series remembers of itself as smaller than it is, at every value it can take, by an amount with a closed form. Subtracting that amount back is one line of arithmetic, and what the line costs is variance.

evaluation · Bias
One experiment finding out where to look. A single run of the fully sequential design: 40 runs, the first 8 placed at the guess K = 1, then the model refitted and the design revised after every 2. The marks are the settings the runs were made at. The horizontal lines are where a design built at the truth K = 3 would have put them — 1.875 and 10.00 — and the rule walks onto them without being told: its estimate of K after the first eight runs was 2.694, and by the end 2.765 against a truth of 3. The whole experiment is 96.5% as efficient as the design that knew the answer, where running all 40 at the guess would have been 81.1%.

The design that stops guessing

Every repair so far protects a guess. The alternative is to run part of the experiment, estimate the parameter from it, and design the rest at the estimate — which recovers most of what a threefold wrong guess costs, and has a best moment to stop guessing that is earlier than anyone expects.

robust · Local design
The row count entered twice, and a penalty is one place. Regret against the best available model as the errors are made persistent. Counting rows more than quadruples; both penalty repairs — the trace, and the scalar effective sample size — are worse than it at every persistence measured; and whitening the sample and keeping the ordinary penalty falls, recovering 86.9% of what counting rows gives up at ρ = 0.85. Doing it at an estimated ρ recovers 80.6%, so having to estimate the dependence from the rows being selected on costs 7.3% of what knowing it is worth. Mallows' forms are drawn beside the logarithmic ones and behave the same, which is what rules the linearisation out.

The repair that was exact and made it worse

A penalty computed from the trace is exactly the optimism it estimates, and selecting with it gives up a fifth more than not correcting anything. The row count entered the criterion twice, and a penalty is the second place.

effective · Forecast
A central composite design, 13 runs. Adding 4 axial runs at ±√2 gives every factor three levels, which is the least that can estimate a squared term. The normal matrix now inverts, so each βᵢᵢ has an estimate of its own — and at exactly this axial distance the design is rotatable, which the next figure measures.

Three levels, and the ring where the design says the same thing

A central composite design puts its axial runs at ±α, and α is not a matter of taste. At F to the quarter the prediction variance depends only on how far a point is from the centre and not at all on which direction it lies in — a property with no simulation in it, exact or absent.

surface · Factorial
A filled value is not an observation. What a 95% interval for the slope actually covers after each way of handling 35.0% missing outcomes, counted over 4000 studies of 200 rows. Dropping the incomplete rows covers 95.93%. Filling with the observed mean covers 13.85%, because the estimate itself has moved. Filling with a fitted value covers 80.85% against a closed prediction of 79.73%: the estimate is right and the reported standard error is short by a factor of 0.6567 against a predicted 0.6500, because the residual sum of squares is divided by the whole sample's degrees of freedom. Adding residual noise recovers the spread and covers 85.78% against a predicted 84.62%, since the interval still ignores the variance of having imputed at all.

One imputation is not an observation

Three ways of filling a missing outcome, under a mechanism that makes dropping the rows beyond reproach. Filling with the observed mean covers 13.85%, filling with a fitted value covers 80.85%, adding noise covers 85.78%, and the thing all three were meant to improve on covers 95.93%.

missing · Missingness
The correction does not arrive at the truth, it passes it. The average decay factor a forecast applies to the last observation, at φ = 0.85 and 50 observations, 3000 series per horizon. The middle curve is φʰ, what the model actually does. Below it is the uncorrected forecast, which uses φ̂ʰ and reverts too fast — 24.8% short at h = 4, 30.0% short at h = 6, 32.7% short at h = 8. Above it is the forecast built on the corrected estimate, which overshoots, and the reason is arithmetic rather than a bad correction: raising an unbiased estimate to a power does not give an unbiased estimate of the power, and the higher the power the more the spread of φ̂ is converted into overshoot.

The repair that moves the wrong number

Correcting the bias in a persistence parameter is one line of arithmetic that works. Feeding the corrected estimate into a forecast repairs the number everybody looks at, makes the forecast worse by squared error at moderate persistence, and improves the interval for a reason that has nothing to do with bias.

evaluation · Bias
A fit takes the low frequencies out of what it leaves behind. The autocorrelation of the errors, of the residuals of a fitted benchmark, and of those residuals rescaled by their own leverage. (I − H) removes the component of the errors lying in a column space that is itself slow-moving, so the residuals are less persistent at every lag — by 5.9% at the first and 26.6% by the fourth. The leverage correction is the standard repair for what a fit does to a residual's size; drawn here against what it does to a residual's dependence, it does nothing.

The residuals are not the errors

A fit removes the part of the errors lying in its own column space, and a persistent design's column space is itself slow — so what is left behind is smoother than what went in, at every lag, by an amount that grows with the lag.

effective · Bootstrap
What the interval covers, after a design that read the data. 800 experiments of 12 runs, all at the same truth. The first pair is a design fixed in advance; the second is one whose settings were chosen from the first stage's own outcomes. If choosing the design from the data broke the inference, the second pair would sit below the first, and it does not — 91.3% against 92.6%. What does move the coverage is the shape of the interval rather than the design: the Wald interval assumes the estimate is normal around its own standard error and falls short under both designs, and the profile interval, computed from the same residual sums of squares with no derivative in it, covers 94.3% and 94.9%.

What a design chosen from the data costs

Two fields on this site measured what happens when a rule reads the data, and the error rate broke both times. A design that reads the data to decide where to put its runs breaks nothing — and the control that proves it also finds what the real shortfall is.

robust · Local design
Six scores, one coverage. What each nonconformity score's interval covers, over 2000 draws with 200 calibration points, against the 95.0249% the rank argument promises. The column runs from 94.30% to 94.90%, a spread of 0.60% against a standard error of a difference of 0.69% — one number, six times. That includes a score aimed five units off the fit and a score that never reads the response at all, because the rank argument does not read the score either: it needs the scores exchangeable and nothing else. Every decision a modeller makes has to show up somewhere else, and the next two readings are where.

The score is the modelling

Six nonconformity scores on the same draws cover within 0.60 points of each other, against a standard error of a difference of 0.69 — one number six times. Their widths run over a factor of 2.361 and their adaptivity over a factor of 8.377.

conformal · Exchangeability
What each correction is worth, exactly. Each variance estimate's expectation under a constant error variance, divided by the variance the slope actually has, at 20 rows on an even design, by two routes: the closed form E[eᵢ²] = σ²(1 − hᵢᵢ) carried through each correction's own weight, and the mean of 20000 counted estimates. The maximum leverage here is 0.1857 and the design's fourth-moment share Σu⁴/(Σu²)² is 0.0897, which is the only thing the closed form reads. HC0 comes out at 0.8603 — short by construction, since its factor is exactly 1 − 1/n − Σu⁴/(Σu²)². HC1 reaches 0.9559, HC2 is exactly 1.0000 at every design and every sample size, and HC3 overshoots to 1.1647. On an even design the four are within a fifth of each other and the choice barely matters.

Three corrections and a leverage

On an even design of twenty rows the four robust corrections read 0.8603, 0.9559, 1.0000 and 1.1647 of the truth and the choice barely matters. Add one point at x = 8 and they read 0.3191, 0.3419, 1.0000 and 5.1127.

sandwich · Misspecification
Two far rows, and the line with one of them deleted. Twenty clean points and two rows near x = 9. The slope is −0.511 with every row, −0.376 with one far row deleted, and 0.495 with both deleted. Deleting one of them barely moves the line, because the other is still there.

Two points that hide each other

One far observation among twenty-one has a Cook's distance of 24.1. Put a second beside it and the two read 0.966 and 0.772, neither crossing 1, while together they reverse the slope and deleting both moves the fit by 53.3.

regression · Leverage
Estimating a weight you already know is worth doing. The variance of an inverse-probability estimate weighted by a propensity fitted from the sample, over the variance of the same estimate weighted by the true propensity, paired on the same 500 samples of 600 units at each of five settings. Every reading is below one: the stabilised estimator keeps 27.8% of its true-weight variance where the assignment is nearly a coin toss and 72.0% where it is nearly decidable, and the unstabilised one 30.0% and 49.3%. Neither estimator is materially biased, so this is a variance rather than a trade. The true weights are right about the population and know nothing about the draw; the fitted weights are the value that sets this draw's own imbalance to zero, and that imbalance was what the variance was made of.

The estimated weight is the better one

The propensity is known exactly here, so it can be weighted by — and estimating it from the same data and weighting by that gives a variance ratio of 0.4769 on paired draws. The reason is a projection: the draw's own imbalance explains 56.33% of the true-weight variance and 0.05% of the estimated-weight one.

weights · Weighting
What a 8-run fraction of 4 factors confounds. The defining relation is I = ABCD, so the resolution is 4. A is estimated as A + BCD; B is estimated as B + ACD; C is estimated as C + ABD; D is estimated as D + ABC. Each of those is an identity about the design rather than an approximation about the data.

The word a fraction costs

A half fraction estimates each main effect as an exact sum of that effect and everything it is confounded with — no error term, no sample-size argument. With every interaction at 0.8 the design reports a true effect of −1 as −0.20, and the design cannot test the assumption that makes the number mean anything.

design · Factorial
What the fit calls the shape, against what it is. One eigenvalue held at −3 and the other swept from −2 to 2, so the truth is a maximum on the left and a saddle on the right and the change happens at exactly zero. At an eigenvalue of −0.25 — a genuine maximum — the fit reports a saddle on 26.4% of studies; at +0.25 — a genuine saddle — it reports a maximum on 25.1%. The standard error of a squared coefficient under this design is 0.3791, and the region of confusion is about that wide either side of zero.

The sign the curvature has

A fitted surface reports a maximum, a minimum or a saddle, and the report is a comparison of two estimated eigenvalues against zero. At a true second eigenvalue of −0.25 the fit calls a genuine maximum a saddle on 26.4% of studies, and at +0.25 it calls a genuine saddle a maximum on 25.1%.

surface · Optimum
What one lost run costs a 16-run factorial fitting 11 coefficients. Every run is worth the same: dropping any one multiplies every coefficient's variance by 1.2000, which is 1 + 1/(N − p) with N = 16 and p = 11, and gives every pair of coefficients a correlation of 0.1667 where the complete design had exactly zero.

The run that did not happen

Lose one run from any orthogonal design and every coefficient's variance is multiplied by exactly 1 + 1/(N − p), and every pair of coefficients acquires a correlation of exactly 1/(N − p + 1) where there was none. The price is set by the design's spare capacity and by nothing else, and a saturated design cannot survive it at all.

design · Factorial

Named alongside it

The objects these essays reach for when they reach for this one.

Monte CarloClosed formDegrees of freedomAutocorrelationModel selectionDependenceInformation criterionExperimental designLeverageMean squared errorOverfittingParameter uncertainty

All concepts