Concept

Model misspecification — where it appears

Fitting or designing for a model the data does not obey, which costs precision even where it leaves the estimate unbiased. It is the premise that decides whether a criterion beats a hold-out, since a criterion uses a derivation and a hold-out uses a measurement.

Named by 28 essays across 15 fields — each of them below, with the objects they name alongside it.

The six best bases of 2 functions, and what each protects. Every cell is R²(g | span B) — the share of the imbalance in that shape a rule balancing that basis removes — computed from exact inner products between Hermite functions and indicators, with nothing simulated. The rows are ordered by their worst cell, which is the number an experimenter who does not know the shape is exposed to. The best row here guarantees 26.8% against every shape in the list, and the worst of the six guarantees 15.1%: the difference between them is entirely which subspace was picked, at the same cost per arrival.

A basis is a subspace

A balancing rule cannot tell one basis from another with the same span, so choosing what to hand it is choosing a subspace — and then what it removes of any outcome shape is a projection, computable exactly, with no trial anywhere in it.

basis · Criterion
The price of each thing the rule is not told. What each rule gives up against the best model available, at a persistence of 0.85 on a fifteen-candidate table, over 400 draws. Reading down: least squares with the ordinary penalty; the whitening at the true ρ; the same at a ρ̂ estimated per candidate; that rule with the term the Gaussian likelihood carries and it omits; a Bartlett-tapered Ω̂ estimated once from the fullest candidate at L = 8; the same estimated per candidate; and the truncated Ω̂, which exists on only 45.0% of draws and is averaged over those. Knowing ρ recovers 89.9% of what counting rows gives up, estimating it 83.4%, and estimating a whole covariance 75.6%.

A covariance with no parameter in it

The whitening that repairs a criterion is told the dependence is a first-order autoregression and left to find one number. A real dependence is not one number, and the obvious estimate of it is not a covariance matrix.

banded · Dependence
Four dependences a single parameter cannot tell apart. Every law here is standardised to a lag-one autocorrelation of 0.8, so a rule told the errors are a first-order autoregression finds the same number in all four and has no way of seeing what separates them. The geometric decay is the world in which estimating a covariance rather than naming it was priced, and found to cost. The five-period moving average has 0.200 at the fourth lag and exactly nothing past it, where the geometric law says 0.328 at the fifth. Long memory at d = 4/9 is still at 0.576 by the twentieth lag, where the geometric law has reached 0.012. The break has no autocorrelation function at all: what is drawn for it is the average over the pairs at each gap, which is what a stationary estimate converges to.

A dependence with a shape

Four ways for errors to repeat, all with the same first lag and nothing else in common. A rule told the errors are a first-order autoregression finds the same number in all four, and is right about one of them.

general · Dependence
The rule is parity, and it runs both ways. At a correlation of 0.5, four combinations of a dictionary and an outcome shape. The joint sign flip (X, Y) → (−X, −Y) leaves the bivariate normal alone at every correlation, so a function that changes sign under it is orthogonal to one that does not. A product of two odd functions is even; a product of an odd and an even one is odd. So an odd dictionary removes exactly none of the first and something of the second, and an even dictionary does the reverse — which it does, to machine precision, in both of the two rows that should be zero. This is one rule where there had been two: that a median split's square is constant, and that a polynomial dictionary contains the products a correlation generates.

A dictionary that is neither

A rule handed two median splits removes none of their interaction; a rule handed two covariates removes none of their product. Those were two results with two explanations, and they are one result with one — and finding it corrected the number underneath both.

dict · Criterion
Four rules, three shapes, and no ordering that survives. The variance of the unadjusted treatment estimate under each rule, as a fraction of the variance a coin gives, over 450 trials of 200 units each. Against a covariate that enters linearly the rule that reads the number nearly halves it. Against a threshold at 1 it removes about a fifth. Against a quadratic every rule here is at or worse than a coin — they are all optimising a criterion that is one over the variance of an estimate in a model this outcome does not obey, and a constraint that helps nothing still costs something. Nothing in a trial says which column it is in.

Balanced on the wrong function

A rule that reads a covariate's numbers halves the variance of the treatment estimate, if the covariate enters the outcome as a straight line. If it enters as a threshold the rule is worth a fifth of that, and if it enters as a curve every rule here is worse than a coin.

shape · Blocking
One wrong model, four designs, four slopes. The slope a straight line converges to when the truth is a quadratic, under four covariate distributions, by two routes: the population projection in closed form, and the mean of 2500 fitted slopes at 200 rows apiece. The even spread over [0, 2] gives 1.6000 and the same spread moved to [1, 3] gives 2.6000, while widening it to [0, 4] gives 2.6000 — the same number as the shifted one, because a symmetric design's target is the truth's tangent slope at the design's own mean and does not read the spread at all. An exponential spread with the SAME mean as the first gives 2.6000. So two studies of one world, each fitting the same wrong model, honestly report slopes 1.0000 apart, and neither is making an error.

What a wrong model estimates

A straight line fitted to a curved truth converges on the tangent at its own design's mean. Two honest studies of one world, fitting the same wrong model, report 2.600000 and 1.600000, and neither is in error.

sandwich · Misspecification
A weight that balances, and one that unbalances. The standardised difference between the arms on each covariate, integrated over the population rather than counted in a sample. Unweighted, the arms differ by 0.8310 on the first covariate and 0.6015 on the second, which is what makes the raw difference of arm means 2.7102 against a true average effect of 1.0000. Weighting each unit by one over its own assignment probability removes both differences exactly — -2.78e-17 and -5.69e-19, which is machine precision and not a small number — because the weighted density of the treated arm is the population's own whatever the propensity is. Weighting by a score fitted without the second covariate balances the first to 0.0035 and pushes the second out to 0.7057, further apart than doing nothing.

A score that balances

Weighting each unit by one over its own assignment probability drives the standardised difference between the arms from 0.8310 to 2.8×10⁻¹⁷ — exactly, not nearly. A score fitted without the second covariate leaves that covariate at 0.7057, further apart than doing nothing at all.

weights · Weighting
A threshold in the tail is a threshold nothing balances. The share of a coin's imbalance in an indicator 1{x > c} that survives a rule which balances the covariate itself. The smooth curve is 1 − ρ² with ρ = φ(c)/√(p(1−p)), a closed form with no trial in it; the points are counted over 500 trials of 200 units at each threshold. At the median the two agree that about a third survives — the removed share is exactly 2/π — and by two standard deviations 86.9% survives. The closed form is exact in the limit and optimistic by a few points at this many units, because the rule balances the sample's mean rather than the population's.

A threshold in the tail

How much of a threshold's imbalance a balanced covariate removes is a correlation, and the correlation is a closed form. At the median it is exactly 2/π — the same 2/π a median split throws away — and two standard deviations out it is an eighth.

shape · Blocking
Four windows, one line, and one that is off it. The optimism measured for each window at a band of 30 lags, against what that window's weights sum to, on 2000 pairs of independent samples of 120 rows. The diagonal is where a window that spent exactly its summed weights would sit. Three of the points are one shape at three levels — the Bartlett window, its square and its cube, whose sums stand in the ratio 6 : 4 : 3 — and they lie on a line through the origin at 0.767 of the diagonal, with 0.033 between the highest and the lowest. Scaling the weights scales the charge by the factor the weights predict, which is what makes the weights the mechanism. The Parzen window has a comparable sum and a different shape, and it sits at 0.871: its weights stay near one over the first few lags, and the first few lags are where the information is. A weight sum treats every lag as equally informative and no sample does.

What a window leaves free

A Bartlett window's weights sum to exactly half its width, which is a candidate for what the band costs. Varying the weights without varying anything else says the weights are the mechanism; varying the shape at the same weight says they are not the arithmetic.

dimension · Criterion
Generality in the wrong direction buys nothing. Regret on a sample whose persistence changes from 0.95 to 0.65 at row 60, over 200 draws. The three stationary rules — told one number, told a window, told an order — are within 0.4 standard errors of each other, and all three stop in the same place: they are general in the lag direction, and the departure is in the other one. Letting the model change once, at a point estimated from the same residuals, is worth 0.05021 more at 4.5 paired standard errors — about as much again as the whole of the first repair. Being told where the break is adds 0.01926, and being told the entire covariance adds 0.02465.

Where the generality runs out

A covariance that changes half way through a sample is not one a window can estimate. One number, a window and an order are worth the same as each other on it — and letting the model change once, at a point nobody can locate, is worth as much again as all three.

general · Dependence
The guarantee, as the basis is allowed more functions. The lower line is the best worst case over the six named shapes for a basis of each size, found by scoring every subset of the dictionary — an exact answer, since the problem is finite. One function guarantees 2.3%, which is nearly nothing; three guarantee 59.0% and the basis that does it is the covariate, its square and its cube, with no indicator in it. The upper line is the same problem with the basis drawn rather than fixed, which is worth 2.09 times as much at two functions and 1.32 at three. The two lines converge because a basis large enough to protect everything has nothing left to randomise over.

Which shapes are worth protecting

Choosing a basis by its worst case is a finite problem with an exact answer. The answer has no tie in it, which a maximin optimum is supposed to have — and the tie comes back, along with twice the guarantee, when the basis is drawn rather than chosen.

basis · Blocking
What a rule reads, against what the outcome uses. The variance of the treatment estimate relative to a coin's, for four things a rule might balance against three shapes the outcome might have, over 350 trials of 200 units. The diagonal is the easy part — a rule that reads the function the outcome uses removes about half the variance. What the table is for is the off-diagonal: reading the covariate alone is worth nothing against a quadratic (0.755), and reading all three is worth nearly as much against every shape as the matching rule is against its own (0.532, 0.493, 0.514).

Three functions of one number

A rule that balances the covariate is exposed to every shape the outcome might have. A rule that balances three functions of it costs two points of variance against the shape the first was built for and takes the worst case from a coin's to about half of it.

shape · Criterion
A class that is a subspace has no guarantee below its own dimension. Each cell is the worst case over every unit-variance function in a class of dimension m, for a rule reading k functions: the smallest squared principal-angle cosine between the two subspaces. Wherever k is less than m the number is zero to machine precision, and that is not a weak guarantee but the absence of one — some direction of the class is orthogonal to the entire basis, and against an outcome in that direction the rule does exactly what a coin does. An experimenter who declines to name the shapes and asks instead to be protected against everything smooth is asking for the cells above the diagonal.

Where the guarantee is exactly zero

An experimenter who declines to name the shapes, and asks instead to be protected against anything in a class, is asking for a number that is not small but zero. Bounding the class is unavoidable, and the two ways of doing it choose different bases.

basis · Randomisation
What each construction carries, against what there was. The autocorrelation of a resampled error series at five lags, averaged over 60 samples of 40 resamples each. Three facts are in the picture. The residuals lie below the errors at every lag, which is the ceiling a multiplier cannot exceed. The blocked multiplier and the fixed-length block lie on top of each other below it — they attenuate identically, because the attenuation is the join — while the stationary bootstrap, whose runs are geometric rather than fixed, sits above them both. And the sieve is the exception in kind rather than in degree: at lag six it carries 0.0638 where the residuals have 0.0300 and the multiplier has -0.0011, because a fitted model extrapolates past the lags it was told about and a truncated sample sequence cannot.

Errors generated from a fitted model

The one construction that is not bounded by the residuals, because a model extrapolates past the lags it was told about and a truncated sample sequence cannot. It is nearly exact where the only defect is dependence, and it pays for it where there are two.

banded · Reference
What each analysis does at a true null, by shape. Four analyses of the same trials — 500 of them at each shape, 120 units, assigned by the rule that reads the covariate. Every rejection is false. The unadjusted analysis is the one that moves: 1.60% against a linear outcome, where the design removed a great deal that the standard error still prices, and 5.20% against a quadratic, where it removed nothing and the standard error is right. Adjusting holds the level in all three columns, and so does the design's own reference distribution, which needs to be told the rule and nothing else.

The analysis and the shape

An unadjusted analysis after a rule that read the covariate is too cautious — by a third against a linear outcome, by nothing at all against a quadratic. And an adjustment for the wrong function recovers almost none of the precision the right one would.

shape · Randomisation
A fit takes the low frequencies out of what it leaves behind. The autocorrelation of the errors, of the residuals of a fitted benchmark, and of those residuals rescaled by their own leverage. (I − H) removes the component of the errors lying in a column space that is itself slow-moving, so the residuals are less persistent at every lag — by 5.9% at the first and 26.6% by the fourth. The leverage correction is the standard repair for what a fit does to a residual's size; drawn here against what it does to a residual's dependence, it does nothing.

The residuals are not the errors

A fit removes the part of the errors lying in its own column space, and a persistent design's column space is itself slow — so what is left behind is smoother than what went in, at every lag, by an amount that grows with the lag.

effective · Bootstrap
One group 6 population widths from the rest. Squared error for each group under partial pooling, with each group's own mean beside it. Seven of the eight are estimated better by pooling. The eighth, which was never from the population, is estimated 6.1 times worse — 13.9 against 2.3.

When borrowing goes wrong

Partial pooling wins on the total and can lose badly on one group. Placed six population widths out, the group that was never from the population is estimated six times worse than by its own mean — and nothing in the output says so.

hierarchical · Pooling
Six scores, one coverage. What each nonconformity score's interval covers, over 2000 draws with 200 calibration points, against the 95.0249% the rank argument promises. The column runs from 94.30% to 94.90%, a spread of 0.60% against a standard error of a difference of 0.69% — one number, six times. That includes a score aimed five units off the fit and a score that never reads the response at all, because the rank argument does not read the score either: it needs the scores exchangeable and nothing else. Every decision a modeller makes has to show up somewhere else, and the next two readings are where.

The score is the modelling

Six nonconformity scores on the same draws cover within 0.60 points of each other, against a standard error of a difference of 0.69 — one number six times. Their widths run over a factor of 2.361 and their adaptivity over a factor of 8.377.

conformal · Exchangeability
Either model is enough; neither is not. The bias of three estimators of an average effect of 1.0000, over 600 samples of 600 units with the assignment rule at strength 1, in each of the four cells made by getting each nuisance model right or wrong. The wrong model in both cases is one that omits the second covariate, which the outcome and the assignment both depend on. The outcome model alone is off by 0.8064 whenever it is the wrong one; weighting alone is off by 0.8190 whenever the propensity model is. The augmented estimator built from both is off by -0.0085, -0.0083 and -0.0016 in the three cells where at least one of them is right, and by 0.8118 in the fourth — which is between its two components rather than better than either.

Either model, but not neither

The augmented estimator's bias is −0.0085, −0.0083 and −0.0016 wherever one nuisance model is right, against components off by 0.8064 and 0.8190. One step past the overlap sweep it is the least biased estimator on the table at 0.0857 and the worst on it at 1.9265.

weights · Weighting
Where the bias lands. The drift in the log variance ratio, fitted across 12 blocks over 4000 trials. E[log λ̂_b] is log λ_b plus ψ(k_B/2) − log(k_B/2) − ψ(k_A/2) + log(k_A/2), which depends on nothing but the degrees of freedom — so the tempting sentence is that it goes into the intercept and leaves the slope alone. It does not, because the blocks alternate between allocations and the alternation is correlated with the covariate being fitted: the lopsided blocks carry 0.5383 of bias and the even ones carry none. Uncorrected the slope reads 1.5597 against a truth of 1.5, which is 8.0 standard errors. Subtracting the two digammas block by block leaves 1.4976.

The bias that lands in the slope

The bias in a log variance estimate depends on nothing but its degrees of freedom, so it goes into the intercept — unless the degrees of freedom alternate with the design, which is exactly what a block-randomised trial makes them do.

blocks · Width
The damage does not stay in the term that was left out. Where each coefficient lands when the model that fills the missing outcomes and the model that analyses them disagree, over 1500 studies of 200 rows at 35.0% missing and 20 imputations. An imputer that omits a covariate the analysis fits attenuates that covariate's coefficient by exactly the missing fraction — -0.1405 counted against a closed -0.1400 — and pushes the coefficient it did impute on the other way by exactly the product of the omitted coefficient, the covariates' correlation and the missing fraction: 0.0402 counted against 0.0420. Both closed forms come out of the same two-by-two solve. Matching models leave both alone, and so does an imputer that knows more than the analysis.

An imputation model the analysis does not contain

A model that fills the gaps without a covariate the analysis fits attenuates that covariate's coefficient by exactly the missing share, 0.4 to 0.26, and moves the one it did carry by exactly γρf, 0.6 to 0.642. The reverse case is supposed to inflate the interval, and at four strengths of the extra knowledge it does not.

missing · Missingness
What a scale that grows across the sample costs. What the interval covers when the noise scale grows across the sample, against how far the departure has gone, over 1500 draws at each setting. The coverage runs from 94.47% at no departure to 83.93% at the end of the sweep, a loss of 11.07%. The rank argument needs the 200 calibration scores and the test score to be exchangeable, and this is one of the three ways that fails. A test built for it reaches 80% power at 4.054, where the coverage is 85.13% — so 9.87% of the loss is inside the region such a test would have missed.

When the order matters

Three ways of breaking exchangeability cost 4.93, 11.07 and 1.07 points of coverage, and the ordering by cost is the reverse of the ordering by how soon a test would have caught them. The departure practitioners check for is the cheapest one.

conformal · Exchangeability
One treatment, a different hazard ratio at every follow-up. The hazard ratio a Cox model converges to, found as the root of its expected score by numerical integration, as the trial runs longer; dropout at 0.1 throughout. The proportional treatment reads 0.5 at every τ. The waning treatment reads 0.5000 at τ = 1, 0.6362 at τ = 3 and 0.7890 at τ = 8 — the same two arms, the same effect in the same first year, and a number that drifts towards one as later, effect-free events are added to the average. The dots are the mean of 400 Cox fits with 400 subjects an arm: 0.5014 at τ = 1, 0.7020 at τ = 2, 0.7635 at τ = 3, 0.8034 at τ = 5, 0.8208 at τ = 8. The crossing treatment reads 0.3429 at τ = 1, exactly 1 at τ = 3 by construction, and 1.0611 at τ = 8: beneficial, null or harmful according to when the trial stopped.

The hazard ratio the follow-up chose

A treatment that halves the hazard for one year and then does nothing has a Cox hazard ratio of 0.5000 if the trial stops at one year, 0.7617 at three and 0.8194 at eight. Nothing about the treatment differs between those numbers. When hazards are not proportional the hazard ratio is an average, and the length of follow-up and the dropout rate choose its weights.

survival · Censoring
Weights that balance a sample by construction. What three sets of weights leave of the standardised difference between the arms on each covariate, as a root mean square over 1200 samples of 600 units. The true propensity leaves 0.1317 and 0.1186 — a sampling error, since it is right about the population and knows nothing of the draw. A likelihood fit leaves 0.0770 and 0.0657, having absorbed part of the draw's imbalance as a side effect of fitting the treatment. Weights fitted so that each arm's weighted means are the sample's leave 1.4e-14 and 1.2e-14, which is the arithmetic's floor rather than a small number: the largest gap between a weighted arm mean and the sample mean in any draw is 9.8e-14.

A weight fitted to balance

Weights fitted so that each arm's weighted covariate means equal the sample's leave a difference of 1.4×10⁻¹⁴ between the arms and give the estimate a third of the variance of weights fitted by likelihood — 0.011883, within a relative 5.8% of the bound no estimator can beat. In the world where the assignment carries a square nobody named, the same exact balance leaves the square further apart than no weighting at all, and where the outcome carries it too the estimate is wrong by 0.6973 with an interval that covers 1.5%.

weights · Weighting
One weighting told the means and one told the second moments, in five worlds. The bias of the fit to balance over 600 samples of 600 units in each world, fitted to the covariates' means and fitted to their means, squares and product. Told the means it is off by -0.0004, -0.0020, 0.0103, 0.6973, 0.2698 in the worlds with no square, a square in the assignment, a square in the outcome, a square in both and a cube in both; told the second moments, by -0.0009, -0.0000, 0.0009, -0.0045, 0.3099. Its interval covers 94.0%, 94.7%, 95.0%, 1.5%, 51.0% and 93.7%, 89.8%, 94.0%, 91.0%, 48.3%.

The moments a balance is told

Weights fitted to balance the covariates' means were wrong by 0.6973 in the world where both the assignment and the outcome carry a square. Told the squares and the product as well, the same construction is off by −0.0045 there and its interval covers 91.0%. The failure moves up a moment rather than away: with a cube in both, the second-moment balance is off by 0.3099 and leaves the cube twice as far apart as no weighting. And where overlap is thin, 37.0% of samples have no such weights at all.

weights · Weighting
A peak where the recorded data have none. The profile log-likelihood of a selection model in cy, the coefficient that lets the chance of being recorded depend on the outcome itself, for one study of 800 rows whose missingness is at random, with residuals normal; every other parameter is maximised at each fixed value. The model assumes the outcome is normal given the covariates. The curve peaks at cy = 0.35, where the fitted slope is 0.839, and the values of cy within the 95% cut run from −0.13 to 0.75; the likelihood-ratio statistic against cy = 0 is 1.47. The study was drawn with cy = 0.00. With the outcome's law left free, every value of cy fits the recorded rows equally well and this curve would be flat: its curvature is the normal assumption.

The assumption that identifies the mechanism

A selection model estimates how strongly an outcome decides whether it is recorded — the quantity two identical datasets showed no statistic can see — and it does so by assuming the outcome is normal. Where that holds and the outcome does decide, it repairs a slope complete cases put at 0.4318 to 0.5795. Where the missingness is at random and the residual is merely skewed, it reports selection that is not there, moves the slope from 0.5971 to 1.0319, and rejects missingness at random in 72.5% of studies.

missing · Missingness
A prior worth 35 observations, moved across the range — truth 0.1, n = 20. The same prior weight centred at each of 33 places. Its interval covers 100.0% where the centre is near the truth and 0.0% at its worst, while the mean width where it covers least is 0.221 against a flat prior's 0.263 on the same data.

When the prior is confident and wrong

A prior worth thirty-five observations, centred in the wrong place, produces a 95% interval that covers nothing at all — and reports a width 5% narrower than an honest one. It takes seventeen thousand observations to repair, not thirty-five, and the worst study to run is the one whose sample size equals the prior's weight, exactly.

bayes · Credible
Twelve groups from two clusters, τ = 1. Every group's truth is in one of two clusters, and the population's total spread is exactly τ = 1, so the analysis recovers τ̂ = 1.515 and every shrinkage weight is what it would be for a single normal population. The estimates are pulled towards the grand mean, which is the middle of the gap — a place 0 of the 12 truths are and 6 of the estimates end up.

One population, or two

Group effects from two clusters rather than one bell, with the same total spread. The analysis recovers the same population spread, uses the same weight for every group, and reports nothing unusual — while 46% of its estimates land in a region holding 6.6% of the truths.

hierarchical · Pooling

Named alongside it

The objects these essays reach for when they reach for this one.

Covariate balanceClosed formNuisance parameterThresholdAllocation ruleAutocorrelationBasis functionsDesign criterionEfficiencyEstimandProjectionTreatment effect

All concepts