Concept

Approximation error — where it appears

The gap between a quantity and the closed form or model standing in for it, as distinct from the sampling error in estimating either. It is not reduced by more data, so a procedure whose approximation error dominates its sampling error stops improving once the sample is large enough for that to be true.

Named by 6 essays across 4 fields — each of them below, with the objects they name alongside it.

A model of the active set, and the active set. Each of the 14 units of one design, at the share of its exchanges the tolerance box blocks — computed from the design's columns and the tolerance under a uniform position in the box, against counted over all 116 admissible assignments. The diagonal is where the two would agree. Over 192 designs they agree about the ordering of the units at a correlation of 0.8141 ± 0.0112, negative on 0.5% of them, and disagree about the level: 0.8442 counted against 0.8170 modelled, a gap of 0.0272 ± 0.0051. An admissible assignment does not sit uniformly in its box, and this is the size of that.

A model and a count

The share of a unit's exchanges a tolerance box refuses can be modelled from the design or counted over the admissible set. They order the units the same way at a correlation of 0.81 and disagree about the level by 0.027.

blocked · Randomisation
Sheppard's arcsine, by two routes. Corr(sign X, sign Y) as the covariates' correlation runs from zero to one, drawn twice. One route is a sixty-four-node quadrature of the orthant probability over the correlation — the general construction, which works at any pair of cut points; the other is (2/π) arcsin ρ, which is elementary and works only at the median. They agree to 3.3e-16 at every one of 81 correlations, which is what licenses the quadrature everywhere else. The curve is above the diagonal at small ρ and below it at large: two signs agree with probability ½ + arcsin(ρ)/π, so a correlation of 0.5 gives exactly ⅓ and a correlation of 0.8 gives 0.5903.

The arcsine that closes it, and the error that was overstated

Two median splits of a correlated pair agree with probability ½ + arcsin(ρ)/π, exactly. And the truncation the field was avoiding falls geometrically in the correlation, not algebraically in the order.

splits · Routes
A weak instrument gives back the problem it was hired for. The counted mean bias of two-stage least squares at 4 instruments and 200 rows, over 2000 draws a setting, against the standard approximation and against the least-squares inconsistency the instrument was brought in to remove. At π = 0.02 the counted bias is 0.3220 ± 0.0142 where least squares is out by 0.3594 — 89.6% of the way back. At π = 0.3 it is 0.0118 against 0.2647. The approximation, the inconsistency over the population first-stage F, tracks the count at the weak end and sits above it in the middle: 0.968, 0.971, 0.918, 0.810, 0.740, 0.722, 0.846 as the ratio of counted to approximated bias.

Weak, and back where it started

A consistent instrumental estimate at two hundred rows and a concentration parameter of 0.32 is biased by 0.3220 ± 0.0142 against a least-squares inconsistency of 0.3594 — 89.6% of the way back to the problem it was hired to solve. Just identified, it has no mean at all, and that is measured as a rate rather than assumed.

instrument · Exclusion
The upper tail of 10 exponential draws: normal, one Edgeworth term, two Edgeworth terms, each against the exact tail. Each curve is an approximation divided by the exact gamma tail, so 1 is exact. Six standard deviations out at n = 10: the normal gives ×0.0000670, one Edgeworth term ×0.00159, two ×0.0167 and the saddlepoint ×1.0007.

A correction that goes below zero

One Edgeworth term takes the normal approximation's error at two standard deviations from 38% to 8% on ten exponential draws, and stretches the range within 10% of the truth from 1.66 to 3.09 standard deviations at a hundred. It also turns negative in the short tail at every sample size — past 3.13 standard deviations at a hundred draws and 9.83 at a hundred thousand — because the region recedes only as the sixth root of n.

expansion · Rate
The upper tail of 10 exponential draws: normal, two Edgeworth terms, saddlepoint, each against the exact tail. Each curve is an approximation divided by the exact gamma tail, so 1 is exact. Six standard deviations out at n = 10: the normal gives ×0.0000670, one Edgeworth term ×0.00159, two ×0.0167 and the saddlepoint ×1.0007.

An approximation built at the threshold

The saddlepoint approximation reads the tail of a sum of five exponential draws to within 0.19% six standard deviations out, where the normal is short by a factor of more than sixty thousand. It is within 2.2% out to ten standard deviations on a single draw, where there is nothing to average, and within 1.1% on a binomial whose expected count is one. It works because it is built where the tail is read rather than at the mean.

expansion · Rate
One of them is mostly leverage. How much of the design's own leverage direction each active-set probe carries, once both are standardised and projected off the rule's span — which is what a probe is, so it is the comparison that matters. Over 189 designs the modelled active set agrees with leverage at |r| = 0.8359 ± 0.0114 and the counted one at 0.4239 ± 0.0216. So the modelled probe is largely leverage under another name and the counted one is genuinely a different direction — and the counted one is the worse probe, at 0.3854 of alignment against 0.4272. What the active set contains beyond leverage points away from where the set splits.

Counting it exactly does not help

If a modelled active set lost because the model was crude, the exact one would win. It is computed at a cost no trial can pay, and it is worse — so the approximation was never what was costing the probe.

blocked · Randomisation

Named alongside it

The objects these essays reach for when they reach for this one.

Covariate balanceExperimental designNormal approximationAssignment mechanismClosed formConnected componentCorrelationDiscretenessEdgeworth expansionExact enumerationImbalanceLeverage

All concepts