Concept

Estimated variance — where it appears

A variance computed from the data rather than known, and therefore itself uncertain — the substitution most plug-in formulas make without saying so. Substituting it for a known one is what turns a normal interval into a t interval, and the degrees of freedom of the substitution are the price.

Named by 30 essays across 13 fields — each of them below, with the objects they name alongside it.

The stopping rule costs more than the weighting does. Coverage over 2000 runs of a trial whose variance ratio drifts by a factor of twenty, at three ways of deciding when to stop. Twelve blocks fixed in advance is the top line and reproduces what a trial of fixed length delivers. Stopping when the reported interval is short enough is the bottom line, and it costs between 3.0% and 5.5% of coverage — including for the rule that is told every block's true ratio, which is what says the shortfall belongs to the stopping and not to the weights. Stopping on a width predicted from the within-arm sums of squares is the middle line, and it is back at the fixed-length values. The standard error on each point is 0.49%.

A width the trial has to stop for

The weighting that covers at 94.9% on twelve blocks covers at 91.5% when the trial stops as soon as its interval is short enough — and so does the rule that is told every block's true variance ratio. The shortfall is the stopping, not the weights.

stop · Stopping
Three intervals, one shortfall. What each of three intervals actually covers, at four rules and two block windows, over 300 samples of 120 rows. All three are built from the same resamples on the same draws, so a difference between them is a difference in what is done with the resampled series. Not one of the twenty-four cells reaches the ninety-five per cent it promises. The studentised interval runs from 75.7% to 92.3%, the percentile interval — the earlier field's — from 80.0% to 89.7%, and a normal interval on the same scale from 81.7% to 89.0%. The standard repair for a percentile interval's shortfall does not repair it.

An interval that carries its scale

A percentile interval inherits the resampled distribution's skewness and its scale error together. The standard repair is one extra variance per resample. It was named and not run, so this runs it.

student · Bootstrap
What x-bar plus or minus 2 sample standard deviations holds, at n = 10. The content of the band is a random variable. Across 20,000 normal samples of 10 it averages 91.1%, its fifth percentile is 74.7%, and it falls short of 95% on 59.9% of samples. The band that is drawn to show where 95% of the data lies.

Two standard deviations of what

The 95.45% inside two standard deviations is a fact about a curve whose centre and width are given. Drawn from ten observations, the same band holds 91.1% on average and less than 95% on 59.9% of samples — and the average is the reading that hides it.

estimated · Bands
One relationship at five designs, residual spread 1.00. Every panel has the same slope of 1, the same intercept of 0 and the same residual standard deviation of 1.00. Only the range of x differs. R-squared runs from 0.021 to 0.849, and the estimated residual spread is 0.9932 in all five.

R² is a property of the design

One line, one residual spread, five studies that differ only in how far apart they placed their x values. R² runs from 0.021 to 0.849 and the estimated residual spread is 0.993 in every one of them. Nothing about the relationship changed.

spread · Summary
Which side a 95% t interval misses on, exponential source. Both tails should be 2.5%. At 8 observations the interval falls short of the mean on 9.75% of samples and overshoots on 0.31%. At 500 they are 3.31% and 2.05%, and the total is 5.36% — which a coverage table reports as very nearly right.

Where the two tails disagree

A 95% t interval on an exponential source at 120 observations covers 94.81%, which reads as very nearly right. It misses below the mean on 4.08% of samples and above on 1.11% — one tail 63% too heavy and the other 56% too light, and the total is the statistic that hides it.

tails · Student
The power trials actually have when sized for 80% from a pilot of 10. Four thousand pilots of 10 observations, each sizing a trial for 80% power at half a standard deviation from its own standard deviation. 55.9% of the trials have less than 80% power and 11.1% less than 50%; the median trial has 76.8%.

The spread a pilot supplies

A trial sized for 80% power from a pilot's standard deviation is sized from an estimate that is too small more often than not. With a pilot of ten, 55.9% of the trials it sizes have less than 80% power and 11.1% less than 50%, although the planned sample is right on average. Sizing from the pilot's 80% upper confidence limit instead leaves 19.8% short, at 1.65 times the sample; from its 90% limit, 10.0% short at 2.12 times.

planned · Power
How wrong the ratio is allowed to be. λ enters only through the weights, so misstating it leaves the estimate unbiased and moves two things — the interval's calibration and its efficiency — both of which are closed forms of the design. Coverage stays at its level over a factor of two in either direction (94.93% at half the truth, 94.27% at twice it) and starts to go at a factor of five. An estimate on hundreds of within-arm degrees of freedom is never wrong by anything like that, which is what makes the feasible rule usable rather than merely definable.

Blinded, and still exact

The one number the exact interval needs is a ratio of within-arm spreads, which is a contrast and contains no mean — so a rule forbidden to look at the effect may compute it, on more degrees of freedom than the interval itself has.

corner · Width
Two promises, and no rule here keeps both. A fixed-width procedure promises two things: that the interval covers at its nominal rate, and that it is no wider than the width asked for. Over 1500 runs of the modelled weighting, a rule that stops when the interval it will report is short enough keeps the width — only 2.0% of runs come out wider than 0.34 — and covers at 91.13% against a nominal 95%. A rule that stops on a width predicted from the within-arm sums of squares covers at 94.80% and comes out wider than promised on 42.3% of runs. The two promises are in conflict because keeping the second one exactly requires conditioning on the very quantity that has to be independent of the stopping time for the first.

Stopping on the arms

The width a trial will report is predictable from quantities the interval is not about. A rule that stops on the prediction covers at 94.5% where one that stops on the interval covers at 91.5, and it costs two blocks and half of the width promise.

stop · Width
What a 95% forecast interval covers, counted. 1200 series of 25 observations from an AR(1) with φ = 0.7, at each horizon, on one set of seeds. The upper line is the interval computed at the true parameters — it covers 95.3% on average, which is the check that σ²Σψ² is the right formula rather than a claim about anything a forecaster can do. The lower line is the same formula fed σ̂² and φ̂: 92.8% at one step and 87.3% at 6. The interval that would cover what it claims is 6.9% wider at one step.

The interval that forgets it estimated

The forecast band is derived for a model whose parameters are known, and then computed by putting estimates into it. Counted, the 95% interval covers 87.3% six steps ahead on twenty-five observations, and the point forecast inside it returns to the mean a third faster than the series does.

forecast · Forecast
What studentising costs. How much wider the studentised interval is than the percentile one, cell by cell, over 300 draws, with what each cell gains in coverage beside it. Averaged over the eight cells the interval is 2.09 times as wide and covers 0.46 points better. At the two rules that choose short blocks the two intervals are within a fifth of each other; at the oracle's length, where a resample holds two or three whole blocks, the studentised interval is 4.37 and 5.04 times as wide. A repair that doubles the width and buys half a point is not one a reader could not have had by widening the interval it replaced.

What studentising costs

Averaged over eight cells the studentised interval is 2.09 times as wide as the percentile one and covers 0.46 points better. At the block lengths the rules choose, the scale it divides by rests on two or three numbers.

student · Bootstrap
Generality in the wrong direction buys nothing. Regret on a sample whose persistence changes from 0.95 to 0.65 at row 60, over 200 draws. The three stationary rules — told one number, told a window, told an order — are within 0.4 standard errors of each other, and all three stop in the same place: they are general in the lag direction, and the departure is in the other one. Letting the model change once, at a point estimated from the same residuals, is worth 0.05021 more at 4.5 paired standard errors — about as much again as the whole of the first repair. Being told where the break is adds 0.01926, and being told the entire covariance adds 0.02465.

Where the generality runs out

A covariance that changes half way through a sample is not one a window can estimate. One number, a window and an order are worth the same as each other on it — and letting the model change once, at a point nobody can locate, is worth as much again as all three.

general · Dependence
What each error is a claim about, and what the claim comes out as. Each variance estimate's average over 20000 draws, divided by the variance the slope actually has across those same draws, at 80 rows with the error variance leaning towards the edges of the design (γ = 0.8). One is a standard error that is right. The model-based estimate reads 0.6081 of the spread, so its standard error is 77.98% of the one it should report; the four robust corrections read 0.9576, 0.9821, 0.9961, 1.0362. Two further routes agree with the count and share none of its arithmetic: n times the counted variance is 4.8905 against a population sandwich of 4.9200, and the counted ratio of the two standard errors is 1.2799 against a closed form of 1.2806.

The bread and the filling

The robust standard error is not a safety margin. At one setting of the error variance it is 1.2806 times the model-based one and at another it is 0.8246 times it, and the sign of a single dial decides which.

sandwich · Misspecification
Three bands called 95%, at n = 20. Half-widths in sample standard deviations: 0.468 for the mean, 2.14 for one future observation, 2.75 to hold 95% of the population. The first two differ by exactly the square root of n + 1, which is 4.58 here.

A tenth as wide, and both of them right

The interval for a mean and the interval for one future observation are both labelled 95%, and at a hundred observations one is 10.05 times the other — exactly the square root of n + 1. Read the narrow one as the wide one and it covers a new value 15.7% of the time.

estimated · Bands
Two studies with an R-squared of 0.85. The left study's points sit 0.50 from the line and the right study's 2.00 — a factor of 4.0. Both report an R-squared of 0.85 and, at the same sample size, the same standard error for the slope. The design was chosen to make it so, and it can always be chosen.

The t statistic wearing different clothes

For a simple regression, t² = (n − 2)R²/(1 − R²), exactly, on every dataset — checked to sixteen significant figures over five hundred fits. So a paper reporting R² and a p-value has reported one number twice, and two studies with the same R² have points four times further from the line.

spread · Summary
The pooled two-sample test's size, with a true null everywhere. Forty units split between two groups, with the second group's variance a stated multiple of the first's, and the two population means equal. A 5% test should reject 5% of the time. The pooled test runs from 0.55% to 18.91% across this region; Welch's runs from 4.63% to 5.51%.

A degrees of freedom that is not a count

The pooled two-sample test's size runs from 0.55% to 18.91% across forty units split five ways against five variance ratios, with a true null in every cell. Welch's runs from 4.63% to 5.51% — bought with a degrees of freedom that is a function of the data, not an integer, and not a count of anything.

tails · Student
Least squares estimates persistence low, by an amount with a formula. 3000 series of 50 observations at each persistence. The lower curve is the counted bias of the least-squares estimate of φ, and the open marks on it are −(1 + 3φ)/n, computed rather than fitted. The upper curve is the bias left after adding that quantity back, evaluated at the estimate rather than at the truth nobody has: -0.0020 at φ = 0.3, -0.0023 at φ = 0.5, -0.0039 at φ = 0.7, -0.0059 at φ = 0.8, -0.0108 at φ = 0.9, -0.0165 at φ = 0.95. The formula is a leading-order expression and it understates the bias where the persistence is nearest one — -0.0882 counted against -0.0770 predicted at φ = 0.95, which is the corner of the parameter space every one of these approximations is worst in.

Correcting the persistence

Least squares estimates how much a series remembers of itself as smaller than it is, at every value it can take, by an amount with a closed form. Subtracting that amount back is one line of arithmetic, and what the line costs is variance.

evaluation · Bias
The weights may not read the block they weight. A weighted least squares decomposition needs weights that are constants, or at least independent of the differences they multiply. One λ̂ pooled across the trial is estimated on hundreds of degrees of freedom and is effectively a constant; a λ̂ estimated inside each block is estimated on that block's own two or three, and is correlated with the difference it weights. Coverage falls from 94.68% to 82.76% — and the interval gets wider while doing it, 0.5163 against 0.3024, which is the signature of weights that are noise.

The condition that cannot be dropped

The weights may not read the block they weight. Estimate the variance ratio inside each block rather than across the trial and the coverage falls to 83% — on an interval that is at the same time seventy per cent wider.

corner · Allocation
Four rules of four change sign. The margin between the two block windows in points of coverage, under each of four rules, on each of three intervals built from the same resamples, over 300 draws. Positive is the tapered window covering better. On the percentile interval the taper wins at all four rules, by 5.33, 1.67, 5.00 and 4.00 points, which is the earlier field's own reading. On the studentised interval the rectangle wins at all four, by 4.33, 7.00, 4.33 and 2.67. And a normal interval, which uses no resampling at all, puts the two within a third of a point at every rule — so the disagreement is manufactured entirely by what is done with the resamples.

The ordering reverses again

One field found two of four rules changing sign between two readings of one resampling. Turn the same resamples into a studentised interval instead of a percentile one and all four change sign.

student · Bootstrap
A wrong weight costs width; a random weight costs level. Five weightings on a trial whose variance ratio drifts by a factor of 20.1 between the first block and the last, over 4000 runs. The rule that knows every λ_b covers at 95.1% and sets the width. One ratio for the whole trial is wrong for every block and costs nothing in level — 94.8% — while being 20% wider; equal weights are calibrated by an identity and 22% wider. The ratio estimated inside each block is the only rule aimed at the quantity that actually varies, and it is the only one that misses the level, at 92.0%: a weight computed from a handful of degrees of freedom is mostly noise, and noise in a weight is not a wrong weight. Modelling the drift across blocks recovers the oracle's width at 94.8%.

A ratio that changes between blocks

A wrong weight costs width and a random weight costs level. The rule aimed at the quantity that actually varies is the only one that misses its own coverage, and the rule that models it across blocks recovers the whole of what knowing it is worth.

blocks · Nuisance
One penalty, read along two dials. How much wider the studentised interval is than the percentile one, at every sample size and every block length on the grid, with the number of whole blocks each cell leaves written beneath. Read across a row and the block length changes; read down a column and the sample size does. The penalty is nearly a function of the block count alone: the cells at 15 blocks read 1.16, 1.20, 1.17, 1.15, 1.13, 1.10 across three sample sizes and three block lengths, while the cells at one block length read anything from 1.10 to 3.95. The largest penalty on the grid is 3.95, at the cell with 3 whole blocks in it.

The count or the length

A block length and a block count are one number read two ways at one sample size. Read at three, the studentised interval's width penalty tracks the count — with an R² of 0.9911 against a closed form that has no length in it — and its coverage tracks the length.

student · Bootstrap
What each correction is worth, exactly. Each variance estimate's expectation under a constant error variance, divided by the variance the slope actually has, at 20 rows on an even design, by two routes: the closed form E[eᵢ²] = σ²(1 − hᵢᵢ) carried through each correction's own weight, and the mean of 20000 counted estimates. The maximum leverage here is 0.1857 and the design's fourth-moment share Σu⁴/(Σu²)² is 0.0897, which is the only thing the closed form reads. HC0 comes out at 0.8603 — short by construction, since its factor is exactly 1 − 1/n − Σu⁴/(Σu²)². HC1 reaches 0.9559, HC2 is exactly 1.0000 at every design and every sample size, and HC3 overshoots to 1.1647. On an even design the four are within a fifth of each other and the choice barely matters.

Three corrections and a leverage

On an even design of twenty rows the four robust corrections read 0.8603, 0.9559, 1.0000 and 1.1647 of the truth and the choice barely matters. Add one point at x = 8 and they read 0.3191, 0.3419, 1.0000 and 5.1127.

sandwich · Misspecification
The fixed-width trial's coverage when the outcomes are not normal, for both stopping rules. normal: stopping on the arms 94.05% after 18.1 blocks, on the report 89.95%; log-normal, skewness 0.95: stopping on the arms 94.70% after 18.5 blocks, on the report 90.80%; log-normal, skewness 2.26: stopping on the arms 94.15% after 19.3 blocks, on the report 90.25%; log-normal, skewness 4.75: stopping on the arms 94.45% after 18.7 blocks, on the report 90.50%; t, five degrees of freedom: stopping on the arms 94.35% after 18.3 blocks, on the report 90.30%; skewness 4.75, arm A only: stopping on the arms 93.80% after 26.0 blocks, on the report 89.90%; skewness 4.75, arm B only: stopping on the arms 93.60% after 14.2 blocks, on the report 89.95%; equal variances, normal: stopping on the arms 94.75% after 11.4 blocks, on the report 90.90%; equal variances, skewness 4.75: stopping on the arms 94.05% after 11.1 blocks, on the report 92.00%.

A width rule on skewed outcomes

The blinded fixed-width rule rests on a within-arm spread being independent of the arm means, which only normal samples guarantee. On outcomes with a skewness of 4.75 the independence fails and the overall coverage barely notices — 93.60% to 94.70% across every shape counted, against 94.05% on normal outcomes. What skew moves is the runs that stop by twelve blocks, which cover about 90% with the skew in one arm, and the trial's length: a variance ratio corrected on normal theory lengthens it from 18.1 blocks to 26.0 with the skew in the first arm and shortens it to 14.2 with the skew in the second.

stop · Width
Four 95% intervals for the mean of 30 exponential observations, split by the side they miss on. Both bars should read 2.5%. The t interval misses below the mean on 6.38% of samples. Widened until its total is exactly 5%, it misses below on 4.69% and above on 0.30%. Hall's transformation misses below on 3.31% and above on 1.96%.

The side a bound is read from

On thirty exponential observations the upper limit of a 95% t interval is exceeded by the true mean 6.38% of the time, against the 2.5% a safety margin set from it assumes. Widen the interval until its total coverage is exactly 95% and the upper limit is still exceeded 4.69% of the time. A symmetric repair fixes the number that is reported and not the one that is used; Hall's transformation, which bends the interval, takes the same rate to 3.31%.

tails · Student
The factor that holds all of the next m observations with 95% probability, from a sample of 10. From 10 observations, the band for one future value has factor 2.371, for ten 3.716, for a hundred 4.942 and for a thousand 6.008. The 95%-content tolerance factor is 3.382 and is passed by m = 10; the Bonferroni stretch of the prediction factor reaches 7.567 at a thousand.

All of the next ten

A warranty, a batch release or a monitoring rule promises something about every one of the next ten observations, not about one. From a sample of ten, the band that holds all ten with 95% probability reaches 3.716 sample standard deviations either side of the mean — already wider than the 3.382 of a tolerance interval for 95% of the population — and it keeps widening: 4.942 for a hundred, 6.008 for a thousand, with no ceiling. A 95% prediction interval, read as the answer, holds all ten 67.9% of the time: more than 0.95 to the tenth power, because the ten succeed and fail together.

estimated · Bands
Where the bias lands. The drift in the log variance ratio, fitted across 12 blocks over 4000 trials. E[log λ̂_b] is log λ_b plus ψ(k_B/2) − log(k_B/2) − ψ(k_A/2) + log(k_A/2), which depends on nothing but the degrees of freedom — so the tempting sentence is that it goes into the intercept and leaves the slope alone. It does not, because the blocks alternate between allocations and the alternation is correlated with the covariate being fitted: the lopsided blocks carry 0.5383 of bias and the even ones carry none. Uncorrected the slope reads 1.5597 against a truth of 1.5, which is 8.0 standard errors. Subtracting the two digammas block by block leaves 1.4976.

The bias that lands in the slope

The bias in a log variance estimate depends on nothing but its degrees of freedom, so it goes into the intercept — unless the degrees of freedom alternate with the design, which is exactly what a block-randomised trial makes them do.

blocks · Width
The price of insurance is noise, not coverage. What a robust standard error costs under a constant error variance — the case where the model-based one is exactly right — at five sample sizes over 20000 draws at the small end. It is not coverage: the leave-one-out interval read against a t on n − 2 covers 95.52% at 20 rows against the model-based 95.06%. It is a wider interval, by a factor of 1.0689 at 20 rows falling to 1.0049 at 250, and it is a variance estimate 2.57 times as variable at 20 rows and 1.85 times at 250 — against a denominator that is exactly V²·2/(n − 2), so only the numerator is counted. The uncorrected estimate is the cheaper of the two at small samples and the dearer at large: 1.31 against 1.76.

Right for the wrong reason

A robust standard error costs no coverage where the risk is absent — 95.52% against 95.06% at twenty rows. It costs a 6.89% wider interval and a variance estimate 2.572 times as variable, and the pre-test that would avoid paying recovers 15.9% of what the insurance is worth.

sandwich · Misspecification
Four intervals at 3 blocks of 32 rows. What four 95% intervals for the mean of a first-order autoregression at 0.7 cover, and how wide they are on average, at 120 rows cut into 3 whole blocks of 32, over 240 draws with 200 resamples each under the rectangle. The normal interval, the block-means variance with 1.96, covers 82.1% at a width of 0.708. The percentile interval covers 82.1% at 0.632 and the studentised one 90.4% at 2.495. The fourth resamples nothing: it is the normal interval with 1.96 replaced by Student's t on 2 degrees of freedom, and it covers 94.2% at 1.555, 0.62 times the studentised interval's width.

The interval with no resampling in it

Replace 1.96 in a normal interval on the block-means variance with Student's t on one fewer degrees of freedom than there are whole blocks, and resample nothing. Across twenty-four cells it covers at least as often as the studentised bootstrap interval at every one, by 0.42 to 10.42 points; it is narrower wherever seven blocks or fewer are left; and at fifteen blocks of 32 it covers 95.0%, which no resampled interval on the grid reaches.

student · Bootstrap
Student's t on 5 degrees of freedom, against the normal. The two-sided 95% critical value is 2.571 for t(5) and 1.960 for the normal — 31% wider. Using the normal at this sample size makes every interval too short by that much.

The correction for not knowing the spread

The t distribution exists because the standard deviation is estimated rather than known. At eight observations, using the normal instead makes every interval 12% too short — and the coverage that follows can be measured rather than argued about.

intervals · Student
What the forecast interval is short by, φ = 0.85, 6 steps ahead. The plug-in interval covers 88.42% against a claimed 95%. Correcting the variance recovers 0.56 points, propagating the persistence's own standard error recovers 0.40, correcting the persistence recovers 2.66, and all three together recover 4.20 — leaving 2.38 points unaccounted for.

What the interval is short by

The forecast interval covers 88.42% where it claims 95%. Correcting the persistence recovers 2.66 points, correcting the innovation variance 0.56, propagating the persistence's own standard error 0.40 — and all three together recover 4.20 of the 6.58, leaving a residual none of the standard repairs reaches.

evaluation · Bias
One estimator, three answers, and only the reference changes. Coverage of the cluster-robust 95% interval for the slope against the number of clusters, at 30 rows in each. The estimator is identical in all three curves; what differs is the number it is compared against. At 5 clusters it covers 75.05% against a normal, 85.30% against a t on 4 degrees of freedom and 87.95% against a t on 3. At 80 clusters the three agree to within a point. The correction costs nothing: the same standard error, a different table.

The reference the sandwich is read against

The cluster-robust interval covers 75.05% at five clusters and 93.58% at eighty. The same estimate read against a t on G − 2 covers 87.95% at five, and the estimator is unchanged — three hundred rows grouped into five clusters cover 74.28% where the same three hundred grouped into seventy-five cover 94.63%.

sandwich · Misspecification

Named alongside it

The objects these essays reach for when they reach for this one.

CoverageDegrees of freedomInterval widthMonte CarloSample sizeVariance ratioBlindingConfidence intervalFixed-width intervalClosed formDependenceBlock bootstrap

All concepts