Concept

Dependence — where it appears

That observations carry information about each other, which every formula dividing by the square root of a sample size assumes away. It reduces the number of independent things a procedure is counting, and both an optimism penalty and a bootstrap fail to the same arithmetic when it is present.

Named by 63 essays across 21 fields — each of them below, with the objects they name alongside it.

Four blocks, and only the ends matter. The weights a block carries, normalised so that the resample keeps the residuals' variance and only their covariances are attenuated. What separates these shapes, for everything that follows, is the value at the two ends and nothing about the middle: the first-order attenuation is −(w(0)² + w(1)²)/(2∫w²), which is -1.0000 for the rectangle, -0.3896 for the trapezoid cut off at half height, and exactly zero for both windows that reach the axis. The half-height trapezoid is in the table to be the case that separates a shape from a boundary value: it is smooth, it is tapered, and it buys none of the order the other two buy.

A block weighted inside itself

The triangle every block resample attenuates by is not a fact about blocks. It is the self-convolution of a rectangle, and a block weighted down towards its own ends has a different one — whose leading term is the squared value at the two ends and nothing else about the shape.

taper · Bootstrap
The price of each thing the rule is not told. What each rule gives up against the best model available, at a persistence of 0.85 on a fifteen-candidate table, over 400 draws. Reading down: least squares with the ordinary penalty; the whitening at the true ρ; the same at a ρ̂ estimated per candidate; that rule with the term the Gaussian likelihood carries and it omits; a Bartlett-tapered Ω̂ estimated once from the fullest candidate at L = 8; the same estimated per candidate; and the truncated Ω̂, which exists on only 45.0% of draws and is averaged over those. Knowing ρ recovers 89.9% of what counting rows gives up, estimating it 83.4%, and estimating a whole covariance 75.6%.

A covariance with no parameter in it

The whitening that repairs a criterion is told the dependence is a first-order autoregression and left to find one number. A real dependence is not one number, and the obvious estimate of it is not a covariance matrix.

banded · Dependence
The expansion that never terminates. The Hermite coefficients of a median split, in magnitude, against the reference j to the power −3/4, anchored at the first one. Every even order is exactly zero because sign is an odd function, and every odd order is not, so no truncation is exact — where a polynomial of degree d is exact at any order past d. Summed, the tail past J falls like 1/√J: sixty orders still leave 6.6% of the variance outside. That statement is what made a cut dictionary's geometry unavailable in closed form, and it is a statement about the function against itself. What it is not is the accuracy of an inner product between two correlated variables, where every term past J carries a factor of ρ^m as well.

A cut is not a polynomial, and it does not have to be

A threshold's expansion never terminates, which is why a balancing dictionary's geometry was closed for powers and taken to draws for cut points. Conditioning on the second variable closes it for both.

splits · Blocking
Where the general fit becomes the parametric one. The band family's objective at the autoregression's own geometric sequence, cut off at each width, on one sample of 60 rows. The horizontal line is the profile likelihood the parametric fit maximises, written independently through a different whitening. At the full width the two are the same number to 3e-14, which is what says the general construction contains the parametric one rather than resembling it. Below 10 lags there is no line at all: the geometric sequence cut off short is not a covariance matrix, so the objective has nothing to evaluate. Between the two the truncation is briefly above the parametric likelihood — a wrong covariance can fit one sample better than the right one, which is the whole reason a width has to be charged for rather than chosen.

A family before a fit

A regression's coefficients and one correlation can be maximised together. Replace the correlation with an estimated covariance and there is nothing left for "jointly" to mean — until a set of covariances is named, and the set turns out not to contain the truth.

family · Dependence
The correction is not a property of the sample. tr(HΩ)/q for each of fifteen candidates, at ρ = 0.7. Two candidates that fit the same number of coefficients need corrections that differ by as much as 1.49, because one of them is fitting the persistent predictors and the other is not — so no single number can be right for both, and the scalar n/n_eff = 5.537 is above every one of them. The four predictors carry persistences 0.9, 0.6, 0.3, 0; at one persistence for every column the whole spread collapses and a scalar looks exactly as good as the trace.

A penalty is a trace

Akaike's 2q is not a count of coefficients. It is the answer a trace collapses to when the rows are independent — and once they are not, the trace is still the right object and is no longer the count.

effective · Order-selection
Three intervals, one shortfall. What each of three intervals actually covers, at four rules and two block windows, over 300 samples of 120 rows. All three are built from the same resamples on the same draws, so a difference between them is a difference in what is done with the resampled series. Not one of the twenty-four cells reaches the ninety-five per cent it promises. The studentised interval runs from 75.7% to 92.3%, the percentile interval — the earlier field's — from 80.0% to 89.7%, and a normal interval on the same scale from 81.7% to 89.0%. The standard repair for a percentile interval's shortfall does not repair it.

An interval that carries its scale

A percentile interval inherits the resampled distribution's skewness and its scale error together. The standard repair is one extra variance per resample. It was named and not run, so this runs it.

student · Bootstrap
The number the comparison was missing. What it costs to choose the tuning parameter for every candidate separately rather than once for the table, under AR(1) at 0.8, paired on the draw. The window's figure is the one the earlier field reported; the order's is the one it named and did not make. They are the same size — 0.00401 against 0.00360, at 2.30 and 1.72 paired standard errors — and matching the lists at eight values leaves them the same size again. The prediction that the longer list would make the order's cost the larger of the two is not what happens; what happens is that the two rules cost the same once they are scored by the same criterion, which took a missing term to arrange.

The comparison that was not made

Choosing a whitening's window separately for every candidate costs 0.00401 of regret. The same question about an order was named and left, because the two lists are different lengths. The order's answer is 0.00360, and matching the lists changes almost nothing.

lists · Order-selection
The eighth was not a constant. The probability that a per-candidate tuning list changes which candidate the table selects, at each of 5 separations between the candidates, over 800 draws apiece. The earlier field reports this flat at about an eighth across list length, on a table and a world it never varies. Vary how much the omitted coefficients are worth — one multiplier, with the table, the list, the law and the sample size all held — and it runs from 17.6% to 1.5%, a factor of 11.75. The world in which every candidate is true is the world in which the tuning list decides most; the world in which one candidate dominates is the world in which it decides nothing.

The eighth that was not a constant

How often a per-candidate tuning list changes which candidate wins is reported flat at about an eighth across list length. Vary how far apart the candidates are instead and it runs from 17.6% to 1.5%.

turnover · Order-selection
A quantile is the dearer reading, everywhere. The error each rule and window delivers on the two error readings, over 400 draws. The lower pair of lines is the implied long-run variance — the instrument the earlier field uses — and the upper pair is the 95% point of the standardised resampled mean, read against the finite-sample truth of 3.889 found by simulating the law directly. The quantile costs more at every one of the eight cells: at the plug-in rule it is 59.1% against 45.3% for the taper. That is not a defect in the bootstrap; a quantile is a statement about the shape of a distribution as well as its scale, and a fixed number of resamples estimates a tail worse than a variance. What matters for the comparison is that the two orderings between the windows are not the same, which the margins figure is about.

The instrument and the reading

Every comparison between two block windows in this collection is an error in an implied long-run variance. Nobody reads a long-run variance. Read on the 95% point a test uses, the same bootstrap costs half as much again.

readout · Bootstrap
Three rules and a target none of them is aimed at. Which block length each rule picks, over 400 samples of 120 rows, for the tapered window. Two of the rules are points: a length written into a protocol is 8.00 on every draw and the rule of thumb is 4.00, because n to the one third does not read the data at all. The plug-in reads the sample's own persistence and lands at 14.36 with a standard deviation of 2.93. The length that would actually have been best on that draw averages 24.57 with a standard deviation of 16.23 and runs from 10 to 48 between its tenth and ninetieth percentiles. The target moves five times as much as the best estimate of it does, which is why no rule can be close to it and why the two that do not try are not merely worse — they are somewhere else.

The length nobody has

Every comparison of block windows in this collection is made at each window's own best block length. That length has a standard deviation of sixteen across draws and averages twenty-five. No rule is aimed at it.

feasible · Bootstrap
Twenty series with a lag-one correlation of 0.8. Every series has a true mean of zero and 60 observations. The marks on the right are the twenty sample means. The variance of that mean is 8.3 times what 60 independent observations would give, so the series is worth about 7 of them.

The observations that repeat each other

Almost every standard error divides by √n, which claims the observations carry independent information. At a lag-one correlation of 0.8 a fifty-point series is worth about six independent observations, and its 95% interval covers 47%.

timeseries · Dependence
The curvature is in the denominator. The optimism a Bartlett band of each width actually costs, divided by that width, on two ways of measuring the width, over 2000 draws at 120 rows. Measured in the weights the band spends — Σ w(k), which is what the earlier field levies its charges on — the reading falls from 0.9528 at two lags to 0.7486 at thirty, so a charge proportional to the summed weights is too dear at one end and too cheap at the other. Measured in the pairs the band uses — Σ w(k)(1 − k/n), because a lag of k is an average over n − k products — the same readings are flat from 4 lags up, at 0.0084 of χ² per width against 0.2359. The correction has no fitted parameter in it: it is a function of the window, the width and the sample size.

The width a band is measured in

A tapered covariance band spends 84% of its own weights at two lags and 74% at thirty. Every charge in the collection is a straight line through the origin in those weights, so it is too dear at one end and too cheap at the other.

curve · Criterion
What each rung is made of. Each pair of searches, over 300 draws, split into the two effects its excess is the difference of. The overlap is what the second search loses by having the first already run at its own answer; the interaction is what the joint search finds by moving the first off it. They subtract to the excess exactly, on every draw, because the pinned supremum cancels. Two disjoint dictionaries of independent columns read an excess of 0.000011 and are made of 0.000514 and 0.000503. A break paired with a dictionary of step columns has an interaction of exactly 0 and is all overlap. And a break paired with an independent column has an overlap of -0.004395 against an interaction of 0.002364, which is what puts its excess below zero.

Two effects in one number

How much two searches over one sample share is measured as the net of two things — ground both of them find, and configurations only the joint search reaches. One extra supremum per draw separates them exactly.

separate · Break point
Two searches find some of the same luck. What each search reports on a sample with no break in it, and what the two report together, on four dependences. The dashed line is the sum of the two — what a rule charging each search separately would levy — and the two together always come in below it: 19.30, 15.44, 14.24, 26.64 short, on 100%, 99%, 99%, 100% of draws. The shortfall is not a rounding. Under AR(1) at 0.8 it is 19.30 of the 34.70 the break search manufactures on its own, which is more than half of it. Two searches over one sample are looking at the same noise, and the second one has less left to find.

Two searches, one sample

A searched break in a regression manufactures 34.7 of likelihood ratio where a count of coefficients says 11.1. A searched window manufactures 84.0. The two together manufacture 99.4, not 118.7.

twice · Break point
The familywise error rate with no correction, α = 0.05. Two routes: the curve is 1 − (1 − α)^m and the points are counted over 6,000 simulated families of true nulls. With twenty tests the chance of at least one false positive is 64.1%.

What the correction corrects

Twenty tests of true nulls produce at least one false positive 64% of the time, and the closed form and the count agree. Bonferroni holds it at 5% and Holm holds it at 5% while finding more. Nobody should still be using Bonferroni.

multiplicity · Multiplicity
8 exponential draws, standardised, against the normal. The source is one-sided and skewed. At n = 8 the standardised sum has skew 0.695, and the theory says 2/sqrt(n) = 0.707 — so the convergence is visible AND its rate is predicted.

Sums of almost anything

The theorem says sums converge on one shape whatever they are sums of, which is remarkable and true. Watching it happen from a one-sided skewed source, with the rate of convergence predicted in advance, is more convincing than watching the shape appear.

normal · Clt
The curvature is in the denominator. The optimism a Bartlett band of each width actually costs, divided by that width, on two ways of measuring the width, over 2000 draws at 120 rows. Measured in the weights the band spends — Σ w(k), which is what the earlier field levies its charges on — the reading falls from 0.9528 at two lags to 0.7486 at thirty, so a charge proportional to the summed weights is too dear at one end and too cheap at the other. Measured in the pairs the band uses — Σ w(k)(1 − k/n), because a lag of k is an average over n − k products — the same readings are flat from 4 lags up, at 0.0084 of χ² per width against 0.2359. The correction has no fitted parameter in it: it is a function of the window, the width and the sample size.

A lag the sample has less of

A sample autocovariance at lag k is an average over n − k products, not n. Count a band's width in the pairs it actually has and the curvature in its charge goes away, on a correction with nothing fitted in it.

curve · Criterion
Two instruments, two block lengths. The block length that would actually have been best on each draw, for each of the two error readings, averaged over 400 samples of 120 rows. For the rectangular window the implied long-run variance wants 18.92 and the 95% point wants 16.05; for the tapered window, 21.82 against 17.74. The quantile wants a shorter block under both windows — a ratio of 0.848 and 0.813. That is the mechanism the whole field turns on: a rule for choosing a block length is a way of guessing a target, and the two instruments do not have the same target. A rule tuned to one is systematically long for the other, and the two windows do not pay the same price for being long.

A length for each instrument

The block length that is best for an implied variance is 18.92; the one best for the 95% point of the same resamples is 16.05. A rule is a way of guessing a target, and there are two targets.

readout · Bootstrap
How far each reference distribution's 95% point falls short. Seven constructions on rows that repeat each other, at a block length of 5 and 200 draws, against the statistic's own 95% point of 3.0224 computed from three thousand draws of the same world. Reading down: a multiplier on every row keeps no dependence at all and is 44% short; a multiplier shared along a block keeps the triangle; a fixed block keeps the same triangle and is 8% closer, which is the pair that says a taper is not what decides this; the stationary bootstrap; the two tapered blocks, both further short than the untapered one at this block length; and errors generated from a fitted model, which is the only construction here not bounded by what the residuals report.

A taper and a critical value

Two constructions whose tapers visibly differ give the same critical value, and two that share a taper exactly do not. Adding a construction whose taper is a decision rather than an accident says which half of that is true.

taper · Reference
Which window is better depends on who chose the block length. The margin between a rectangular block and a tapered one, on 400 samples of 120 rows, under four rules for choosing the block length. At the length that would actually have been best on each draw the taper is ahead by 2.12 points of a 35.8% error, at 22.1 paired standard errors; at a length estimated from the sample's own persistence it is ahead by 1.89. At the length this field's own figures use — eight — the rectangle is ahead by 2.04, and at the rule of thumb by 4.54. Every rule sees the same draws. What separates them is the length: the two rules that lose to the rectangle pick 4.00 and 8.00 where the best available is 24.57, and a tapered window at a quarter of the right length has thrown away most of what it was weighting.

An ordering that depends on the rule

The tapered block beats the rectangular one at the best available block length and at one estimated from the data. At a length written into a protocol, and at the rule of thumb, the rectangle wins — at every sample size measured.

feasible · Bootstrap
The one candidate an effective sample size is right about. n/n_eff with the finite-sample inflation Σ(1 − |k|/n)ρ^|k| is not an approximation to tr(HΩ) for a fit with only an intercept — it is that trace, to machine precision, because the hat matrix of a constant column is 1/n everywhere and its trace against Ω is the mean of Ω. The quoted limit form n(1 − ρ)/(1 + ρ) is not even right about that one. And the average correction the table's fifteen candidates actually need is 3.318 per parameter, well below the scalar, so applying it to all of them over-charges every one.

One number for a table of candidates

An effective sample size is a real quantity, it is exactly right about one thing, and that thing is a mean. Substituted into Akaike's criterion it changes nothing at all, because the penalty it is meant to fix has no sample size in it.

effective · Dependence
Sheppard's arcsine, by two routes. Corr(sign X, sign Y) as the covariates' correlation runs from zero to one, drawn twice. One route is a sixty-four-node quadrature of the orthant probability over the correlation — the general construction, which works at any pair of cut points; the other is (2/π) arcsin ρ, which is elementary and works only at the median. They agree to 3.3e-16 at every one of 81 correlations, which is what licenses the quadrature everywhere else. The curve is above the diagonal at small ρ and below it at large: two signs agree with probability ½ + arcsin(ρ)/π, so a correlation of 0.5 gives exactly ⅓ and a correlation of 0.8 gives 0.5903.

The arcsine that closes it, and the error that was overstated

Two median splits of a correlated pair agree with probability ½ + arcsin(ρ)/π, exactly. And the truncation the field was avoiding falls geometrically in the correlation, not algebraically in the order.

splits · Routes
Three charges, and only one of them is a test. What each of three thresholds does to the same decision, under AR(1) at 0.8, against the size of a genuine break in the mean at row 60. A chi-square on the 5 coefficients a split adds — 11.07 — declares a break on 73.6% of samples that have none: it is not a test at all. The break search's own 95% point, 69.6, carried into a rule that also chooses its window, fires on 0.0% of null samples and on 0.0% of samples with the largest break measured — the natural way of combining two published corrections does not lose a little power, it switches the test off. The calibrated charge, 27.2, holds 5.6% at no break and reaches 29.2% at the largest.

The charge that is not a sum

Charging two searches what each costs on its own is conservative, and conservative here means the test never fires. At the largest break measured it declares nothing, on every draw, while a calibrated threshold reaches 29%.

twice · Break point
How much memory a fit takes out, candidate by candidate. Under AR(1) at 0.8, the lag-one autocorrelation a candidate's residuals report, computed exactly for each candidate on 200 draws. The upper line is the law at 0.8000. A candidate that is an intercept alone reports 0.7773 — which is exactly what a sample of 120 errors reports, because an intercept annihilates the sample mean and nothing else, and the two arithmetics agree to the last bit. Every predictor after that takes more out, down to 0.7341 at the fullest candidate. That is the collision this field is about: the rule every whitening here uses estimates its nuisance once, from the fullest candidate, so that the criteria stay comparable — and the fullest candidate is the one whose residuals report the least.

The fit that takes the memory out

A candidate's residuals report less dependence than its errors do, and how much less is arithmetic rather than noise. The rule used for a good reason reads the series that has lost the most.

together · Dependence
Most of the rise is the optimiser's, and under one law it is not. The rise in log-likelihood from the tapered plug-in to the maximum over the same eight-lag band, beside what the same optimiser produces on a sample generated from the plug-in's own covariance — where the family is correctly specified by construction and there is nothing to find. Under AR(1) at 0.8 the raw rise is 5.72 and the manufactured baseline is 4.79, leaving 0.93 at 1.8 standard errors; under long memory the excess is 0.14, at 0.2. Under the moving average it is 11.87 at 19.4 standard errors, on every draw. The taper is a shrinkage, and it costs nothing where the sequence decays smoothly and a great deal where it stops dead.

The plug-in and the maximum

A tapered covariance estimate sits five and a half log-likelihood units below the maximum of the likelihood it is substituted into. Four fifths of that is what the optimiser would have found if nothing were missing.

family · Dependence
The term that cancels, and the term that does not. The volume each candidate's whitening moves — log|Ω̂| — for a sieve of order 4 on one sample of 120 rows. Estimated once from the fullest candidate and used for the whole table, it is the same number for every candidate, so it drops out of every difference the criterion reads: that is why nothing in this collection has ever needed to carry it. Estimated from each candidate's own residuals it ranges over 23.87, which is more than a parameter is worth, and the criteria being compared are then fits made under different error models with no term saying so. The window's rule has carried this term since the estimated-covariance field and the sieve's never had it.

The volume a whitening moves

A sieve's whitening has a determinant and this collection's criterion for it never carried one. Shared across a table the term cancels exactly, which is why nothing ever noticed; used per candidate it is worth more than a parameter and the whole comparison turns on it.

lists · Criterion
A window with two wrong ends. The regret of a rule whitened by a Bartlett-tapered Ω̂, as the window widens, against three rules that need no window at all. At L = 0 the estimate is the identity and the rule is exactly least squares — 0.08556, the same number to five places. It falls to 0.01754 at L = 20 and rises again by L = 30, because a quarter of the sample's lags are then being estimated from it. The automatic bandwidth a practitioner would reach for, 4(n/100) to the power 2/9, which at this sample size is 4, gives 0.02963 — 69% above the best available window. The rule told the dependence is an AR(1) sits at 0.01440 throughout, which is the price of not knowing the form.

The window that has to be chosen, and the term that was dropped

An estimated covariance has a bandwidth in it, and both ends of the dial are wrong for different reasons. The rule a practitioner would reach for is two thirds worse than the best window there is.

banded · Order-selection
Twenty tests, 10 of them real — what each procedure holds. no correction: familywise 40.8%, false discovery 5.3%, power 85%. Bonferroni: familywise 2.8%, false discovery 0.5%, power 49%. Holm: familywise 3.6%, false discovery 0.6%, power 53%. Benjamini–Hochberg: familywise 20.0%, false discovery 2.6%, power 75%.

Two different promises

Bonferroni bounds the chance of any false positive. Benjamini–Hochberg bounds the share of the findings that are false. Both are called correcting for multiple comparisons, and one of them lets the familywise rate reach 20%.

multiplicity · Multiplicity
Two factors, opposite directions. The two factors the cost of a per-candidate tuning parameter is a product of, as the candidates are pulled apart, over 800 draws at each of 5 separations. How often the candidates disagree about the tuning parameter rises from 31.8% to 88.8%; the share of those disagreements that change which candidate the table selects falls from 51.6% to 1.7%. So the setting where the candidates quarrel most about the tuning parameter is the setting where the quarrel matters least, and a sweep that reads the rate and stops has read the factor pointing the wrong way.

Two factors pointing opposite ways

As the candidates on a table are pulled apart, they quarrel about the tuning parameter three times as often and the quarrel decides the winner thirty times less often. A sweep that reads the first factor has read the one pointing the wrong way.

turnover · Order-selection
Two independent random walks, 100 steps. Nothing connects these two series: each is generated from its own independent draws. Regressing one on the other gives a slope with t = -10.9, R² = 0.55 and p = 0.0e+0 — a result that would be reported as a finding by any standard output.

Two walks and a finding

Regress one random walk on another, independently generated, and the slope is significant 76.7% of the time with a median R² of 0.17. Nothing connects the two series, nothing in the output says so, and more data makes it worse.

timeseries · Spurious
Both halves grow; the difference does not. The control pair's two components and their difference, against how much each of its two searches can find, over 1200 draws at each dictionary size. Two disjoint sets of independent columns are additive at every size — the excess stays inside a standard error or two of zero throughout — and it is not because there is nothing there. The overlap grows from 0.000112 at two columns to 0.000870 at ten, a factor of 7.76, and the interaction grows with it, staying within a factor of two of the overlap at every size. Two searches competing for one residual sum share ground and find configurations neither has alone, in almost equal measure, and their difference is what the earlier field's scale calls zero.

What a zero is made of

Two disjoint dictionaries of independent columns read an excess of 0.000116 and are made of an overlap of 0.000583 and an interaction of 0.000467. The control the whole scale is anchored on reads zero because two effects cancel.

separate · Break point
What studentising costs. How much wider the studentised interval is than the percentile one, cell by cell, over 300 draws, with what each cell gains in coverage beside it. Averaged over the eight cells the interval is 2.09 times as wide and covers 0.46 points better. At the two rules that choose short blocks the two intervals are within a fifth of each other; at the oracle's length, where a resample holds two or three whole blocks, the studentised interval is 4.37 and 5.04 times as wide. A repair that doubles the width and buys half a point is not one a reader could not have had by widening the interval it replaced.

What studentising costs

Averaged over eight cells the studentised interval is 2.09 times as wide as the percentile one and covers 0.46 points better. At the block lengths the rules choose, the scale it divides by rests on two or three numbers.

student · Bootstrap
What a search costs is not a property of that search. The likelihood ratio a searched break in the regression reports, two ways, on every law. On its own — the whole rule being a split of the sample, at no whitening — it averages 34.70 under AR(1) at 0.8, against the 11.07 a chi-square on the five coefficients a split adds would use as a threshold. Inside a rule that also chooses a window from a list of eight, the same search adds only 15.40 — less than half. Most of what a break search finds under correlated errors is the correlation, and a whitening chosen from the same sample has taken it already. A charge measured for one search, carried into a rule that makes two, is not conservative in some harmless direction: it is measuring a different quantity.

A charge that depends on the rule

The break search's charge is 34.7 on its own and 15.4 once a window has been chosen from the same sample. Most of what a break search finds under correlated errors is the correlation, and a whitening has taken it already.

twice · Break point
A line in the right width beats two curves. How far each candidate charge sits from the measured optimism across the plateau, in units of each width's own standard error, over 2000 draws. The straight line through the origin in the band's summed weights — which is what the earlier field levies — misses by 0.2382 per width. The same straight line in the pairs the band actually uses, Σ w(k)(1 − k/n), misses by 0.0095. A fitted power law misses by 0.0293 and a fitted decaying rate by 0.0172, both on one fitted constant more. The deferral this field answers asked for a curve; the answer is a line, in a variable with nothing fitted in it.

A line that beats two curves

A deferral asked for a curve. Fitted against the same measurements, a straight line in a variable nobody had to fit describes the plateau better than either curve does with a constant more — and for three windows out of four it does not.

curve · Criterion
What a longer list actually changes. How often the five candidates choose different tuning parameters, at a true null where every one of them contains the truth, so a disagreement is manufactured rather than discovered. The order's list is an interval of integers, and thinning it moves the rate smoothly from 0% at two values to 39% at thirteen. The window's is not an interval — it runs 0, 1, 2, 4, 8, 12, 20, 30 — so a thinned window list jumps depending on whether it happens to keep the width the criterion wants, between 0% and 42% with no order to it. So "the same length" was never quite the same thing for the two rules, and it is a smaller effect than the field it was invoked to explain.

A list is not a rule

How often five candidates disagree about a tuning parameter runs from nothing at two values on the list to two draws in five at thirteen. What the disagreement costs does not move at all.

lists · Order-selection
The split depends on the order. How much two searches share, measured both ways round, over 300 draws. Pinning the first search at its own answer and searching the second gives one overlap; pinning the second and searching the first gives another. A break paired with an independent column reads -0.004395 one way and 0.002163 the other, at 8.70 paired standard errors and on opposite sides of zero. The excess the two components subtract to is the same in both orders by construction, so what changes is only how it is attributed. There is no order-free way to say which of two searches found ground both can reach, and the two orders bracket it.

A split that depends on the order

Run the second search first and pin that instead, and the same draw gives a different overlap and a different interaction — with the same difference. And one pair has no second order at all.

separate · Break point
Flat along a row, apart between them. The probability that a per-candidate tuning list changes the winner, at three list lengths on three candidate tables, over 800 draws in each of the nine cells. Along a row — the reading the earlier field takes — it moves by a factor of at most 1.21, so that field's invariant survives on every table. Down a column it moves by up to 1.98. The list length is the dial that does not move this number and the table is one that does, and the earlier field varied only the first.

A table and a list

A nested ladder of candidates differing by one coefficient was predicted to turn over more often at every list length. It turns over less at every one, and its list changes the winner half as often.

turnover · Order-selection
Least squares estimates persistence low, by an amount with a formula. 3000 series of 50 observations at each persistence. The lower curve is the counted bias of the least-squares estimate of φ, and the open marks on it are −(1 + 3φ)/n, computed rather than fitted. The upper curve is the bias left after adding that quantity back, evaluated at the estimate rather than at the truth nobody has: -0.0020 at φ = 0.3, -0.0023 at φ = 0.5, -0.0039 at φ = 0.7, -0.0059 at φ = 0.8, -0.0108 at φ = 0.9, -0.0165 at φ = 0.95. The formula is a leading-order expression and it understates the bias where the persistence is nearest one — -0.0882 counted against -0.0770 predicted at φ = 0.95, which is the corner of the parameter space every one of these approximations is worst in.

Correcting the persistence

Least squares estimates how much a series remembers of itself as smaller than it is, at every value it can take, by an amount with a closed form. Subtracting that amount back is one line of arithmetic, and what the line costs is variance.

evaluation · Bias
Four rules of four change sign. The margin between the two block windows in points of coverage, under each of four rules, on each of three intervals built from the same resamples, over 300 draws. Positive is the tapered window covering better. On the percentile interval the taper wins at all four rules, by 5.33, 1.67, 5.00 and 4.00 points, which is the earlier field's own reading. On the studentised interval the rectangle wins at all four, by 4.33, 7.00, 4.33 and 2.67. And a normal interval, which uses no resampling at all, puts the two within a third of a point at every rule — so the disagreement is manufactured entirely by what is done with the resamples.

The ordering reverses again

One field found two of four rules changing sign between two readings of one resampling. Turn the same resamples into a studentised interval instead of a percentile one and all four change sign.

student · Bootstrap
The row count entered twice, and a penalty is one place. Regret against the best available model as the errors are made persistent. Counting rows more than quadruples; both penalty repairs — the trace, and the scalar effective sample size — are worse than it at every persistence measured; and whitening the sample and keeping the ordinary penalty falls, recovering 86.9% of what counting rows gives up at ρ = 0.85. Doing it at an estimated ρ recovers 80.6%, so having to estimate the dependence from the rows being selected on costs 7.3% of what knowing it is worth. Mallows' forms are drawn beside the logarithmic ones and behave the same, which is what rules the linearisation out.

The repair that was exact and made it worse

A penalty computed from the trace is exactly the optimism it estimates, and selecting with it gives up a fifth more than not correcting anything. The row count entered the criterion twice, and a penalty is the second place.

effective · Forecast
The reversal is a property of the instrument. The margin between the two block windows under each of four rules, on three readings of the same resampled means, signed so that a positive bar is the tapered window winning. On the implied long-run variance the taper wins at the best available block length and at one estimated from the sample and loses at a length written into a protocol and at the rule of thumb — which is the reversal the earlier field's whole argument turns on, at 1.48 and 4.52 points. On the 95% point a test actually reads, the taper wins at all four, by 6.13 to 7.08 points. On the coverage the interval actually delivers, the taper wins at all four again, by 2.50 to 5.75 percentage points. Two of the four rules change sign between the first reading and the other two, and the two that change are exactly the two the earlier field's recommendation is about.

The reversal that was the instrument's

On an implied variance the rectangle wins at a protocol length and at the rule of thumb. On the 95% point a test reads, and on the coverage an interval delivers, the taper wins at all four rules.

readout · Bootstrap
Two constructions on one triangle, and a third that is not. Three resamplings that all keep runs of neighbours, on the same residuals at a block length of 5, with the lags running past ℓ so that the tapers separate. A blocked multiplier never moves a residual; a fixed-length moving block moves every one; and they attenuate identically, worst gap 1.4 standard errors, both sitting on γ_resid(k)(1 − k/ℓ)⁺ and both exactly zero past ℓ — so the attenuation is the block boundary rather than the multiplier. The third is the stationary bootstrap, whose runs are geometric rather than fixed: its taper is γ_resid(k)(1 − 1/ℓ)^k, it agrees with the other two at the first lag and at no other, and at lag 6 it still carries 0.0081 where they carry -0.0005.

The triangle that was not the multiplier's

A resampling that leaves each residual on its own row can keep only what the residuals have, times a triangle. A construction that moves every one of them has the same triangle — and the one in this collection's own table has a different taper entirely.

banded · Bootstrap
The guarantee that survives a correlation, and the one that does not. What a balancing rule handed both main effects removes of the pure interaction between them, as the covariates become dependent. For median splits it is exactly zero at every correlation, because sign(x)² = 1: the interaction sign(X)sign(Y) is orthogonal to sign(X) and to sign(Y) whatever ρ is. For the product of the raw covariates it is 4ρ²/(1+ρ²)² — 64.00% by ρ = 0.5, rising to all of it at perfect correlation. A cut away from the median sits between them and is not small: 23.01% at a cut of one. The zero is not a fact about interactions. It is a fact about a dictionary whose functions square to a constant, which a polynomial one does not.

The zero that survives a cut

A rule holding both main effects removes half of a pure interaction between correlated powers and exactly none between correlated median splits. The guarantee that a correlation destroyed was never about interactions.

splits · Criterion
The argument is a third of the size of the thing it is inside. Three quantities on one scale, in points of the error in a block resample's implied long-run variance, at 120 rows. The gap between the two windows at the best available block length — the whole subject of the comparison this field inherited — is 2.12 points. What the best rule a practitioner could actually run gives up against that same best length is 7.26, a factor of 3.42. What the rule of thumb gives up is 26.01. So the ordering between windows is worth establishing and is not worth arguing about, and the sentence that follows from it is not use the taper but estimate the block length, because that is where the points are.

What choosing the length costs

The gap between two block windows at the best available length is 2.12 points. What the best rule a practitioner could run gives up against that same length is 7.26. The argument is a third of the size of the thing it is inside.

feasible · Bootstrap
What differencing fixes, and what it costs, 100 steps. The first pair is the false-positive rate for two independent random walks: 77% on the levels, 4.9% on the differences. The second pair is how much of a real relationship survives: R² falls from 0.91 to 0.33. The same operation does both.

What differencing costs

Differencing takes the false-positive rate between two unrelated walks from 76.7% to 4.9%, and takes a genuine relationship's R² from 0.91 to 0.33. Applied to a series that did not need it, it doubles the variance and installs a correlation of −0.5 that the data never had.

timeseries · Spurious
What each construction carries, against what there was. The autocorrelation of a resampled error series at five lags, averaged over 60 samples of 40 resamples each. Three facts are in the picture. The residuals lie below the errors at every lag, which is the ceiling a multiplier cannot exceed. The blocked multiplier and the fixed-length block lie on top of each other below it — they attenuate identically, because the attenuation is the join — while the stationary bootstrap, whose runs are geometric rather than fixed, sits above them both. And the sieve is the exception in kind rather than in degree: at lag six it carries 0.0638 where the residuals have 0.0300 and the multiplier has -0.0011, because a fitted model extrapolates past the lags it was told about and a truncated sample sequence cannot.

Errors generated from a fitted model

The one construction that is not bounded by the residuals, because a model extrapolates past the lags it was told about and a truncated sample sequence cannot. It is nearly exact where the only defect is dependence, and it pays for it where there are two.

banded · Reference
An AR(1) at φ = 0.5, 200 observations. The bars are the measured correlations; the curve is φᵏ, which is what an AR(1) must have. The band is ±1.96/√n, where an independent series would stay. The first bar is 0.53 against a band of ±0.14.

The check before the standard error

One number decides whether every interval in an analysis is trustworthy, and the check for it flags a lag-one correlation of 0.5 nine times in ten — and one of 0.2 only one time in five, where the interval already covers 88.6% instead of 95%.

timeseries · Dependence
Where the taper's case begins, and it is not where the algebra says. The block length at which a trapezoidal block's implied variance stops being more biased than a rectangular one's, against the length of the sample. Computed exactly — from the law's own autocovariances, with no sampling in it — the answer is 19.2 and does not depend on the sample at all. What a sample of 120 rows reports is 13.3, and the reported crossing walks out towards the exact one as the sample grows: 13.3, 15.0, 16.4, 18.0. The mechanism is that the autocovariances the window is applied to are themselves attenuated, worst at the longest lags, and the window that discards those lags loses less of them.

The error no window repairs

Every block window's best estimate of a long-run variance is wrong by about forty per cent at a hundred and twenty rows, and the largest part of that is not a bias at all. Choosing the window moves a twentieth of it.

crossing · Bootstrap
A fit takes the low frequencies out of what it leaves behind. The autocorrelation of the errors, of the residuals of a fitted benchmark, and of those residuals rescaled by their own leverage. (I − H) removes the component of the errors lying in a column space that is itself slow-moving, so the residuals are less persistent at every lag — by 5.9% at the first and 26.6% by the fourth. The leverage correction is the standard repair for what a fit does to a residual's size; drawn here against what it does to a residual's dependence, it does nothing.

The residuals are not the errors

A fit removes the part of the errors lying in its own column space, and a persistent design's column space is itself slow — so what is left behind is smoother than what went in, at every lag, by an amount that grows with the lag.

effective · Bootstrap
What 20 clusters of 20 correlated observations do to a 95% interval. Each study has 400 observations arranged as 20 clusters of 20. The lower points are the counted coverage of the usual interval, which treats them as 400 independent observations; the curve through them is 2Φ(1.96/√deff) − 1 with deff = 1 + 19ρ, computed before any data was drawn. At ρ = 0.81 the interval covers 36% rather than 95%. The upper points treat the cluster as the unit and need no variance components at all.

Two levels at once

A third level of grouping adds no new arithmetic and produces one number — the design effect — that decides how many independent observations a clustered study is worth. It is the same quantity the time-series field computes for autocorrelated data, arrived at from a completely different picture.

multilevel · Levels
The scale moves the width; the curve does not. The band width each charge picks, averaged over 150 draws of 120 rows. The two conventions — a unit a lag and half a log n a lag — pick 6.08 and 3.65 lags. The four charges derived from the measured optimism pick 15.05, 15.60, 15.47 and 15.44, against a best width on the draw of 14.13. So the scale a charge is levied on moves the width by a factor of 4.27 and the shape of the charge moves it by 3.6%. None of the six is an estimate of the draw's own best width: the correlations are -0.006, -0.003, 0.017, -0.001, 0.016, -0.004.

What a better charge buys

Four charges derived from the same measurements pick band widths within six per cent of each other and deliver errors within two per cent of the gap any of them leaves. The scale a charge is levied on decides the width; the shape of the charge decides nothing.

curve · Criterion
Fitting them together is worth something under one law. Four fits of the same regression under four dependences: least squares, the two-step plug-in every whitened rule in this collection runs, the coefficients and the band maximised together, and a whitening at the law's own covariance that nobody has. Under the moving average — the one law the band family contains — the joint fit beats the two-step by 0.0077 at 3.3 paired standard errors. Under the autoregression, long memory and the break it is a tie: 0.4, 1.0, 0.3 standard errors. That is the same ordering the likelihood gap gave, arrived at through the coefficients rather than through the objective.

What fitting them together buys

Maximising over the coefficients and the covariance together beats the two-step under one of four dependences and ties under the other three. It is the one the band family contains, and the likelihood said so before any coefficient was compared.

family · Dependence
What the interval actually covers. The coverage of the two-sided interval each rule and window builds, over 400 samples of 120 rows, against the 95% it promises. Not one of the eight reaches it: the best is 91.0% and the worst is 80.8%, on a promise of 95%. So the first thing this instrument says is that the choice between the two windows is a choice inside a range that is already four to fourteen points short, which neither of the other two readings can express at all. The second is the ordering: the tapered window covers better under every one of the four rules, by 5.00, 2.50, 5.75 and 4.00 points — including at a length written into a protocol and at the rule of thumb, where the implied variance says the rectangle wins.

What the interval covers

Eight rules and windows, and not one of them reaches its promised 95%. The range is 80.8% to 91.0%, and the choice between two block windows is a choice inside a shortfall that is four times larger.

readout · Bootstrap
One penalty, read along two dials. How much wider the studentised interval is than the percentile one, at every sample size and every block length on the grid, with the number of whole blocks each cell leaves written beneath. Read across a row and the block length changes; read down a column and the sample size does. The penalty is nearly a function of the block count alone: the cells at 15 blocks read 1.16, 1.20, 1.17, 1.15, 1.13, 1.10 across three sample sizes and three block lengths, while the cells at one block length read anything from 1.10 to 3.95. The largest penalty on the grid is 3.95, at the cell with 3 whole blocks in it.

The count or the length

A block length and a block count are one number read two ways at one sample size. Read at three, the studentised interval's width penalty tracks the count — with an R² of 0.9911 against a closed form that has no length in it — and its coverage tracks the length.

student · Bootstrap
The same dial, on a list that steps by one. The probability that a per-candidate tuning list changes which candidate the table selects, at each of 5 separations, for both tuning parameters at a matched list length of 8. The sieve order runs from 13.1% to 2.6%, a factor of 5.00; the whitening window, from 16.4% to 1.5%, a factor of 11.75. What is held is the number of options, the table, the law, the sample size and the seeds; what cannot be held is the size of a step, since an integer step and a geometric step are different amounts of change. The dial moves both, and it moves them by 2.35 times as much on one as on the other.

A step that is not a ratio

Run the separation sweep on a tuning list of integers rather than a geometric ladder and the two factors still point opposite ways. The invariant does not survive: along a row of integers the probability moves by 2.163 where along the geometric ladder it moves by 1.208.

turnover · Order-selection
The false discovery rate of twenty correlated tests, against the correlation. BH, every null true: 5.08% at 0, 4.86% at 0.3, 3.70% at 0.6, 2.34% at 0.9. BH, 10 of 20 real: 2.55% at 0, 2.53% at 0.3, 2.26% at 0.6, 1.66% at 0.9. BY, every null true: 1.46% at 0, 1.31% at 0.3, 1.03% at 0.6, 0.69% at 0.9. BY, 10 of 20 real: 0.72% at 0, 0.75% at 0.3, 0.64% at 0.6, 0.50% at 0.9. 20,000 families at each correlation.

False discoveries that arrive together

Correlate twenty tests and Benjamini–Hochberg still holds its false discovery rate — 1.66% at a correlation of 0.9 with ten real effects, against 2.55% when the tests are independent. What changes is how the errors come. A family of true nulls reports anything 2.34% of the time instead of 5.08%, and when it does, it reports 16.56 false findings out of twenty.

multiplicity · Multiplicity
The ceiling a multiplier cannot reach past. A wild-type resampling forms e*_t = e_t·w_t with the multiplier independent of the residual, so what comes out has autocovariance γ_resid(k)·γ_w(k) — the residuals' own, multiplied by the multiplier's. Since |γ_w| ≤ 1 the reference distribution's dependence is bounded above by the residuals', and the residuals' is already below the errors'. The two shortfalls compose. For a block of ℓ the multiplier's autocorrelation is exactly the triangle (1 − k/ℓ)⁺, drawn here as the dashed prediction against the realised resamples at ℓ = 5; the bound is attained only at ℓ = n, where the reference distribution is built from one sign.

What a multiplier cannot keep

Two reasons were named for the quarter a blocked resampling falls short, and taking either away makes the gap larger. What is left is a bound — a multiplier can only take dependence out, and the residuals' own is already below the errors'.

effective · Reference
Largest where least is needed. What the pairs correction supplies against what each window's measured profile needs, across this field's plateau, over 2000 draws at 120 rows. Both are stated as the multiplicative rise the charge per unit of width has to take between four lags and thirty. What the correction supplies is arithmetic — (1 − μ(4)/n)/(1 − μ(30)/n), where μ is the mean lag of the weight the band adds — and it runs 1.0795, 1.0580, 1.0456, 1.0539 for the four windows. What the measurement needs runs 1.1076, 1.2928, 1.2296, 1.6550. The two orderings are opposite: the plain Bartlett window has the longest mean lag, so it gets the biggest correction, and the flattest profile, so it needs the smallest. They coincide to 0.9746 of each other, and nowhere else does the correction account for more than 85.0% of the fall.

What the correction assumes

A correction with nothing fitted in it repairs one window of four. The reason is that its size is set by where a window puts its weight and the curvature it must repair is set by something else — and for one window at one sample size the two happen to agree.

curve · Criterion
Storey's estimate of the share of true nulls over twenty thousand families, independent and correlated at 0.6. The true share is 0.5. Independent tests: mean 0.610, spread 0.160, below half the truth in 0.92% of families. Correlated at 0.6: mean 0.609, spread 0.240, below half the truth in 9.33%.

Estimating how many nulls are true

Benjamini–Hochberg at 5% delivers 2.55% when half of twenty nulls are false, because it cannot tell how many are. Storey's estimate of that share, read off the p-values above one half, spends the rest and finds 81.93% of the real effects instead of 74.70% on independent tests. Correlated at 0.9, the same procedure reports a finding in 19.29% of families in which every null is true.

multiplicity · Multiplicity
Four intervals at 3 blocks of 32 rows. What four 95% intervals for the mean of a first-order autoregression at 0.7 cover, and how wide they are on average, at 120 rows cut into 3 whole blocks of 32, over 240 draws with 200 resamples each under the rectangle. The normal interval, the block-means variance with 1.96, covers 82.1% at a width of 0.708. The percentile interval covers 82.1% at 0.632 and the studentised one 90.4% at 2.495. The fourth resamples nothing: it is the normal interval with 1.96 replaced by Student's t on 2 degrees of freedom, and it covers 94.2% at 1.555, 0.62 times the studentised interval's width.

The interval with no resampling in it

Replace 1.96 in a normal interval on the block-means variance with Student's t on one fewer degrees of freedom than there are whole blocks, and resample nothing. Across twenty-four cells it covers at least as often as the studentised bootstrap interval at every one, by 0.42 to 10.42 points; it is narrower wherever seven blocks or fewer are left; and at fifteen blocks of 32 it covers 95.0%, which no resampled interval on the grid reaches.

student · Bootstrap
What 240 observations are worth, by which question is asked. 8 rows and 10 columns with 3 observations in each cell — 240 in all, each belonging to one row and one column, neither nested in the other. the overall mean: variance 0.1703 against a naive 0.0060, a design effect of 28.4 and 8.5 effective observations; a difference between two rows: variance 2.0682 against a naive 0.0960, a design effect of 21.5 and 11.1 effective observations; a difference between two columns: variance 1.0845 against a naive 0.1200, a design effect of 9.0 and 26.6 effective observations.

Two groupings that cross

Pupils belong to a school and to a neighbourhood, and neither is nested in the other. There is then no design effect: the overall mean is worth 8.5 independent observations out of 240, a row difference 11.1 and a column difference 26.6, and which grouping matters depends on the question rather than on the study.

multilevel · Levels
Reading the draw changes what is charged, not what is tracked. The correlation between the band width each rule picks and the best band width on the same draw, over 400 draws. The three fixed charges read -0.069, -0.066, 0.012. The three that read the sample read -0.012, -0.019, -0.041. None of the six is distinguishable from nothing. The statistic the first plug-in reads does vary — the draw's own summed squared autocorrelation runs from 2.06 to 10.19 with a mean of 4.36 — so the failure is not that the charge stopped moving. It is that what it moves with carries no information about which width this draw wanted.

A charge that reads the draw

Three charges built to read the sample track the best band width on their own draw at −0.012, −0.019 and −0.041, deliver more error than the fixed rule they are calibrated to, and pick a width half again as variable. The statistic moves; the answer does not.

curve · Criterion
The rows are held fixed; only the clusters move. Counted coverage of four 95% intervals for a slope, at five cluster counts with the row count held at 300 throughout and a within-cluster correlation of 0.1, over 6000 draws apiece, with the sizes equal. The interval that counts rows covers 53.42% at 5 clusters — a second closed form says 2Φ(z/√D) − 1 = 54.44% for a design effect of 6.900, and reads nothing about clusters at all. The cluster-robust interval read against a normal covers 74.43% there and 94.20% at 100 clusters; read against a t on G − 1 it covers 85.08% and 94.47%. The number of independent things is the cluster count, and every quantity here is blind to how many rows were typed.

The count that is not the rows

Three hundred rows in five clusters of sixty carry 6.9000 times the variance an independent-rows calculation reports, and the interval that counts rows covers 53.42%. The same five unequal sizes laid out two ways give design effects of 9.3158 and 5.4652.

sandwich · Misspecification

Named alongside it

The objects these essays reach for when they reach for this one.

Monte CarloClosed formTaperingLong-run varianceModel selectionBlock bootstrapInformation criterionResamplingAutocorrelationSelection effectDegrees of freedomEstimation error

All concepts