Concept

Eigenvalue — where it appears

A number by which a matrix scales one of its own characteristic directions, unchanged in direction by the matrix acting on it. The design criteria are all functions of the information matrix's eigenvalues, and which function is the whole of the difference between them.

Named by 13 essays across 8 fields — each of them below, with the objects they name alongside it.

A proposal that moves more, refused more often. The two halves of the trade, both exact, on the 410 admissible assignments of twelve units. The integrated autocorrelation time of an imbalance the rule was never handed falls from 7.30 at one swap to 3.97 at three, and the acceptance rate falls with it, from 58.8% to 40.8%. A rejected proposal costs one evaluation and leaves the chain where it was, so acceptance is not the price of anything and the ranking by acceptance is the reverse of the ranking by cost. Past three the family folds: exchanging k of six from each arm is the complement of exchanging six − k, so k = 5 has the same 36 proposals as k = 1 and k = 6 has 1.

A proposal that moves more than two units

The walk's autocorrelation is a fact about its step size and not about its acceptance rate. Exchanging three units from each arm mixes nearly twice as fast as exchanging one, and is refused a third more often.

blocks · Randomisation
The end of the family is the member the algorithm cannot reach. Every member of the Φₚ family optimised on the same 11² candidates, and two eigenvalues of each answer. The upper curve is the smallest eigenvalue of the information matrix — the quantity E-optimality maximises — which rises from 0.0993 at the D end to 0.1994 at p = 64. The lower curve is the gap between that eigenvalue and the next one up, which falls from 0.0611 to 0.0011. A smallest eigenvalue is not differentiable where it is repeated, and the family is driving the gap to zero: the one criterion here whose meaning fits in a sentence is the one whose optimum sits on a corner of its own surface. The candidate grid is on a slider and it answers a narrower question than it looks. On this square region, refining an odd grid from seven to eleven moves nothing at all — the optimum's support is the corners, the edge midpoints and the centre, and every odd grid from five up contains all of them. An even grid has no centre point and cannot reach the answer at any member of the family. On a disc, where the boundary passes through no grid point, refinement does move it.

The family behind the letters

A, D and E are not three ideas. They are three points of one family with a single dial, and running the dial from one end to the other doubles the smallest eigenvalue of the information matrix while closing the gap above it fifty-three-fold — which is the family driving its own last member to the place where it stops being differentiable.

criteria · Criterion
Three series and one relation between them. Above, three series generated from Δy = Πy₋₁ + ε with Π of rank 1. Below, the combination y1 −y2. It stays inside a band of 9.5 while the series themselves travel 28.4. The count of combinations that behave this way is the rank of Π, and it is what every method in the field sets out to estimate.

Three series and a count

A pair of series is either tied together or it is not, so its whole inference is one test with one answer. Three can carry none, one or two relations at once — and the thing being estimated stops being a slope and becomes an integer, read off the gap in a spectrum whose top eigenvalue holds at 0.25 while the rest fall like 1/n.

systems · Rank
Stationary is not the same as convergent. How far each k-swap walk is from uniform after t steps, started at the least balanced admissible assignment of 410. Every one of these chains has a symmetric proposal and rejects by standing still, so every one of them is doubly stochastic and every one preserves the uniform distribution exactly. Only five of the six get there. Exchanging all six units of each arm is a single proposal — the complement — and the admissible set is closed under complement, so the walk takes it every time and oscillates between two assignments for ever: after 160 steps it has visited 1 state and sits 0.9976 from uniform. Its stationary distribution is a fact about the matrix; its limit does not exist.

Stationary is not convergent

A walk that exchanges every unit in each arm preserves the uniform distribution exactly and never gets near it. Every doubly stochastic matrix has the same stationary distribution; only some of them have a limit.

blocks · Randomisation
Four windows, one line, and one that is off it. The optimism measured for each window at a band of 30 lags, against what that window's weights sum to, on 2000 pairs of independent samples of 120 rows. The diagonal is where a window that spent exactly its summed weights would sit. Three of the points are one shape at three levels — the Bartlett window, its square and its cube, whose sums stand in the ratio 6 : 4 : 3 — and they lie on a line through the origin at 0.767 of the diagonal, with 0.033 between the highest and the lowest. Scaling the weights scales the charge by the factor the weights predict, which is what makes the weights the mechanism. The Parzen window has a comparable sum and a different shape, and it sits at 0.871: its weights stay near one over the first few lags, and the first few lags are where the information is. A weight sum treats every lag as equally informative and no sample does.

What a window leaves free

A Bartlett window's weights sum to exactly half its width, which is a candidate for what the band costs. Varying the weights without varying anything else says the weights are the mechanism; varying the shape at the same weight says they are not the arithmetic.

dimension · Criterion
The trace statistic under the null, and the 5% point it needs. 600 systems of 3 unrelated random walks, each put through the reduced-rank regression, with the statistic for "rank ≤ 0" collected. The 5% point is 31.91. There is no standard table to look that up in: the distribution depends on the number of common trends under the null and is not a chi-square, so the value is simulated on one set of seeds and applied on another — exactly the position the pair's residual test was in one field ago.

Counting what is still wandering

The statistic that turns a spectrum into an integer has one name and three distributions. Its 5% point is 8.12, 18.64 or 31.74 depending only on how many series are left wandering under the null being tested — and read against the wrong one of those three, it calls unrelated random walks cointegrated most of the time.

systems · Rank
A class that is a subspace has no guarantee below its own dimension. Each cell is the worst case over every unit-variance function in a class of dimension m, for a rule reading k functions: the smallest squared principal-angle cosine between the two subspaces. Wherever k is less than m the number is zero to machine precision, and that is not a weak guarantee but the absence of one — some direction of the class is orthogonal to the entire basis, and against an outcome in that direction the rule does exactly what a coin does. An experimenter who declines to name the shapes and asks instead to be protected against everything smooth is asking for the cells above the diagonal.

Where the guarantee is exactly zero

An experimenter who declines to name the shapes, and asks instead to be protected against anything in a class, is asking for a number that is not small but zero. Bounding the class is unavoidable, and the two ways of doing it choose different bases.

basis · Randomisation
What the fit calls the shape, against what it is. One eigenvalue held at −3 and the other swept from −2 to 2, so the truth is a maximum on the left and a saddle on the right and the change happens at exactly zero. At an eigenvalue of −0.25 — a genuine maximum — the fit reports a saddle on 26.4% of studies; at +0.25 — a genuine saddle — it reports a maximum on 25.1%. The standard error of a squared coefficient under this design is 0.3791, and the region of confusion is about that wide either side of zero.

The sign the curvature has

A fitted surface reports a maximum, a minimum or a saddle, and the report is a comparison of two estimated eigenvalues against zero. At a true second eigenvalue of −0.25 the fit calls a genuine maximum a saddle on 26.4% of studies, and at +0.25 it calls a genuine saddle a maximum on 25.1%.

surface · Optimum
The bounded error and the unbounded one. How the sequential trace procedure's answer is distributed, against the sample length, for a three-series system with 2 genuine relations. Over-counting — claiming a stationary combination that is a random walk — reads 4.9%, 7.2%, 5.7%, 6.2%, 5.9%, 4.2% across the six lengths, never far from the 5% of a single test. Under-counting reads 69.5%, 40.2%, 14.0%, 0.5%, 0.0%, 0.0%. The procedure is described as a 5% rule and the 5% applies to one of those columns.

The rank is a decision

The sequential procedure's 5% bounds one of its two errors. Over-counting reads between 4.2% and 7.2% at every sample length from fifty observations to three hundred; under-counting reads 69.5% at fifty and 0.0% at three hundred, and nothing in the procedure bounds it.

systems · Rank
The ridge, when the fitted optimum is outside the region. One fitted surface. Its stationary point is at a radius of 2.289 and the fit calls the shape a maximum. The ridge is the best setting at each radius, found by the Lagrange condition (B̂ − μI)x = −ĝ/2; the fitted response rises along it from 59.93 at the centre to 62.10 at the edge. The true optimum is at (0.4, 0.3).

When the best setting is outside the region

On a flat surface at twice the noise the fitted optimum lands outside the experimental region on 24.9% of studies and more than three coded units out on 11.8%. The answer is a ridge — the best setting at each radius, with a closed form — and the two obvious rules for using it turn out to be within four per cent of each other.

surface · Optimum
Where each criterion's optimum puts the information. the D-optimal design's smallest eigenvalue is 0.09927, attained once; the A-optimal design's smallest eigenvalue is 0.16516, attained once; the I-optimal design's smallest eigenvalue is 0.17541, attained once; the E-optimal design's smallest eigenvalue is 0.19999, attained 2 times. A criterion that reads the smallest eigenvalue has no derivative where that eigenvalue is repeated, and the E-optimal design is exactly there.

The criterion with no derivative

E-optimality maximises the smallest eigenvalue of the information matrix, and at its own optimum that eigenvalue is attained twice — which is exactly where the function has a corner. The multiplicative search this field's other three criteria use assumes a derivative that is not there, and stops at 37.2% of the optimum.

optimality · Equivalence
One of these converges and the other does not. Two measurements on the same fits, against the sample length, for a system with 2 genuine relations. The distance from the fitted plane to the true plane falls from 0.1438 at 100 observations to 0.0075 at 1600 — halving with each doubling, which is the 1/n rate this field's estimates converge at. The angle between the leading fitted relation and the leading generating one reads 29.6° and 29.0° at those same lengths, and is flat in between. The plane is an estimate; the relation inside it is not.

A space is not a relation

The fitted plane approaches the true one at rate 1/n — 0.1438 at a hundred observations and 0.0075 at sixteen hundred. The angle between the leading fitted relation and the leading generating one reads 29.6° and 29.0° at those same lengths, and never moves.

systems · Rank
The law is the eigenvalues, and nothing else. The mean and the skewness of n(ĝ − g) at the stationary point, measured over 40,000 draws, against the closed forms ½ Σλ and 2√2 Σλ³ ⁄ (Σλ²)^(3⁄2). The worst disagreement anywhere is 0.028. In one variable the second-order law is a single χ² and its sign is the sign of g″; here it is a weighted sum with the Hessian's eigenvalues as weights, so a bowl and a valley differ in both moments and a saddle has both equal to zero.

A flat point with more than one direction

At a stationary point of a function of several means the second-order law is ½ Z′HZ, so the bias is half the Hessian's trace — 2.008 for a bowl, 5.028 for a valley, and −0.006 for a saddle, where the eigenvalues cancel. The saddle's coverage is the worst of the three.

normal · Clt

Named alongside it

The objects these essays reach for when they reach for this one.

Experimental designMonte CarloClosed formCointegrating rankCommon trendReduced-rank regressionCointegrationCovariate balanceInformation matrixRandom walkSaddle pointStationary point

All concepts