Series

Criterion — the series

33 essays on one idea, from the one that introduces it to the one that assumes the rest.
  1. Where a D-optimal design puts its runs. The D-optimal measure over 121 candidate settings on a square region. It keeps 9 of them and discards the rest, and the 9 it keeps are the settings a catalogue would have offered without any of this arithmetic. What the search adds is the weights: 0.1458, 0.0962, 0.0802, which nine equal runs cannot express.

    A design is a number

    A standard design is taken from a catalogue and then measured. Turn the arithmetic round and a design becomes the answer to an optimisation — and over 121 candidate settings the search keeps nine of them, which are exactly the nine a catalogue would have offered, at weights nine equal runs cannot express.

    part 1 · optimality
  2. Four criteria on 5 designs, on a square region. Each design scored as an efficiency — its value over the best attainable — so four criteria in as many different units sit on one scale where 1 is the optimum. Rows are ordered by D. D picks 13-run exchange; A picks face-centred composite; G picks 13-run exchange; I picks face-centred composite. Every design has been scaled to just fit the region first, because a design run at settings the region does not contain is not a competitor on it. The disagreement is the point: the letter is a choice, and it is almost never reported as one.

    Four letters and two camps

    D, A, G and I are four ways of turning one matrix into one number, and they do not agree. The design that wins on D is the worst thing here on I. And the same two designs swap places entirely when the region changes from a square to a disc — on all four criteria at once.

    part 2 · optimality
  3. Adding a run can make the design worse. The D-efficiency of the best N-run design at each size, against the optimal measure. It is not a rising curve. 13 runs reaches 99.77% and 14 falls to 99.44%, because the optimal weights are real numbers and N runs is an integer approximation to them, so how good a design can be depends on how well N divides. A Wald interval behaves the same way: a larger sample sometimes makes its coverage worse, for exactly this reason.

    The design that has to be integers

    The optimal design is a set of real weights and an experiment is a set of runs, so the theory's answer is never available. Thirteen runs reach 99.77% of it and fourteen reach 99.44% — adding a run makes the design worse per run, and the search that finds it does not always find the same one.

    part 3 · optimality
  4. The end of the family is the member the algorithm cannot reach. Every member of the Φₚ family optimised on the same 11² candidates, and two eigenvalues of each answer. The upper curve is the smallest eigenvalue of the information matrix — the quantity E-optimality maximises — which rises from 0.0993 at the D end to 0.1994 at p = 64. The lower curve is the gap between that eigenvalue and the next one up, which falls from 0.0611 to 0.0011. A smallest eigenvalue is not differentiable where it is repeated, and the family is driving the gap to zero: the one criterion here whose meaning fits in a sentence is the one whose optimum sits on a corner of its own surface. The candidate grid is on a slider and it answers a narrower question than it looks. On this square region, refining an odd grid from seven to eleven moves nothing at all — the optimum's support is the corners, the edge midpoints and the centre, and every odd grid from five up contains all of them. An even grid has no centre point and cannot reach the answer at any member of the family. On a disc, where the boundary passes through no grid point, refinement does move it.

    The family behind the letters

    A, D and E are not three ideas. They are three points of one family with a single dial, and running the dial from one end to the other doubles the smallest eigenvalue of the information matrix while closing the gap above it fifty-three-fold — which is the family driving its own last member to the place where it stops being differentiable.

    part 4 · criteria
  5. Where a Ds-optimal design puts its runs. The Ds-optimal measure over 121 candidate settings on a square region. It keeps 9 of them and discards the rest, and the 9 it keeps are the settings a catalogue would have offered without any of this arithmetic. What the search adds is the weights: 0.2500, 0.1250, 0.0625, which nine equal runs cannot express.

    The two terms anybody wanted

    D-optimality estimates all six parameters of a quadratic as precisely as possible. Nobody wants that. An experimenter looking for a maximum wants the two curvature terms, and the design that gives them is not the D-optimal one — it is a quarter of the runs at the centre, exactly, and the D-optimal design is 75.3% efficient for the question that was actually asked.

    part 5 · criteria
  6. Two designs for one model, and the weights are not equal. Where the runs go, for the same two-parameter model at K = 1 and a ceiling of T = 10. The D-optimal design for both parameters is the familiar one: half the runs at 0.8333 and half at the ceiling. The design for the half-saturation constant alone moves the lower setting down to 0.6040 — where the response curve is still bending, which is where K is visible — and, unlike every design in the two fields before this one, it does not split the runs evenly: the weights are exactly 1/√2 and 1 − 1/√2, 0.7071 and 0.2929, at every K, V and T. The equal weights of D-optimality were a consequence of asking about both parameters at once, and nobody had to notice while that was the only question being asked.

    An efficiency that is a ratio

    A design chosen for a model is not a design chosen for the parameter somebody wanted. Asking for one of two parameters moves the runs, unbalances the weights, and costs the other question exactly 15.07% — at every setting, because it is algebra.

    part 6 · guarantee
  7. Four designs, and what each of them guarantees. The worst Ds-efficiency each design achieves anywhere in a rectangle of parameter values 4 times wide in each coordinate. The design built at the guess guarantees 7.3%: it is perfect where it was built and nearly useless at one corner. The D-optimal design at the same guess guarantees 25.6% — it answers the wrong question everywhere and is therefore not concentrated on being right anywhere. The third is the one this field exists to measure: maximin over the first parameter's whole range, with the second held at its guess. It guarantees 7.0%, which is no better than the design that protects nothing. Protecting both is worth 45.9%, and it needs 3 settings to do it.

    The worst case in two directions

    A design that protects a range of one parameter is robust. Protect the range of one parameter while holding the other at a guess and the design is still robust, still has a guarantee, and guarantees no more than a design that protects nothing at all.

    part 7 · blind
  8. What a rule reads, against what the outcome uses. The variance of the treatment estimate relative to a coin's, for four things a rule might balance against three shapes the outcome might have, over 350 trials of 200 units. The diagonal is the easy part — a rule that reads the function the outcome uses removes about half the variance. What the table is for is the off-diagonal: reading the covariate alone is worth nothing against a quadratic (0.755), and reading all three is worth nearly as much against every shape as the matching rule is against its own (0.532, 0.493, 0.514).

    Three functions of one number

    A rule that balances the covariate is exposed to every shape the outcome might have. A rule that balances three functions of it costs two points of variance against the shape the first was built for and takes the worst case from a coin's to about half of it.

    part 8 · shape
  9. The six best bases of 2 functions, and what each protects. Every cell is R²(g | span B) — the share of the imbalance in that shape a rule balancing that basis removes — computed from exact inner products between Hermite functions and indicators, with nothing simulated. The rows are ordered by their worst cell, which is the number an experimenter who does not know the shape is exposed to. The best row here guarantees 26.8% against every shape in the list, and the worst of the six guarantees 15.1%: the difference between them is entirely which subspace was picked, at the same cost per arrival.

    A basis is a subspace

    A balancing rule cannot tell one basis from another with the same span, so choosing what to hand it is choosing a subspace — and then what it removes of any outcome shape is a projection, computable exactly, with no trial anywhere in it.

    part 9 · basis
  10. The zero was a fact about independence. What a balancing rule handed every main effect of both covariates removes of a pure interaction, as the covariates are allowed to move together. At ρ = 0 it is exactly nothing — at machine precision, at any number of main effects — which is the independent-covariate result and is correct. It is not small anywhere else: the product of the two covariates loses 64.0% of itself by ρ = 0.5, because h₁h₁ = h₀ + √2·h₂ and Mehler pairs h₂ with h₂ at ρ². Four interactions are drawn and none of them keeps the zero.

    A zero that was an assumption

    A rule handed every main effect of both covariates removes exactly none of a pure interaction. That is true at machine precision, it is a fact about independence, and it dies as the square of the correlation.

    part 9 · joint
  11. The guarantee that survives a correlation, and the one that does not. What a balancing rule handed both main effects removes of the pure interaction between them, as the covariates become dependent. For median splits it is exactly zero at every correlation, because sign(x)² = 1: the interaction sign(X)sign(Y) is orthogonal to sign(X) and to sign(Y) whatever ρ is. For the product of the raw covariates it is 4ρ²/(1+ρ²)² — 64.00% by ρ = 0.5, rising to all of it at perfect correlation. A cut away from the median sits between them and is not small: 23.01% at a cut of one. The zero is not a fact about interactions. It is a fact about a dictionary whose functions square to a constant, which a polynomial one does not.

    The zero that survives a cut

    A rule holding both main effects removes half of a pure interaction between correlated powers and exactly none between correlated median splits. The guarantee that a correlation destroyed was never about interactions.

    part 10 · splits
  12. The rule is parity, and it runs both ways. At a correlation of 0.5, four combinations of a dictionary and an outcome shape. The joint sign flip (X, Y) → (−X, −Y) leaves the bivariate normal alone at every correlation, so a function that changes sign under it is orthogonal to one that does not. A product of two odd functions is even; a product of an odd and an even one is odd. So an odd dictionary removes exactly none of the first and something of the second, and an even dictionary does the reverse — which it does, to machine precision, in both of the two rows that should be zero. This is one rule where there had been two: that a median split's square is constant, and that a polynomial dictionary contains the products a correlation generates.

    A dictionary that is neither

    A rule handed two median splits removes none of their interaction; a rule handed two covariates removes none of their product. Those were two results with two explanations, and they are one result with one — and finding it corrected the number underneath both.

    part 11 · dict
  13. The area under the window is what the band actually costs. The three windows' weight sequences at a width of 30 lags, drawn against the lag as a share of the window. A truncated window applies a weight of one to every lag inside it and zero outside, which is why its sum is the width and why every conventional charge is right for it — and it is a covariance matrix on almost no sample, so it cannot be used. The Bartlett window falls linearly to zero and its weights sum to exactly 15.000000000000004, which is half the width, at every width: Σ(1 − k/(L+1)) over k = 1 … L is L − L/2. The Parzen window sums to 11.13 here, three eighths of the width, and it gets there by holding a weight near one over the first few lags and then falling faster. A plug-in estimate multiplied by a weight below one is a shrunk estimate, and a shrunk estimate is worth less than a free one — which is the whole of why a charge levied per lag is a charge for parameters the window has already spent.

    The charge nobody derived

    A band of lags is charged one log-likelihood unit apiece, because that is what a regression coefficient costs. A band's numbers are not regression coefficients, and measuring what they actually cost puts the convention out by a factor of nearly three.

    part 12 · dimension
  14. How much of one search the other has already found. Five pairs of searches on one sample, on a scale whose zero and one are both fixed by construction. Zero is two searches over disjoint sets of independent columns: they remove shares of the residual sum that add, at 0.8 standard errors from exactly additive, and they read 0.004. One is a break search paired with a step column it contains, which reads exactly one on every draw because the step adds nothing at all. Between them: two dictionaries of step columns cut a few rows apart read 0.125, and the pair the earlier field measured — a break and a whitening window, both reading the same residual series — reads 0.762, three quarters of the way to one search containing the other. And below zero, a break paired with a search over independent columns reads -0.306: the joint search finds configurations neither half of it contains, so charging the two separately under-charges.

    Two searches that share nothing

    Two searches over independent columns remove shares of the residual sum that add exactly. On the scale a chi-square point is quoted on they look super-additive by a fifth of a unit, and none of it is overlap.

    part 12 · apart
  15. What the worst case is worth, one function at a time. The smallest share each dictionary removes, over seven outcome shapes, at a correlation of 0.5. A rule balancing the mean of each covariate has a worst case of exactly zero — against the square, and against both products. Adding a median split to it, which is the second thing every trial balances, leaves the worst case at exactly zero, because a median split is odd and so is a mean. Adding the square instead moves it to 6.8%, and the extra functions after that move it to 7.4%. The worst case is decided by which parities the dictionary contains rather than by how many functions are in it.

    What the extra function buys

    A rule balancing the mean of each covariate has a worst case of exactly zero. Adding the median split — the other thing every trial balances — leaves it at exactly zero, and one square moves it.

    part 12 · dict
  16. One zero holds and one does not. Three rules, at a correlation of 0.5, against the skewness of the covariate. A rule balancing the mean of each covariate removes exactly nothing of their product when the marginal is symmetric — including the heavy-tailed symmetric one at skewness zero, which is what says the guarantee needs symmetry rather than normality — and removes up to 29.7% when it is not. A rule balancing a median split of each removes exactly nothing of the product of the splits under every marginal here, to 1e-30: both sides are functions of the sign of the latent normal, and a monotone transformation moves neither. A rule balancing a threshold at a value on the covariate's own scale removes between 4.9% and 22.5% — it never had a zero to lose, under any marginal at all.

    A zero that rests on a symmetry

    A balancing rule removes exactly none of an interaction between two odd functions, at every correlation. The argument needs the joint sign flip to preserve the law, and no real covariate is symmetric about anything.

    part 13 · skew
  17. What the second search finds, alone and afterwards. For four of the pairs, what the second search removes on its own and what it removes once the first has already run. The gap between the two is the overlap in absolute terms. Where the searches share nothing the two readings are the same: an independent column removes 0.0261 alone and 0.0260 afterwards. Where one contains the other they are 0.1387 and exactly zero. The pair the earlier field measured sits between: a whitening window removes 0.5033 alone and 0.3120 after a break search has run. This is the earlier field's own reading of its pair, on the share scale rather than in log-likelihood units, and it is the number a rule that runs both searches actually has to charge for.

    A search that is already the other

    A break search shifts every coefficient after a row, so a step column is one of the directions it can move in. Paired with a dictionary of them it reads exactly one, on every draw, and that fixes the top of the scale.

    part 13 · apart
  18. Four windows, one line, and one that is off it. The optimism measured for each window at a band of 30 lags, against what that window's weights sum to, on 2000 pairs of independent samples of 120 rows. The diagonal is where a window that spent exactly its summed weights would sit. Three of the points are one shape at three levels — the Bartlett window, its square and its cube, whose sums stand in the ratio 6 : 4 : 3 — and they lie on a line through the origin at 0.767 of the diagonal, with 0.033 between the highest and the lowest. Scaling the weights scales the charge by the factor the weights predict, which is what makes the weights the mechanism. The Parzen window has a comparable sum and a different shape, and it sits at 0.871: its weights stay near one over the first few lags, and the first few lags are where the information is. A weight sum treats every lag as equally informative and no sample does.

    What a window leaves free

    A Bartlett window's weights sum to exactly half its width, which is a candidate for what the band costs. Varying the weights without varying anything else says the weights are the mechanism; varying the shape at the same weight says they are not the arithmetic.

    part 13 · dimension
  19. One zero is arithmetic and one is a symmetry. What two balancing rules remove of the interaction they are aimed at, on five joint laws of the ranks matched at a Spearman correlation of 0.4, with a normal covariate throughout. A rule holding a median split of each covariate removes exactly nothing of the product of the splits under every one of them, including the two that are not symmetric under reflection — and the reason is not a symmetry at all: a centred median split takes the values ±½, so its square is a quarter identically, and the interaction is orthogonal to both main effects whatever the joint law is. A rule holding the mean of each removes exactly nothing under the three radially symmetric copulas and 7.71% under the two that are not. Bars at the floor are exact zeros; the axis cannot draw 9e-32.

    A zero that is arithmetic

    A median split's exact zero was explained by a symmetry of the latent normal. It holds under a Clayton copula, which has no such symmetry, because a centred median split squares to a quarter identically.

    part 14 · copula
  20. Where the two kinds of cut sit. Six covariates, each a monotone transformation of the same latent normal. The vertical line at zero is where every median split sits, on every one of them, because a monotone map preserves order: the median of the covariate is the image of the median of the latent normal. The marks on the curves are where a threshold at 1 on the covariate's scale falls — 1.000, 0.881, 0.875, 0.783, 0.713, 0.337 — and none of them is at zero. That is the whole of the difference. A function of the sign of the latent normal is odd, and a rule made of odd functions removes exactly nothing of an interaction between two of them; a threshold anywhere else is neither odd nor even and removes something.

    A split survives what a mean does not

    The two things every trial balances come apart on a skewed covariate. A median split is a function of the sign of the latent normal whatever the marginal is; a mean is not, and its exact zero is gone at a skewness of one.

    part 14 · skew
  21. Cut the charge and the width follows it. The band width each charge picks, averaged over 400 draws of 120 rows under AR(1) at 0.8, with the standard deviation across draws beside it. Schwarz's charge — half a log n a lag, which is 2.39 here — picks 3.67. Akaike's picks 6.02. Charging the numbers the window actually leaves free, which is half the width, picks 10.12; charging what the optimism measures, 0.767 of that, picks 14.15. A charge and the width it buys are very nearly reciprocal, which is what a likelihood rising at a fixed rate a lag implies and is why the four answers span a factor of 3.86. The width that was actually best on the draw averages 13.90 and moves by 10.30 from draw to draw — three times as much as any rule's answer does.

    A width that moves and an error that does not

    Four charges give four widths a factor of four apart and four errors half a per cent apart. The derived charge wins, significantly, by a quarter of what was on offer — and none of the four is an estimate of anything.

    part 14 · dimension
  22. The ladder is the same ladder under every law. Each pair's overlap under each of the four laws, over 400 draws apiece. The scale is fixed at both ends by construction: two searches over disjoint sets of independent columns read -0.001, 0.013, -0.002, 0.004, and a break search paired with a step column it already contains reads exactly one under every law. Between them the pair that reads one residual series twice runs 0.763, 0.805, 0.577, 0.752 — lowest under long memory, where a whitening has most to do and the break search has least left to find that the whitening has not taken. The rung that moves most is the pair of step dictionaries, from -0.148 under a moving average to 0.419 under a break; and the pair that is negative is negative under all four.

    Three quarters of the way to one search

    The pair that started this reads 0.762 on a scale whose one is containment. And the pair that shares nothing but its response reads −0.306, so the sign the earlier field found does not transport at all.

    part 14 · apart
  23. The mean's zero is the copula's symmetry. Five copulas, each at a Spearman rank correlation of 0.4, with a normal covariate throughout — so nothing here is about the marginal, which is the whole of the earlier field. Horizontally: how far the copula's density is from its own reflection through the centre of the unit square, measured rather than read off the family's name. Vertically: what a rule balancing the mean of each covariate removes of their product. The three copulas at zero on the horizontal axis remove exactly nothing, to thirty decimal places. The two that are not symmetric remove 7.71%. A guarantee that held for six marginals turns out to have needed something the marginals could not have told anybody about.

    The symmetry the marginals could not show

    A mean's interaction zero needs the covariate to be symmetric and the copula to be symmetric under reflection. Six marginals could only ever test one of those, and the other is broken by the commonest kind of dependence there is.

    part 15 · copula
  24. A cut at a quantile, and a cut at a value. Two rules that read identically in a protocol. One splits each covariate at its median; the other splits it at 1 on the covariate's own scale — a dose, a temperature, a clinical threshold. At a correlation of 0.5 the first removes exactly nothing of the interaction between its own two splits, under every marginal here, because a median split is a function of the sign of the latent normal whatever the marginal is. The second removes what the bars show, and it does so on a normal covariate too: the threshold sits at 1.000 on the latent scale rather than at zero, so it is 59.4% odd and 40.6% even. The exact zero was never about the cut; it was about the cut being at the median.

    The cut that is not a quantile

    A protocol that says split the covariate at a threshold and one that says split it at the median read the same and are different rules. One has an exact guarantee under every marginal and the other has none under any.

    part 15 · skew
  25. Which tail the threshold is in. What a rule balancing a threshold at 1 on each covariate's own scale removes of the interaction between the two thresholds, on five copulas matched at a Spearman rank correlation of 0.4 with a normal covariate throughout. This rule never had a zero to lose — the earlier field establishes that under every marginal — so what is left is a size, and the size depends on where the dependence lives. A Clayton copula, whose density piles up in the lower tail, leaves 5.33%; the same copula turned over, so that it piles up in the upper tail where the threshold is, leaves 33.36%. Same rank correlation, same Kendall tau, same marginal, same threshold: 6.26 times the leak, decided by which end of the distribution the dependence and the cut are both in.

    Which tail the cut sits in

    The same copula and its reflection have the same rank correlation, the same Kendall tau and the same marginals. A balancing rule holding a threshold at a dose leaves 5.33% under one and 33.36% under the other.

    part 16 · copula
  26. A guarantee that stops being a number. The worst case of each dictionary over six outcome shapes, at a correlation of 0.5, against the skewness of the covariate. Under a symmetric marginal every rule made of odd functions has a worst case of exactly zero, and the rule holding a mean and a median split of each covariate — the two things every trial balances — is one of them. Under skew that zero becomes 0.74%, 1.83%, 2.24%, 2.49%: small numbers, each of which depends on a marginal nobody stated. The guarantee has not improved by becoming positive. It has stopped being a guarantee, because it can no longer be written down without the covariate's distribution in it.

    Balancing a skewed covariate

    The worst case of the rule every trial runs goes from exactly zero to somewhere between a quarter of a per cent and two and a half. Which is small, and is a number that cannot be stated without the covariate's distribution in it.

    part 16 · skew
  27. The term that cancels, and the term that does not. The volume each candidate's whitening moves — log|Ω̂| — for a sieve of order 4 on one sample of 120 rows. Estimated once from the fullest candidate and used for the whole table, it is the same number for every candidate, so it drops out of every difference the criterion reads: that is why nothing in this collection has ever needed to carry it. Estimated from each candidate's own residuals it ranges over 23.87, which is more than a parameter is worth, and the criteria being compared are then fits made under different error models with no term saying so. The window's rule has carried this term since the estimated-covariance field and the sieve's never had it.

    The volume a whitening moves

    A sieve's whitening has a determinant and this collection's criterion for it never carried one. Shared across a table the term cancels exactly, which is why nothing ever noticed; used per candidate it is worth more than a parameter and the whole comparison turns on it.

    part 18 · lists
  28. The curvature is in the denominator. The optimism a Bartlett band of each width actually costs, divided by that width, on two ways of measuring the width, over 2000 draws at 120 rows. Measured in the weights the band spends — Σ w(k), which is what the earlier field levies its charges on — the reading falls from 0.9528 at two lags to 0.7486 at thirty, so a charge proportional to the summed weights is too dear at one end and too cheap at the other. Measured in the pairs the band uses — Σ w(k)(1 − k/n), because a lag of k is an average over n − k products — the same readings are flat from 4 lags up, at 0.0084 of χ² per width against 0.2359. The correction has no fitted parameter in it: it is a function of the window, the width and the sample size.

    The width a band is measured in

    A tapered covariance band spends 84% of its own weights at two lags and 74% at thirty. Every charge in the collection is a straight line through the origin in those weights, so it is too dear at one end and too cheap at the other.

    part 19 · curve
  29. The curvature is in the denominator. The optimism a Bartlett band of each width actually costs, divided by that width, on two ways of measuring the width, over 2000 draws at 120 rows. Measured in the weights the band spends — Σ w(k), which is what the earlier field levies its charges on — the reading falls from 0.9528 at two lags to 0.7486 at thirty, so a charge proportional to the summed weights is too dear at one end and too cheap at the other. Measured in the pairs the band uses — Σ w(k)(1 − k/n), because a lag of k is an average over n − k products — the same readings are flat from 4 lags up, at 0.0084 of χ² per width against 0.2359. The correction has no fitted parameter in it: it is a function of the window, the width and the sample size.

    A lag the sample has less of

    A sample autocovariance at lag k is an average over n − k products, not n. Count a band's width in the pairs it actually has and the curvature in its charge goes away, on a correction with nothing fitted in it.

    part 20 · curve
  30. A line in the right width beats two curves. How far each candidate charge sits from the measured optimism across the plateau, in units of each width's own standard error, over 2000 draws. The straight line through the origin in the band's summed weights — which is what the earlier field levies — misses by 0.2382 per width. The same straight line in the pairs the band actually uses, Σ w(k)(1 − k/n), misses by 0.0095. A fitted power law misses by 0.0293 and a fitted decaying rate by 0.0172, both on one fitted constant more. The deferral this field answers asked for a curve; the answer is a line, in a variable with nothing fitted in it.

    A line that beats two curves

    A deferral asked for a curve. Fitted against the same measurements, a straight line in a variable nobody had to fit describes the plateau better than either curve does with a constant more — and for three windows out of four it does not.

    part 21 · curve
  31. The scale moves the width; the curve does not. The band width each charge picks, averaged over 150 draws of 120 rows. The two conventions — a unit a lag and half a log n a lag — pick 6.08 and 3.65 lags. The four charges derived from the measured optimism pick 15.05, 15.60, 15.47 and 15.44, against a best width on the draw of 14.13. So the scale a charge is levied on moves the width by a factor of 4.27 and the shape of the charge moves it by 3.6%. None of the six is an estimate of the draw's own best width: the correlations are -0.006, -0.003, 0.017, -0.001, 0.016, -0.004.

    What a better charge buys

    Four charges derived from the same measurements pick band widths within six per cent of each other and deliver errors within two per cent of the gap any of them leaves. The scale a charge is levied on decides the width; the shape of the charge decides nothing.

    part 22 · curve
  32. Largest where least is needed. What the pairs correction supplies against what each window's measured profile needs, across this field's plateau, over 2000 draws at 120 rows. Both are stated as the multiplicative rise the charge per unit of width has to take between four lags and thirty. What the correction supplies is arithmetic — (1 − μ(4)/n)/(1 − μ(30)/n), where μ is the mean lag of the weight the band adds — and it runs 1.0795, 1.0580, 1.0456, 1.0539 for the four windows. What the measurement needs runs 1.1076, 1.2928, 1.2296, 1.6550. The two orderings are opposite: the plain Bartlett window has the longest mean lag, so it gets the biggest correction, and the flattest profile, so it needs the smallest. They coincide to 0.9746 of each other, and nowhere else does the correction account for more than 85.0% of the fall.

    What the correction assumes

    A correction with nothing fitted in it repairs one window of four. The reason is that its size is set by where a window puts its weight and the curvature it must repair is set by something else — and for one window at one sample size the two happen to agree.

    part 23 · curve
  33. Reading the draw changes what is charged, not what is tracked. The correlation between the band width each rule picks and the best band width on the same draw, over 400 draws. The three fixed charges read -0.069, -0.066, 0.012. The three that read the sample read -0.012, -0.019, -0.041. None of the six is distinguishable from nothing. The statistic the first plug-in reads does vary — the draw's own summed squared autocorrelation runs from 2.06 to 10.19 with a mean of 4.36 — so the failure is not that the charge stopped moving. It is that what it moves with carries no information about which width this draw wanted.

    A charge that reads the draw

    Three charges built to read the sample track the best band width on their own draw at −0.012, −0.019 and −0.041, deliver more error than the fixed rule they are calibrated to, and pick a width half again as variable. The statistic moves; the answer does not.

    part 24 · curve

All series