How this site is made

The figure library

Every picture here is generated from code at build time and asserts something about the numbers it just drew. This page lists the generators by the field they were written for, with the title each gives itself and how many essays call it.

No figure on this site is a drawing that was made once and saved. Each one is a function: it takes parameters, runs the simulation or sums the closed form, and returns SVG — so the same generator draws the coverage of a 95% interval at twenty observations and at two hundred without either picture being redrawn.

Two things follow from that, and both are the reason this page exists. The first is that the collection can keep growing without the illustrations drifting apart: a generator is written once and every essay that calls it inherits the same line weights, the same colour roles and the same behaviour in dark mode. The second is sharper, and particular to a site whose figures are samples. A generator asserts what is true of its own arguments before it draws anything, so a figure that is drawn is a claim that has passed; and because the generator is seeded, it is byte-identical on every build. There are 80 of them so far, called 587 times across 436 essays, in 80 families.

A generator is not one picture. It is one object with a set of views — the windows a block resample can carry, the order of the bias each of them buys, the critical value any of it reaches — and an essay walks a generator through the views its argument needs rather than collecting a picture from somewhere else for each paragraph. Every view a generator draws is on that generator's own page.

Each name below carries the title the generator gives itself — the one it computes from the numbers it has just drawn and puts on its own caption strip. Nothing here is transcribed, so a title on this page that differs from the title on the figure would mean this page is stale rather than that the description is wrong.

The distribution itself

Sums converge on it, which is the most-quoted theorem in the subject. What the quoting leaves out is the rate: the middle converges quickly and the tail does not, and the tail is where the approximation is actually read. The same theorem carried through a function stops being a normal limit where the function is flat, and carried through a ratio it forbids every interval that is always finite. 1 generator, 40 views

Intervals, counted

An interval that claims 95% is making a checkable statement about a procedure. Build every possible sample and count. The interval taught first fails, the failure is worst where proportions are most often reported, and more data does not monotonically help. 1 generator, 47 views

Tests, and the second number

A p-value alone cannot be read: the same 0.04 means different things at different sample sizes, and nothing at all without knowing how many analyses were available. Under a real effect it is a draw from a distribution several orders of magnitude wide, a replication of a result at 0.05 succeeds exactly half the time under any model centred on that result, and the standard ways of combining several p-values disagree about which evidence counts. Every figure here carries the number that makes it interpretable. 1 generator, 18 views

Reversals that are not errors

Simpson's reversal, the base rate, regression to the mean. Each is normally taught with one famous table. A table is a point; these are swept, so how much of the space behaves that way and how large it can get both have answers — including why two correct analyses of one baseline disagree, how far a group enrolled on a noisy reading falls with nothing done to it, and why the correlation predicts the regression of the extremes only when the population is normal. 1 generator, 18 views

The prior, doing visible work

A credible interval says the thing everyone wants a confidence interval to say, and it needs a prior to do it. So the prior is treated as a component with a measurable effect: it is worth a stated number of observations, and the interval it produces has a coverage that can be summed over the sample space like any other. 1 generator, 16 views

Regression, and what the summary hides

A slope, a standard error and an R² can all be computed from data the model is grossly wrong about, and none of them says so. Leverage is a number known before the outcome is looked at, influence is a number, and whether a residual plot looks bad is a question with a calibrated answer. Two wrong rows placed together hide from every single-row diagnostic, and a loss that bounds a large residual does nothing about a far one. 1 generator, 23 views

Corrections, and what each controls

Bonferroni bounds the chance of any false positive; Benjamini–Hochberg bounds the share of the findings that are false. Both get called correcting for multiple comparisons and they are different promises, so every procedure here is made to report both rates and the power each one costs. 1 generator, 8 views

When the data stops early

A subject still event-free when a study ends is not missing and not observed — it is known to exceed something, which is a third state most tools have no slot for. The estimator that gives it that slot recovers the curve, and everything read off the curve afterwards has a condition attached: the interval printed around its end stops covering where the curve is read, a dropout that carries information leaves a record identical to one that does not, one minus the curve is not a risk once something else can end observation first, and a hazard ratio is an average whose weights the length of follow-up chose. 1 generator, 9 views

Stopping rules

A p-value is defined relative to a sampling plan, so two experiments with identical data and different stopping rules have different p-values. That sounds philosophical and it is arithmetical: testing five times at the nominal level rejects a true null 14% of the time. 1 generator, 8 views

Groups that borrow

Eight hospitals are neither one hospital nor eight unrelated problems, and the estimate between those two answers is not a compromise. It is a weighted average whose weight — se²/(se² + τ²) — is decided by how well each group is measured and by nothing else, and the population spread in it is estimated from the data rather than assumed. 1 generator, 19 views

Decided before the data

Every other field here repairs an analysis after the fact. These figures are about the arrangement that makes the repair unnecessary: blocking removes exactly the variance the blocks carry, randomisation buys a known reference distribution rather than balance, and varying every factor at once estimates each of them from every run. 1 generator, 13 views

When the observations repeat each other

Every standard error on this site divides by √n, which claims the observations carry independent information. In time order they usually do not: at a lag-one correlation of 0.8 a fifty-point series is worth about six independent observations, its 95% interval covers 47%, and two series that wander are called related three times out of four. 1 generator, 7 views

The spread, and its own uncertainty

Partial pooling estimates the population spread from the group means and then uses it as though it were known. It is not known: eight groups leave τ anywhere in a range that spans a factor of several, and on a third of eight-group datasets the usual estimate of it is exactly zero. Carrying that uncertainty rather than dropping it is the difference between an interval that covers 95% and one that covers 79%. 1 generator, 7 views

Hierarchy past one number

The same weighted average, on quantities that are not means. A slope, where how much a group borrows is decided by the arrangement of its x values rather than by how many it has. A proportion, where the pooling has to happen on a scale the data does not live on. A second level of grouping, whose arithmetic turns out to be the effective sample size of the time-series field arrived at from the other direction. 1 generator, 11 views

Series that move together

Two random walks regressed on each other are called related three times in four, which is why the time-series field ends in a warning. The exception it names and does not measure is here: when the pair is genuinely tied, the fitted relation converges at rate 1/n rather than the usual 1/√n, and the test that separates the two cases cannot be read against any table that exists. 1 generator, 7 views

Three series, and a count

A pair is either tied or it is not, so its whole inference is one test with one answer. Three series can carry none, one or two relations at once, and the quantity being estimated stops being a slope and becomes an integer — read off the gap in a spectrum, against a critical value that depends on how many things are left wandering and on nothing else. 1 generator, 7 views

The surface between the corners

A two-level factorial answers which factors matter and is structurally unable to answer what setting is best: every run sits at a corner, every squared term is 1 there, and the column that would estimate curvature is a copy of the intercept. What it takes to see a curve, how far a fitted gradient can be trusted, and why the location of an optimum is a ratio of estimates rather than an estimate. 1 generator, 17 views

A design chosen rather than looked up

Every design here so far came from a catalogue and was then measured. Turn the arithmetic around and a design is the answer to an optimisation: maximise a functional of X′X and see what comes back. What comes back is the catalogue's own nine settings, re-weighted — and an exact theorem that says when a search is finished without knowing what it was searching for. 1 generator, 10 views

Splitting the units

The only design decision that costs nothing: the same units, the same measurements, the same analysis, and a different variance. Sample the noisier arm more, the expensive one less, and the shared control by the square root of the number of arms — and then notice that every one of those rules is a function of a quantity nobody has, and measure what happens when it is estimated instead. 1 generator, 8 views

Designs that change while they run

The stopping-rule field is a fixed design looked at more than once. Here the design itself is a function of the data — how many units, which arm the next one goes to, which arms survive the interim — and the question stops being what the rule spends and becomes what it leaves behind. The estimate from an arm chosen for being ahead is ahead by more than it should be, and the unbiased estimate is the one that throws away the data the choice was made on. 1 generator, 5 views

The reference distribution the design supplies

Hold the outcomes fixed, re-run the rule that assigned them, and count. A p-value built that way needs no assumption about the outcomes, no large-sample argument and no nuisance parameter — it needs to be told the rule, which is the one thing the experimenter already knows. It repairs the adaptive design the previous field leaves broken, and it is not free. 1 generator, 6 views

The observation that has not happened

Two fields here stop at estimation. A forecast is the other question — not what the parameter is but what the next observation will be — and the band round it is computed by substituting estimates into a formula derived for the truth. Counted, that 95% band is not 95%, the shortfall grows with the horizon, and the model has to beat two benchmarks that estimate nothing at all. 1 generator, 4 views

The criterion, and what it assumes

D-optimality gave the catalogue back, which was reassuring, and left two things unsaid. A, D and E are one family and the letter is a choice — the design that wins the first is nearly worst at the last. And for a model whose information depends on its own parameters, a design is optimal only at a guess about the answer, which makes the cost of guessing wrong a quantity with a closed form. 1 generator, 3 views

Balancing on what was recorded first

The adaptive field's rules read outcomes and every one of them broke something. These read baseline covariates, and each holds one imbalance flat while the others grow behind it exactly as a coin's do. What follows is the opposite of the outcome-adaptive story at every step: the unadjusted analysis is too cautious rather than too eager, and the exact test that cost nineteen points of power there costs almost nothing here. 1 generator, 4 views

Comparing two forecasters

The forecast field ranks forecasts by mean squared error and stops. Whether one forecaster is really better is a test, its terms are dependent because neighbouring forecasts overlap, and where one model contains the other it fails completely — declaring the smaller model significantly better while the null it claims to test is true, and more certainly the more data it is given. 1 generator, 10 views

A design that assumes less

A design for a non-linear model is optimal only at a guess about the answer. Two ways out: protect the worst parameter value in a range rather than the average, which is an optimum that sits on a tie rather than a slope; or stop guessing, run part of the experiment, and design the rest at the estimate — where the interesting question turns out to be what that does to the interval afterwards. 1 generator, 3 views

More arms than two

Minimisation balances a trial by making the arms' counts even inside every factor level. With two arms there is one way to measure how uneven two counts are. With three there are several, they are all called minimisation, and they send different patients to different arms — while the ratio a trial was designed to deliver quietly disappears unless the score was told about it. 1 generator, 3 views

The best of a set, and what the search costs

Comparing two forecasters is a test. Comparing eight is a multiplicity problem on top of a dependence problem, and the two do not separate: eight windows of one series carry the multiplicity of about two independent comparisons and eight separate problems carry eight, so a correction that charges for the number of models is wrong in both directions. The reference distribution has to be over the whole set — and once it is, the models nobody would have run turn out to cost more than the ones that were close. 1 generator, 3 views

What the design is asked to guarantee

Two halves of one question that were never separated: which parameters a design is for, and how much precision is enough. Protecting one parameter of a non-linear model over a range of its own values is a worst case of a ratio of determinants, and it is not a special case of either problem it is made of. Letting the experiment stop when it is precise enough is the adaptation with a theorem against it — the rule stops when its own noise estimate is low, so the interval it produces is short. 1 generator, 3 views

Balancing what has no levels

Every balancing rule in the two fields before this one reads a level. Age, blood pressure and a baseline score have none, and the first thing that happens to them is that somebody invents some — a choice with a cost available in closed form before any data exists: a median split can see exactly 2/π of a normal covariate, so a rule that balances its two halves perfectly still leaves three fifths of a coin's imbalance. A rule that reads the number instead does not beat that by a factor; it beats it by a rate. 1 generator, 3 views

The set field compares forecasts nobody estimated. A specification search compares a benchmark with a table of variants that all contain it, and two things change at once: every variant is behind before the search begins, by an amount with a closed form in the shape of the table, and the reference distribution for the winner can no longer be resampled from the data — it has to be generated from a model. The repair the nested case asks for, applied row by row, takes a table from its nominal level to a quarter. 1 generator, 3 views

What the procedure may not read

Two restrictions that were cheap until now. A design for a non-linear model may not read the parameter values, and where the model has two of them the guess is a point in a rectangle: the weights that were exactly 1/√2 become a function, and a design robust in one coordinate turns out to guarantee no more than one built at a single point. A stopping rule may not read the mean it will report — and there is an exact way to obey that, at a price in the width of the interval rather than in the number of observations. 1 generator, 3 views

The shape the covariate enters by

Every balancing rule in the two fields before this one optimises one over the variance of the treatment estimate in a model where the covariate enters linearly — which is not an assumption about the analysis, it is what the criterion is. Balancing the mean of a covariate removes exactly 2/π of the imbalance in a median split of it, a quarter of a threshold in the tail, and nothing at all from a quadratic. The ranking of the rules reverses between shapes, and against one of them every rule here is worse than a coin. 1 generator, 3 views

A search with no fixed point

The search field compares a benchmark with variants that all contain it. Take the fixed point away — let the benchmark be one candidate among sixteen, selected by the same data as its rivals — and the closed form for what every candidate is behind by survives intact, because it was never about nesting but about counting parameters. Three readings of one true null give 2.5%, 8.0% and 76.8%; Bonferroni is three times too strong on one table and not strong enough on another; and searching the benchmark makes the table harder to reject with rather than easier. 1 generator, 3 views

Choosing what the rule reads

A balancing rule cannot tell one basis from another with the same span, so choosing what to hand it is choosing a subspace, and everything about it is exact geometry: what a rule removes of a shape is that shape's squared multiple correlation on the span. The maximin basis over a list of shapes is then a finite problem with an exact answer — and over a whole subspace the guarantee is exactly zero rather than small, which is why the list cannot be avoided. Randomising which functions the rule reads is worth twice the best fixed choice. 1 generator, 3 views

The block size as a schedule

The blinded rule's block size is a dial between an interval's degrees of freedom, a stopping rule's, and the overshoot. Letting it change during the run keeps the coverage exact — the argument never needed the sizes to be equal — and then runs into an identity: the two kinds of degree of freedom add to N − 1 on every run, so a schedule cannot make more of both. What it can do is spend each where it is worth most, which is worth a few per cent and is not the free lunch the dial looked like. 1 generator, 3 views

Scoring a search without spending data

A hold-out is expensive: every row spent scoring is a row not spent fitting. There are two standard ways of not paying — an information criterion, which predicts the hold-out from the in-sample numbers, and a resampling, which builds a reference distribution from one sample. Neither avoids what it replaces. The criterion's penalty *is* the displacement, computed rather than counted, and it beats a rolling hold-out at every split there is; both it and the resampling are undone by the same defect, which is rows that repeat each other, because both are counting independent things and there are fewer of those than there are rows. 1 generator, 3 views

When the set is too large to walk

Two constructions that were each measured by enumerating everything, at the size where enumerating everything stops being possible. A balancing dictionary over two covariates is an outer product, every inner product in it is still closed form, and a rule holding every main effect of both removes exactly nothing of any interaction. A trial of two hundred units has assignments that cannot be walked — and the acceptance rate turns out not to depend on the size of the trial at all, so the exhaustion a sixteen-unit count runs into is a fact about sixteen units. 1 generator, 3 views

A promise about two arms

The fixed-width interval whose coverage is exact was built for one mean, and almost nothing anybody runs an experiment for is one mean. The construction survives with the harmonic effective size in place of the block size, and four things do not: a unit of that size costs four observations rather than one, the second variance puts Neyman's allocation in reach of a blinded rule, the conservation identity comes up one degree of freedom per block short, and there is an exact estimator under either of two conditions and none under both. 1 generator, 3 views

Counting what is independent

Akaike's penalty is 2q because the optimism of a fit is a trace and the trace collapses to the parameter count when the rows are independent. Repair the trace and the criterion gets worse, on either scale, because the row count entered twice and a penalty is the second place: whiten the fit and the ordinary penalty is correct again, which recovers 88% of what counting rows gives up. Estimating the dependence from the rows being selected on costs 8% of that. And a multiplier resampling cannot keep more dependence than the residuals have, which is a ceiling rather than a tuning problem. 1 generator, 3 views

When the two are not independent

The dictionary that is an outer product needs the covariates independent, and the count that is a rate times a binomial coefficient needs the draws independent. Neither holds in a real trial. A linearisation turns the missing fourth-order expectation into Mehler's formula applied to products, which makes the whole geometry closed at any correlation — and the result it was built to test does not survive: what a rule removes of a pure interaction is exactly nothing at independence and 4ρ²/(1+ρ²)² everywhere else. A walk on the admissible set is exactly uniform, needs a burn-in, and is dearer than hunting until almost nothing is admissible. 1 generator, 3 views

The weights the corner needs

A fixed-width interval about a difference is exact when the arms share a variance or the allocation ratio is constant, and exact under neither when both fail. It is exact there too, with h_b(λ) = (1/m_A + λ/m_B)⁻¹ — the inverse variances, written as a function of the variance ratio alone, which is a contrast and so is readable by a blinded rule. The estimated precision weights that had no theorem behind them turn out to be that rule at an estimated ratio. And an interval's own scale estimate is right in exactly two cases: inverse-variance weights, and equal ones. 1 generator, 3 views

Estimating the dependence, not naming it

Whitening a sample repairs a criterion, and the whitening that repairs it is told the dependence is a first-order autoregression and left to find one number. A real dependence has no parameter. The obvious estimate — the sample autocovariances, cut off at some lag — is not a covariance matrix on half the draws there are, so the rule built on it does not exist; the tapered estimate that is always a covariance matrix costs a further seven points of what the repair is worth. And the two constructions named as escaping the resampling's ceiling turn out to be one construction and one identity: a moving block attenuates exactly as a blocked multiplier does, because the attenuation is the join. 1 generator, 3 views

A cut point, at a correlation

The geometry of a balancing dictionary over two dependent covariates is closed for polynomials and was taken to be asymptotic for cut points, because a threshold's Hermite coefficients never terminate. Conditioning on the second variable closes it exactly: every mixed inner product is ρ^j times a one-variable answer and the only two-dimensional object left is an orthant probability, which at the median is (2/π) arcsin ρ. The truncation that was feared falls geometrically in the correlation rather than algebraically in the order — and the interaction guarantee a correlation destroys for powers survives it exactly for median splits, because a two-valued function squares to a constant. 1 generator, 3 views

What a block may vary

Two questions from two corners of the collection with the same answer. A walk over the admissible assignments that exchanges more than one unit per arm mixes faster and is refused more often, and the trade is exactly computable: the gain is a factor of six where a hunt is fifty times cheaper anyway, and nothing at all where the comparison is actually decided. And a variance ratio that drifts between blocks cannot be estimated inside the block it weights — but it can be modelled across them, once the bias in a log variance estimate is subtracted, because that bias depends on the degrees of freedom and the degrees of freedom alternate with the allocation. 1 generator, 3 views

The shape a dependence has

Estimating a covariance rather than naming it was priced on errors that really were a first-order autoregression, where generality can only cost. Here are four laws with the same first lag and nothing else in common: a moving average that stops, a memory that does not, a break in the middle. The cost of generality where it is not needed and the benefit of it where it is turn out to be the same size — and against a covariance that changes with position rather than with gap, one number, a window and an order are worth exactly the same as each other. Both tuning parameters are then chosen from the sample, and every criterion available for choosing them is aimed at something else. 1 generator, 4 views

A block, weighted inside itself

Every resampling in this collection attenuates the dependence it is trying to keep, and until now every one of those attenuations was inherited rather than chosen. The construction the long-run-variance literature actually uses weights the residuals down towards each block's own ends — and what that buys turns out to be exactly the squared value of the window at the two ends and nothing else about its shape. It is an asymptotic gain that has not arrived at any block length a hundred and twenty rows can afford, and at the lengths that are available a tapered block is very nearly a shorter plain one. 1 generator, 3 views

What a dictionary buys and what it costs

A balancing rule is a list of functions, and the list decides two things that pull against each other. What it protects against is decided by parity: an interaction between two odd functions is even, an odd function is orthogonal to an even one at every correlation, and the two things every trial balances — a mean and a median split — are both odd, so their worst case is exactly zero however many of them are held. What it costs is the assignments it leaves, and past a certain thinness those stop being one set: the admissible assignments split into an arrangement and its mirror image, no sequence of admissible single swaps joins them, and the walk that samples them is uniform on half the reference distribution for ever. 1 generator, 4 views

When a fixed width is reached

A fixed-width interval promises a precision rather than a sample size, so the trial stops when it has enough — and what the stopping rule is allowed to read decides whether the interval means anything. A rule that stops when its own interval is short enough is stopping on the spread it is about to quote, at a correlation of 0.938, and every weighting loses three to five points of coverage including the one told every true variance ratio. The width a set of weights will produce is predictable from the within-arm sums of squares alone, which are independent of every difference the interval is about; a rule that stops on the prediction is back at nominal, costs two and a half blocks, and gives up half of the width promise it was keeping. 1 generator, 4 views

Fitted together, or fitted after

Every whitening in this collection is a two-step rule: fit a candidate, read the dependence off what is left over, whiten, fit again. The step nobody examined is the first one. A fit removes memory as well as signal, and how much is arithmetic rather than noise — the residuals of a hundred and twenty rows report a lag-one coefficient of 0.73 where the errors report 0.78 and the law says 0.80. Estimating the coefficient with the line instead of after it recovers most of that, and iterating the two-step rule to its fixed point recovers nearly as much — because a fixed point of the sum of squares is not a maximum of the likelihood, and the difference between them is real, one-sided, and worth nothing at all in the decision the number feeds. The order the same criterion picks moves by nearly two, which is worth a great deal more. 1 generator, 3 views

Paying for a search

A two-regime whitening chooses its change point by searching a profile, and then reads a criterion that counts parameters. Nothing charges for the search. Nothing can, in the usual way: under no break the two regimes share a coefficient and the break point is not identified at all, so the likelihood ratio is a supremum over a parameter that exists only under the alternative and no count of restrictions describes its distribution. Searching a hundred and twenty rows for a change point that is not there manufactures about five units of ratio where one parameter costs two — and a rule that counts the break as free splits a stationary sample on 99% of draws. A second break costs as much again, on a sample that has only ever had one. 1 generator, 3 views

What a chain cannot report

A thin admissible set falls into an arrangement and its mirror image, and the walk that samples it is uniform on half the reference distribution for ever while every diagnostic passes. That was found by enumerating fourteen units, and enumeration stops at about twenty-four. The test that does not enumerate is two chains — one from an assignment, one from its complement — compared on a statistic that changes sign under the complement, with a third chain from the same starting point to say whether a large reading is a fact about the set or about the length of the run. It agrees with the enumeration at every tolerance where the answer is known. At two hundred units it reports something else: the walk reaches the whole set at the tolerances a trial would use, and past a point the diagnostic stops agreeing with itself. 1 generator, 3 views

A guarantee that needed a symmetry

What a balancing rule can remove of an interaction is decided by parity, and parity is a statement about a symmetry of the law rather than about the covariate. Keep the dependence and change the marginals — a Gaussian copula under a monotone transformation — and the two things every trial balances come apart. A median split is a function of the sign of the latent normal whatever the marginal is, so its exact zero survives every transformation to the last digit. A mean is odd only when the marginal is symmetric, and its zero is gone at a skewness of one. The separating case is a marginal that is heavy-tailed and symmetric, where every zero holds exactly: it is not normality the guarantees needed. And a threshold at a value on the covariate's own scale never had one at all. 1 generator, 3 views

Where a taper's case begins

A tapered block was measured against a plain one on the exact bias each implies, and the answer was that the taper's advantage has not arrived at any block length a hundred and twenty rows can afford — with the crossing at ℓ = 20 and a difference there of three tenths of a point, too small for the critical-value table to resolve. Two of those three statements are about the wrong quantity. What a sample reports at ℓ = 20 is four points rather than three tenths, because the autocovariances the window is applied to are themselves attenuated and the window that discards the long lags loses less of them; and the comparison is made at a shared block length where each window has its own best one. Read at each window's own setting, on the error rather than on the bias, the ordering reverses at a hundred and twenty rows. 1 generator, 3 views

A covariance with no parameter

A regression's coefficients and its errors' dependence can be fitted together when the dependence is one number. When it is an estimated covariance there is nothing for “jointly” to mean — until a family is named, and then the family's own width is the parameter. Three things the measurement says, two of them the opposite of the guess: the plug-in's shortfall is the taper's rather than the data's and is nothing under two of four laws; nothing in the likelihood chooses the width, because nested families buy about a unit of it a lag and that is what a criterion charges; and a fit-only objective does not run away, because a unit diagonal fixes the trace. 1 generator, 3 views

How long the list is

A tuning parameter chosen per candidate rather than once for the table costs something, and the window's figure was measured while the order's was not — because their lists are different lengths and matching them changes what each rule is. Matched at every length, the two cost the same. What separated them was not the list at all: a per-candidate whitening is a different error model for every candidate, the Gaussian likelihood has a term that says so, and the sieve's rule in this collection never carried it. 1 generator, 2 views

Two searches over one sample

A break point that has been looked for costs more than a count of parameters says. So does a window chosen from a list, and nothing here had charged for it. A rule that does both is running two searches over one sample, and the charges do not add: the pair manufactures a fifth less than the two apart. What follows is sharper than a correction to a sum — most of what a break search finds under correlated errors is the correlation, so once a whitening has been chosen from the same sample the break's own charge more than halves, and carrying the published one across switches the test off entirely. 1 generator, 2 views

The diagnostic after the trial

The test for whether a balanced-assignment walk can reach the whole admissible set is run on a covariate function, before any outcome exists. Run instead on the difference in arm means it is the same test — under the sharp null the outcome is a fixed column — and it is about the statistic the p-value is actually built from. Two findings: an outcome is a probe nobody chose, and on a set that is genuinely split three in ten of them see nothing at all; and the published two-sided p-value is exactly right on half a reference distribution, while a one-sided one is not. 1 generator, 3 views

The other half of the dependence

What a balancing rule can remove of an interaction was measured across six marginals with one joint law of the ranks held fixed. Change the copula instead and the two exact zeros come apart for two different reasons. A median split's zero is arithmetic — a centred median split squares to a quarter identically — and holds under every copula there is. A mean's zero needs the copula to be symmetric under reflection as well as the covariate to be symmetric: it is exactly nothing under a Gaussian, t or Frank copula and 7.71% under a Clayton, at the same rank correlation and with a normal covariate throughout. 1 generator, 2 views

A block length chosen from the data

Two block windows were compared at each window's own best block length, which is the argmin of a quantity that needs the truth. Re-run on rules a practitioner could actually run, the ordering reverses: the taper wins at the best available length and at one estimated from the sample's own persistence, and the rectangle wins at a length written into a protocol and at the rule of thumb, at every sample size measured. And the whole argument is a third of the size of what estimating the length costs. 1 generator, 2 views

A charge for a covariance's own dimension

The two charges that pick a band's width were derived for a regression coefficient, and a band's numbers are neither free nor entered the same way. Measured as an optimism against a second independent sample, a Bartlett band costs 0.374 of a log-likelihood unit a lag where a criterion charges one. The window's own weights are the mechanism — scale them and the charge scales with them, to within two per cent across three windows — and they are not the arithmetic: the level is three quarters of the sum, a different shape at the same sum costs more, and the charge rises with the law's persistence. Levying the right one moves every width in the table by a factor of four and the error by half a per cent. 1 generator, 3 views

The rate and the size of a disagreement

A sweep of what it costs to let every candidate choose its own tuning parameter reported a product and called it a cost. Separated — and the decomposition is exact, because a draw on which nothing disagreed carries exactly zero — the rate rises by half across the list and what a disagreement is worth does not move at all. A third quantity explains why: a disagreement costs something only when it changes which candidate the table selects, that happens on about an eighth of draws, and the eighth does not depend on the list. 1 generator, 2 views

Two searches over different features

A break search and a window search on one sample manufacture less likelihood together than separately, and both read the same residual series. Put four more pairs beside them on a scale whose zero is two searches over independent columns and whose one is a search that contains the other, and the shortfall is a property of the pair: 0.00 for the control, 0.76 for the pair that reads one series twice, exactly 1 for containment — and −0.31 for a break paired with an independent column, where charging the two separately under-charges rather than over-charging. 1 generator, 2 views

A probe chosen rather than picked

A diagnostic that reports on what a balancing rule was not handed is run through a column that is 92% inside the span the rule balanced — because orthogonality in the population is not orthogonality on fourteen units. Projecting the probe off that span costs one least-squares fit and triples the separation it carries; choosing the direction from the design's own leverage is worth another factor of four over a random one; and the projection-pursuit direction the argument invites is worse than random. On a short chain the raw probe misses 44% of the sets that are split and the projected one misses 12%. 1 generator, 2 views

Both halves of the dependence at once

One field varies the marginal with the copula held Gaussian and another varies the copula with the marginal held normal, and the natural guess is that the two leaks compound. They do not add in either direction: eleven of twenty cells cancel and nine compound, a mildly skewed covariate under a lower-tail copula leaks 0.002% where adding the two gives 16.6%, and a heavy-tailed symmetric covariate that leaks exactly nothing on its own doubles what an asymmetric copula leaks. A median split's zero survives all thirty combinations at under 10⁻¹⁶. 1 generator, 2 views

The block length read on a quantile

An ordering between two block windows reverses depending on who chose the block length — measured on an implied long-run variance, which is the instrument that makes the sweep affordable and is not what anybody reads. Read on the 95% point a test uses, the tapered window wins under all four rules; read on the coverage the interval delivers, it wins under all four again. Two of the four rules change sign, and they are the two the recommendation was about. No rule and no window reaches its promised coverage: the eight cells run from 80.8% to 91.0%. 1 generator, 2 views

A charge that is not a straight line

The measured charge for a tapered covariance band falls from 95% of its summed weights at two lags to 75% at thirty, so every rule that levies it as a straight line through the origin is too dear at one end and too cheap at the other. The deferral asked for a curve; the answer is that the curvature is in the denominator. Counted in the pairs the band actually uses — a lag of k is an average over n − k products — the same readings are flat from four lags up, at a spread of 0.029 against 0.081, on a correction with no fitted parameter. It repairs the one window the earlier field measured and leaves the other three wanting a curve. 1 generator, 4 views

What decides whether a tuning list decides

The probability that a per-candidate tuning list changes which candidate a table selects is reported flat at about an eighth across list length, on a table and a world that are never varied. Vary how far apart the candidates are — one multiplier on the omitted coefficients, everything else held — and it runs from 17.6% to 1.5%, while the disagreement rate it is a factor of rises from 31.8% to 88.8%. The world where the candidates quarrel most about the tuning parameter is the world where the quarrel matters least, and a nested table turns over less rather than more. 1 generator, 3 views

Overlap and complementarity, separated

How much two searches over one sample share is measured as the net of two effects: ground both of them find, and configurations the joint search reaches that neither slice contains. Pin the first search at its own answer and search the second, and the two separate exactly — the pinned supremum cancels, so the split adds back to the original number on every draw. The control the whole scale is anchored on reads zero because its two components are several times larger and cancel, and the split depends on which of the two searches is pinned while their difference does not. 1 generator, 2 views

A probe from what the rule blocks

A balancing rule breaks the admissible set into pieces by blocking exchanges, so the quantity a diagnostic should be aimed at is the constraint's active set rather than the design's leverage — which is a heuristic about the same thing. Built from the design and the tolerance alone it is a real probe, well ahead of a random direction; it is also behind leverage at 4.4 paired standard errors. Counting the active set exactly, at a cost no trial can pay, makes it worse rather than better, so the approximation was never what cost it. 1 generator, 3 views

The same table at seven correlations

Every cell of the copula-by-marginal table is measured at one rank correlation, and eleven of its twenty cells cancel there. Sweep the correlation from 0.1 to 0.7 and four of the twenty change the sign of their answer, all four from compounding to cancelling, all four at the most skewed covariates. The near-perfect cancellation that is the earlier field's headline is a crossing: the cell passes through zero at a Spearman of 0.38, two hundredths from where it was read, and is two orders of magnitude larger by 0.7. 1 generator, 3 views

The interval, studentised

A percentile interval inherits the resampled distribution's skewness and its scale error together, and the standard repair is to resample a t-statistic so each resample carries its own scale. It is one extra variance per resample, and it repairs nothing: not one of the eight cells reaches its promise, the interval is twice as wide for half a point of coverage, and at the block lengths the rules choose the scale rests on two or three whole blocks. And the ordering between the two windows reverses at all four rules, on resamples that are the same resamples. 1 generator, 4 views

A forecast that is a probability

The field before this one ranks point forecasts by squared error. A probability forecast can be held to something stronger and stranger: say 30% often enough and about three in ten of those days should happen, which is a claim anybody can check by counting. Calibration alone passes a forecaster that issues the base rate every time, the diagram that draws it charges a blameless forecaster in proportion to how finely it was binned, and an honest record of a hundred forecasts shows more apparent miscalibration than the amount routinely read as evidence of a problem. A score that is not proper pays a forecaster to answer only 0 or 1, and what that liar gives up in ranking and resolution is computed rather than described. 1 generator, 8 views

Coverage without a distribution

Every interval counted here so far has been found short of what it promised. This one is not, and the reason is that it is not a statement about the data at all: the rank of a future observation among a set of exchangeable calibration scores is uniform, so an interval built at the right order statistic covers a stated share of the time at any sample size, under any distribution, around any model however wrong. What it does not say is where that coverage sits — and everything the procedure does not know turns up in the answer to that rather than in the total. 1 generator, 7 views

Weighting one sample into another

A weight turns the sample that was assigned into the sample a coin would have assigned, and the exchange is exact: integrated over the population, the standardised difference on every covariate goes to machine zero whatever the assignment rule was. What it charges is observations — a treated arm worth 94% of itself where the assignment is nearly a toss-up and 12% of itself where it is nearly decidable — and past that point trimming does not repair the estimate, it replaces the question. Weighting by an estimate of the probability turns out twice as precise as weighting by the probability itself, and weights fitted to balance the covariates directly reach the precision no estimator can beat — while staying exactly balanced, and silent, on every moment they were not told about. 1 generator, 6 views

A standard error for a model that is wrong

Every standard error a regression prints is a statement about a model that is true, and the repair for a model that is not is one of the most-quoted lines in applied work. It is priced here rather than recommended: the robust error reads a spread the model-based one does not, and its 95% interval covers 88.73% at twenty rows, is the worse of the two under mild heteroskedasticity until fifty, and is the smaller of the two whenever the error variance sits in the middle of the design rather than at its edges. The four corrections differ in one number — what each does with a point's leverage — and on a design with a single far-out point that number takes one of them to 0.3191 of the truth and another to 5.1127 times it. 1 generator, 7 views

A variable that moves one thing only

An instrument identifies a causal effect by assuming that a path no data can check is exactly zero, and it charges for that assumption in variance. Both prices are computed rather than described: a violation nobody can see is divided by the first stage, so the strength an instrument needs is 2.78 times the violation it is assumed not to have, and the conventional interval turns out to fail not when one instrument is weak but when many are — a failure each row's own first-stage leverage explains, and leaving that row out repairs at a price in width that depends on how strong the instruments are. What is left identified is the effect among the units the instrument actually moved, which at the standing compliance profile is 1.1000 against a population average of 0.5000. 1 generator, 6 views

The value that is not there

A missing value is missing conditional on something, and which something decides everything that follows. Dropping the incomplete rows leaves a regression slope exactly right when the chance of being observed depends on the regressor, however strongly, and wrong by a quarter of itself when it depends on the outcome — and the data cannot tell those two cases apart. Filling the gaps in is not free either: one filled value is not an observation, and the arithmetic that makes several of them into one is two corrections rather than the one everybody quotes. 1 generator, 8 views

The tail past the last observation

The distribution a sum converges on has one limit; the distribution a maximum converges on has three, and every extreme-value analysis is a claim about which. The claim is also an extrapolation: a hundred-block level read from fifty blocks sits above the largest reading in the record half the time. So each measurement here reports not what the estimate is but how much of it came from data. 1 generator, 7 views

What conditioning on a variable does

Putting a covariate into the regression is one arithmetic operation, and it is the right thing to do in one of the three worlds it could have come from. Here the three worlds are built to share a covariance matrix entry for entry, so the estimate they disagree about by 0.348 is a number no sample of any size can settle, and the rule that controls for everything measured is measured: it leaves a larger bias than controlling for nothing on 65.5% of four thousand structures. 1 generator, 7 views

The other ways in

FieldsThreadsSeriesConceptsAll essaysSearch