Fields
The distribution, and what a claim about it means 14 Models fitted to data 13 Many questions at once 29 Series in time 6 Where the runs go 7 Who gets which arm 15 When to stop, and what it costs 8 What makes any of it checkable 1
The distribution, and what a claim about it means
The law that sums converge on, the interval whose coverage can be summed rather than estimated, the test's second number, and the reversals that follow from arithmetic rather than from anybody's error.
The distribution itself
Sums converge on it, which is the most-quoted theorem in the subject. What the quoting leaves out is the rate: the middle converges quickly and the tail does not, and the tail is where the approximation is actually read. The same theorem carried through a function stops being a normal limit where the function is flat, and carried through a ratio it forbids every interval that is always finite.
Intervals, counted
An interval that claims 95% is making a checkable statement about a procedure. Build every possible sample and count. The interval taught first fails, the failure is worst where proportions are most often reported, and more data does not monotonically help.
Tests, and the second number
A p-value alone cannot be read: the same 0.04 means different things at different sample sizes, and nothing at all without knowing how many analyses were available. Under a real effect it is a draw from a distribution several orders of magnitude wide, a replication of a result at 0.05 succeeds exactly half the time under any model centred on that result, and the standard ways of combining several p-values disagree about which evidence counts. Every figure here carries the number that makes it interpretable.
The reference distribution the design supplies
Hold the outcomes fixed, re-run the rule that assigned them, and count. A p-value built that way needs no assumption about the outcomes, no large-sample argument and no nuisance parameter — it needs to be told the rule, which is the one thing the experimenter already knows. It repairs the adaptive design the previous field leaves broken, and it is not free.
Reversals that are not errors
Simpson's reversal, the base rate, regression to the mean. Each is normally taught with one famous table. A table is a point; these are swept, so how much of the space behaves that way and how large it can get both have answers — including why two correct analyses of one baseline disagree, how far a group enrolled on a noisy reading falls with nothing done to it, and why the correlation predicts the regression of the extremes only when the population is normal.
The tail past the last observation
The distribution a sum converges on has one limit; the distribution a maximum converges on has three, and every extreme-value analysis is a claim about which. The claim is also an extrapolation: a hundred-block level read from fifty blocks sits above the largest reading in the record half the time. So each measurement here reports not what the estimate is but how much of it came from data.
Coverage without a distribution
Every interval counted here so far has been found short of what it promised. This one is not, and the reason is that it is not a statement about the data at all: the rank of a future observation among a set of exchangeable calibration scores is uniform, so an interval built at the right order statistic covers a stated share of the time at any sample size, under any distribution, around any model however wrong. What it does not say is where that coverage sits — and everything the procedure does not know turns up in the answer to that rather than in the total.
The interval that holds observations, not a mean
Every interval counted before this one is about a parameter. These are about the values themselves — what a band actually holds, and how often. Two sample standard deviations around a sample mean of ten observations holds 91.1% of the population on average, and less than the advertised 95.45% on 59.9% of samples, so the average is the reading that hides it. The interval between a sample's smallest and largest reading holds a share whose distribution is Beta(n − 1, 2) for any continuous population at all, and buying with no assumption what normality buys at ten observations costs ninety-three of them. That number is the exchange rate between an assumption and data, priced in observations. And a band for all of the next m observations has no ceiling at the tolerance factor: from ten observations, holding all of the next ten needs 3.716 standard deviations, already past the 3.382 that holds 95% of the population, while a 95% prediction interval holds all ten only 67.9% of the time.
When the stratified answer and the pooled one disagree
Simpson's reversal is normally taught once, on one table, as a warning about confounding. None of the reversals here is confounded. A trial randomised by a coin, on a population where the treatment helps in both groups, reports it losing overall on 3.40% of trials at eighty units — and stratifying the randomisation takes that to zero at every size. An odds ratio of exactly 2.5 in all five strata pools to 1.789, because an odds ratio is not a weighted average of odds ratios, while the risk difference on the same table is exactly its stratum value. And where the grouping variable is something the treatment caused, both answers are right and answer different questions.
Shape, and what it does to a two-sample test
A t interval survives a skewed source in total and not in its tails. On an exponential source at 120 observations a 95% interval covers 94.81% while missing below the mean on 4.08% of samples and above it on 1.11%. The pooled two-sample test's size runs from 0.55% to 18.91% across forty units with a true null in every cell, and Welch's stays between 4.63% and 5.51% — on normal populations. On skewed ones Welch's tails are decided by one number, the skewness of the difference of the two means: two identical exponentials balance at twenty and twenty and reject low on 7.16% of samples at eight and thirty-two. And the side a decision reads needs its own repair: widened until its total coverage is exactly 95%, a t interval's upper limit on thirty exponential observations is still exceeded on 4.69% of samples, where Hall's transformation takes it to 3.31%.
Two tests, a threshold, and the rate they are read against
A predictive value is computed from a sensitivity, a specificity and a prevalence, and each of the three is softer than it is quoted as being. Two positives from a 90/95 test on a one-in-a-thousand condition are worth 24.49% if the tests' errors are independent, 10.16% at a correlation of 0.1, and 3.16% at 0.5 — barely more than the 1.77% one positive was worth. The sensitivity and specificity are not two properties of a test but one property read at a threshold somebody chose, and the threshold that minimises harm moves by three standard deviations of the score across the prevalences such a test is used at. The prevalence itself is usually estimated from the same test's positive rate, which at one in a thousand reads 5.09%.
Past the first term of the normal approximation
The Berry–Esseen bound guarantees how far a standardised sum can be from the normal, and on an exponential source it is 8.62 times the real worst error at every sample size — a constant set by a fair coin, whose error is a half-step, and at a hundred draws larger than the 2.5% tail it would be asked to certify. One Edgeworth term takes the normal's error at two standard deviations from 38% to 8% on ten draws, and turns negative in the short tail at every sample size, beyond a threshold that recedes only as the sixth root of n. The saddlepoint approximation is built at the threshold rather than at the mean, and reads a tail six standard deviations out to within 0.19% on five draws and within 2.2% on one.
A proportion's interval near the boundary, and the coin
Wilson's interval for a proportion wobbles a point or two around 95% away from the boundary and has a hole near it: its worst coverage is 83.50% at ten trials and 83.81% at a thousand, at an expected count of 0.1765 each time, because below that count the interval built on one success lies wholly above the truth. The intervals that never fall below 95% pay in average coverage, and the two of them pay for different promises — Clopper–Pearson holds each side under 2.5% and is 10.4% wider than Wilson at thirty trials, Blaker holds only the total and is 4.6% wider. Only an interval that adds a uniform random draw to the count covers exactly 95% at every proportion, at 0.9% more width than Wilson, and it is not used because two analysts with the same data would report different intervals.
An interval read beside something else
A 95% interval is almost never read alone. Read against a replication's estimate it captures an equal-sized one 83.42% of the time, because both estimates are uncertain, and ten replications of one study miss together: none of them lands outside 31.40% of the time against 16.32% if they were independent. Read against a second interval, two that just touch mark a difference with p = 0.0056 rather than 0.05, standard-error bars that touch mark p = 0.157, and in a figure of ten identical groups some pair of 95% intervals fails to overlap one time in seven. Read against a second verdict, two studies of one effect at 50% power disagree about significance half the time, and when they do the difference between them is significant in 9.75% of cases.
Models fitted to data
A line through a cloud, a prior doing visible work, groups that are neither the same nor separate, and observations that stop before the thing being measured happens.
Regression, and what the summary hides
A slope, a standard error and an R² can all be computed from data the model is grossly wrong about, and none of them says so. Leverage is a number known before the outcome is looked at, influence is a number, and whether a residual plot looks bad is a question with a calibrated answer. Two wrong rows placed together hide from every single-row diagnostic, and a loss that bounds a large residual does nothing about a far one.
The prior, doing visible work
A credible interval says the thing everyone wants a confidence interval to say, and it needs a prior to do it. So the prior is treated as a component with a measurable effect: it is worth a stated number of observations, and the interval it produces has a coverage that can be summed over the sample space like any other.
The spread, and its own uncertainty
Partial pooling estimates the population spread from the group means and then uses it as though it were known. It is not known: eight groups leave τ anywhere in a range that spans a factor of several, and on a third of eight-group datasets the usual estimate of it is exactly zero. Carrying that uncertainty rather than dropping it is the difference between an interval that covers 95% and one that covers 79%.
Groups that borrow
Eight hospitals are neither one hospital nor eight unrelated problems, and the estimate between those two answers is not a compromise. It is a weighted average whose weight — se²/(se² + τ²) — is decided by how well each group is measured and by nothing else, and the population spread in it is estimated from the data rather than assumed.
Hierarchy past one number
The same weighted average, on quantities that are not means. A slope, where how much a group borrows is decided by the arrangement of its x values rather than by how many it has. A proportion, where the pooling has to happen on a scale the data does not live on. A second level of grouping, whose arithmetic turns out to be the effective sample size of the time-series field arrived at from the other direction.
What partial pooling does to one group, to the set, and to a ranking
Partial pooling minimises the total squared error across groups, and three uses of its estimates are not the total. One group: with each group's standard error equal to the population's spread, every group truly more than 1.73 widths from the centre — 8.33% of a normal population — is estimated worse than by its own mean, and among eight groups with the spread estimated the truly most extreme one is worse off in 61.6% of datasets; capping each group's shift at one standard error improves the total, the truly extreme group and the observed extreme group at once. The set: posterior means spread 0.707 as widely as the truth and count 0.234% of groups beyond two widths against a true 2.28%. The ranking: in a league table of a hundred groups of unequal size, raw means fill 62.1% of the top ten with small groups, posterior means 13.4%, and no ranking recovers more than 5.47 of the true ten.
When the data stops early
A subject still event-free when a study ends is not missing and not observed — it is known to exceed something, which is a third state most tools have no slot for. The estimator that gives it that slot recovers the curve, and everything read off the curve afterwards has a condition attached: the interval printed around its end stops covering where the curve is read, a dropout that carries information leaves a record identical to one that does not, one minus the curve is not a risk once something else can end observation first, and a hazard ratio is an average whose weights the length of follow-up chose.
What conditioning on a variable does
Putting a covariate into the regression is one arithmetic operation, and it is the right thing to do in one of the three worlds it could have come from. Here the three worlds are built to share a covariance matrix entry for entry, so the estimate they disagree about by 0.348 is a number no sample of any size can settle, and the rule that controls for everything measured is measured: it leaves a larger bias than controlling for nothing on 65.5% of four thousand structures.
The value that is not there
A missing value is missing conditional on something, and which something decides everything that follows. Dropping the incomplete rows leaves a regression slope exactly right when the chance of being observed depends on the regressor, however strongly, and wrong by a quarter of itself when it depends on the outcome — and the data cannot tell those two cases apart. Filling the gaps in is not free either: one filled value is not an observation, and the arithmetic that makes several of them into one is two corrections rather than the one everybody quotes.
A variable that moves one thing only
An instrument identifies a causal effect by assuming that a path no data can check is exactly zero, and it charges for that assumption in variance. Both prices are computed rather than described: a violation nobody can see is divided by the first stage, so the strength an instrument needs is 2.78 times the violation it is assumed not to have, and the conventional interval turns out to fail not when one instrument is weak but when many are — a failure each row's own first-stage leverage explains, and leaving that row out repairs at a price in width that depends on how strong the instruments are. What is left identified is the effect among the units the instrument actually moved, which at the standing compliance profile is 1.1000 against a population average of 0.5000.
A standard error for a model that is wrong
Every standard error a regression prints is a statement about a model that is true, and the repair for a model that is not is one of the most-quoted lines in applied work. It is priced here rather than recommended: the robust error reads a spread the model-based one does not, and its 95% interval covers 88.73% at twenty rows, is the worse of the two under mild heteroskedasticity until fifty, and is the smaller of the two whenever the error variance sits in the middle of the design rather than at its edges. The four corrections differ in one number — what each does with a point's leverage — and on a design with a single far-out point that number takes one of them to 0.3191 of the truth and another to 5.1127 times it.
What a summary of a scatter is a property of
R-squared is read as a property of a relationship and it is a property of a design. Five studies differing only in how far apart they placed their readings report it from 0.021 to 0.849, with the same line and an estimated residual spread of 0.993 in every one. For a simple regression the squared t statistic is (n − 2) times R-squared over (1 − R-squared), exactly, checked to sixteen significant figures across five hundred fits — so a report carrying both has carried one number twice. A coefficient that is zero only under independence does exist: distance correlation separates the four datasets built to share a correlation by 0.10, and the rank coefficient that guarantees nothing separates them by 0.49.
What a diagnostic plot is showing
A quantile plot is read by eye against a band drawn for one point at a time, and a genuinely normal sample of forty has forty chances to leave it: 45.0% of them do. The band a reader is actually using is that one widened by a factor of 1.502, and nothing draws it. The quantity is wrong as well as the width — a t interval needs the sampling distribution of the mean to be normal rather than the data, and the source that leaves the band on every sample of forty covers 94.80% while one that leaves it on 57% covers 95.73%. And what is plotted are not the errors: a design whose leverages run from 0.045 to 0.663 gives residuals whose spreads differ by a factor of 1.68 with the model exactly right.
Many questions at once
What a correction controls and what it does not, the best of a set against the cost of having searched for it, a table of fitted models where the benchmark is one candidate among many, the dependence a criterion is told to assume and has to estimate instead, what a break point costs once it has been looked for rather than named, what a joint fit means for a covariance with no parameter in it, and what happens to any of these comparisons when the tuning parameter has to be chosen from the data as well.
Corrections, and what each controls
Bonferroni bounds the chance of any false positive; Benjamini–Hochberg bounds the share of the findings that are false. Both get called correcting for multiple comparisons and they are different promises, so every procedure here is made to report both rates and the power each one costs.
The best of a set, and what the search costs
Comparing two forecasters is a test. Comparing eight is a multiplicity problem on top of a dependence problem, and the two do not separate: eight windows of one series carry the multiplicity of about two independent comparisons and eight separate problems carry eight, so a correction that charges for the number of models is wrong in both directions. The reference distribution has to be over the whole set — and once it is, the models nobody would have run turn out to cost more than the ones that were close.
Searching among fitted models
The set field compares forecasts nobody estimated. A specification search compares a benchmark with a table of variants that all contain it, and two things change at once: every variant is behind before the search begins, by an amount with a closed form in the shape of the table, and the reference distribution for the winner can no longer be resampled from the data — it has to be generated from a model. The repair the nested case asks for, applied row by row, takes a table from its nominal level to a quarter.
A search with no fixed point
The search field compares a benchmark with variants that all contain it. Take the fixed point away — let the benchmark be one candidate among sixteen, selected by the same data as its rivals — and the closed form for what every candidate is behind by survives intact, because it was never about nesting but about counting parameters. Three readings of one true null give 2.5%, 8.0% and 76.8%; Bonferroni is three times too strong on one table and not strong enough on another; and searching the benchmark makes the table harder to reject with rather than easier.
Scoring a search without spending data
A hold-out is expensive: every row spent scoring is a row not spent fitting. There are two standard ways of not paying — an information criterion, which predicts the hold-out from the in-sample numbers, and a resampling, which builds a reference distribution from one sample. Neither avoids what it replaces. The criterion's penalty *is* the displacement, computed rather than counted, and it beats a rolling hold-out at every split there is; both it and the resampling are undone by the same defect, which is rows that repeat each other, because both are counting independent things and there are fewer of those than there are rows.
Counting what is independent
Akaike's penalty is 2q because the optimism of a fit is a trace and the trace collapses to the parameter count when the rows are independent. Repair the trace and the criterion gets worse, on either scale, because the row count entered twice and a penalty is the second place: whiten the fit and the ordinary penalty is correct again, which recovers 88% of what counting rows gives up. Estimating the dependence from the rows being selected on costs 8% of that. And a multiplier resampling cannot keep more dependence than the residuals have, which is a ceiling rather than a tuning problem.
Estimating the dependence, not naming it
Whitening a sample repairs a criterion, and the whitening that repairs it is told the dependence is a first-order autoregression and left to find one number. A real dependence has no parameter. The obvious estimate — the sample autocovariances, cut off at some lag — is not a covariance matrix on half the draws there are, so the rule built on it does not exist; the tapered estimate that is always a covariance matrix costs a further seven points of what the repair is worth. And the two constructions named as escaping the resampling's ceiling turn out to be one construction and one identity: a moving block attenuates exactly as a blocked multiplier does, because the attenuation is the join.
The shape a dependence has
Estimating a covariance rather than naming it was priced on errors that really were a first-order autoregression, where generality can only cost. Here are four laws with the same first lag and nothing else in common: a moving average that stops, a memory that does not, a break in the middle. The cost of generality where it is not needed and the benefit of it where it is turn out to be the same size — and against a covariance that changes with position rather than with gap, one number, a window and an order are worth exactly the same as each other. Both tuning parameters are then chosen from the sample, and every criterion available for choosing them is aimed at something else.
A block, weighted inside itself
Every resampling in this collection attenuates the dependence it is trying to keep, and until now every one of those attenuations was inherited rather than chosen. The construction the long-run-variance literature actually uses weights the residuals down towards each block's own ends — and what that buys turns out to be exactly the squared value of the window at the two ends and nothing else about its shape. It is an asymptotic gain that has not arrived at any block length a hundred and twenty rows can afford, and at the lengths that are available a tapered block is very nearly a shorter plain one.
Fitted together, or fitted after
Every whitening in this collection is a two-step rule: fit a candidate, read the dependence off what is left over, whiten, fit again. The step nobody examined is the first one. A fit removes memory as well as signal, and how much is arithmetic rather than noise — the residuals of a hundred and twenty rows report a lag-one coefficient of 0.73 where the errors report 0.78 and the law says 0.80. Estimating the coefficient with the line instead of after it recovers most of that, and iterating the two-step rule to its fixed point recovers nearly as much — because a fixed point of the sum of squares is not a maximum of the likelihood, and the difference between them is real, one-sided, and worth nothing at all in the decision the number feeds. The order the same criterion picks moves by nearly two, which is worth a great deal more.
Paying for a search
A two-regime whitening chooses its change point by searching a profile, and then reads a criterion that counts parameters. Nothing charges for the search. Nothing can, in the usual way: under no break the two regimes share a coefficient and the break point is not identified at all, so the likelihood ratio is a supremum over a parameter that exists only under the alternative and no count of restrictions describes its distribution. Searching a hundred and twenty rows for a change point that is not there manufactures about five units of ratio where one parameter costs two — and a rule that counts the break as free splits a stationary sample on 99% of draws. A second break costs as much again, on a sample that has only ever had one.
Where a taper's case begins
A tapered block was measured against a plain one on the exact bias each implies, and the answer was that the taper's advantage has not arrived at any block length a hundred and twenty rows can afford — with the crossing at ℓ = 20 and a difference there of three tenths of a point, too small for the critical-value table to resolve. Two of those three statements are about the wrong quantity. What a sample reports at ℓ = 20 is four points rather than three tenths, because the autocovariances the window is applied to are themselves attenuated and the window that discards the long lags loses less of them; and the comparison is made at a shared block length where each window has its own best one. Read at each window's own setting, on the error rather than on the bias, the ordering reverses at a hundred and twenty rows.
A covariance with no parameter
A regression's coefficients and its errors' dependence can be fitted together when the dependence is one number. When it is an estimated covariance there is nothing for “jointly” to mean — until a family is named, and then the family's own width is the parameter. Three things the measurement says, two of them the opposite of the guess: the plug-in's shortfall is the taper's rather than the data's and is nothing under two of four laws; nothing in the likelihood chooses the width, because nested families buy about a unit of it a lag and that is what a criterion charges; and a fit-only objective does not run away, because a unit diagonal fixes the trace.
How long the list is
A tuning parameter chosen per candidate rather than once for the table costs something, and the window's figure was measured while the order's was not — because their lists are different lengths and matching them changes what each rule is. Matched at every length, the two cost the same. What separated them was not the list at all: a per-candidate whitening is a different error model for every candidate, the Gaussian likelihood has a term that says so, and the sieve's rule in this collection never carried it.
Two searches over one sample
A break point that has been looked for costs more than a count of parameters says. So does a window chosen from a list, and nothing here had charged for it. A rule that does both is running two searches over one sample, and the charges do not add: the pair manufactures a fifth less than the two apart. What follows is sharper than a correction to a sum — most of what a break search finds under correlated errors is the correlation, so once a whitening has been chosen from the same sample the break's own charge more than halves, and carrying the published one across switches the test off entirely.
A block length chosen from the data
Two block windows were compared at each window's own best block length, which is the argmin of a quantity that needs the truth. Re-run on rules a practitioner could actually run, the ordering reverses: the taper wins at the best available length and at one estimated from the sample's own persistence, and the rectangle wins at a length written into a protocol and at the rule of thumb, at every sample size measured. And the whole argument is a third of the size of what estimating the length costs.
A charge for a covariance's own dimension
The two charges that pick a band's width were derived for a regression coefficient, and a band's numbers are neither free nor entered the same way. Measured as an optimism against a second independent sample, a Bartlett band costs 0.374 of a log-likelihood unit a lag where a criterion charges one. The window's own weights are the mechanism — scale them and the charge scales with them, to within two per cent across three windows — and they are not the arithmetic: the level is three quarters of the sum, a different shape at the same sum costs more, and the charge rises with the law's persistence. Levying the right one moves every width in the table by a factor of four and the error by half a per cent.
The rate and the size of a disagreement
A sweep of what it costs to let every candidate choose its own tuning parameter reported a product and called it a cost. Separated — and the decomposition is exact, because a draw on which nothing disagreed carries exactly zero — the rate rises by half across the list and what a disagreement is worth does not move at all. A third quantity explains why: a disagreement costs something only when it changes which candidate the table selects, that happens on about an eighth of draws, and the eighth does not depend on the list.
Two searches over different features
A break search and a window search on one sample manufacture less likelihood together than separately, and both read the same residual series. Put four more pairs beside them on a scale whose zero is two searches over independent columns and whose one is a search that contains the other, and the shortfall is a property of the pair: 0.00 for the control, 0.76 for the pair that reads one series twice, exactly 1 for containment — and −0.31 for a break paired with an independent column, where charging the two separately under-charges rather than over-charging.
A probe chosen rather than picked
A diagnostic that reports on what a balancing rule was not handed is run through a column that is 92% inside the span the rule balanced — because orthogonality in the population is not orthogonality on fourteen units. Projecting the probe off that span costs one least-squares fit and triples the separation it carries; choosing the direction from the design's own leverage is worth another factor of four over a random one; and the projection-pursuit direction the argument invites is worse than random. On a short chain the raw probe misses 44% of the sets that are split and the projected one misses 12%.
Both halves of the dependence at once
One field varies the marginal with the copula held Gaussian and another varies the copula with the marginal held normal, and the natural guess is that the two leaks compound. They do not add in either direction: eleven of twenty cells cancel and nine compound, a mildly skewed covariate under a lower-tail copula leaks 0.002% where adding the two gives 16.6%, and a heavy-tailed symmetric covariate that leaks exactly nothing on its own doubles what an asymmetric copula leaks. A median split's zero survives all thirty combinations at under 10⁻¹⁶.
The block length read on a quantile
An ordering between two block windows reverses depending on who chose the block length — measured on an implied long-run variance, which is the instrument that makes the sweep affordable and is not what anybody reads. Read on the 95% point a test uses, the tapered window wins under all four rules; read on the coverage the interval delivers, it wins under all four again. Two of the four rules change sign, and they are the two the recommendation was about. No rule and no window reaches its promised coverage: the eight cells run from 80.8% to 91.0%.
A charge that is not a straight line
The measured charge for a tapered covariance band falls from 95% of its summed weights at two lags to 75% at thirty, so every rule that levies it as a straight line through the origin is too dear at one end and too cheap at the other. The deferral asked for a curve; the answer is that the curvature is in the denominator. Counted in the pairs the band actually uses — a lag of k is an average over n − k products — the same readings are flat from four lags up, at a spread of 0.029 against 0.081, on a correction with no fitted parameter. It repairs the one window the earlier field measured and leaves the other three wanting a curve.
What decides whether a tuning list decides
The probability that a per-candidate tuning list changes which candidate a table selects is reported flat at about an eighth across list length, on a table and a world that are never varied. Vary how far apart the candidates are — one multiplier on the omitted coefficients, everything else held — and it runs from 17.6% to 1.5%, while the disagreement rate it is a factor of rises from 31.8% to 88.8%. The world where the candidates quarrel most about the tuning parameter is the world where the quarrel matters least, and a nested table turns over less rather than more.
Overlap and complementarity, separated
How much two searches over one sample share is measured as the net of two effects: ground both of them find, and configurations the joint search reaches that neither slice contains. Pin the first search at its own answer and search the second, and the two separate exactly — the pinned supremum cancels, so the split adds back to the original number on every draw. The control the whole scale is anchored on reads zero because its two components are several times larger and cancel, and the split depends on which of the two searches is pinned while their difference does not.
A probe from what the rule blocks
A balancing rule breaks the admissible set into pieces by blocking exchanges, so the quantity a diagnostic should be aimed at is the constraint's active set rather than the design's leverage — which is a heuristic about the same thing. Built from the design and the tolerance alone it is a real probe, well ahead of a random direction; it is also behind leverage at 4.4 paired standard errors. Counting the active set exactly, at a cost no trial can pay, makes it worse rather than better, so the approximation was never what cost it.
The same table at seven correlations
Every cell of the copula-by-marginal table is measured at one rank correlation, and eleven of its twenty cells cancel there. Sweep the correlation from 0.1 to 0.7 and four of the twenty change the sign of their answer, all four from compounding to cancelling, all four at the most skewed covariates. The near-perfect cancellation that is the earlier field's headline is a crossing: the cell passes through zero at a Spearman of 0.38, two hundredths from where it was read, and is two orders of magnitude larger by 0.7.
The interval, studentised
A percentile interval inherits the resampled distribution's skewness and its scale error together, and the standard repair is to resample a t-statistic so each resample carries its own scale. It is one extra variance per resample, and it repairs nothing: not one of the eight cells reaches its promise, the interval is twice as wide for half a point of coverage, and at the block lengths the rules choose the scale rests on two or three whole blocks. And the ordering between the two windows reverses at all four rules, on resamples that are the same resamples.
The analyses that were available and not run
A correction divides by the number of analyses, and that number is the one quantity nobody measures. Twenty analyses of one dataset are worth 11.37 independent ones at a correlation of 0.6 and 2.58 at 0.95, so the threshold controlling a stated error rate can be counted rather than assumed. What the correction costs is charged in the estimate rather than the error rate: a larger statistic is a more selected one, and what survives averages 1.69 times the truth after correction against 1.35 times before it. Naming the analysis in advance is the alternative, and it is worth exactly what correcting all twenty is worth when the chance of having named the right one is 38%.
Series in time
The independence every other field assumes, two series that move together, three that carry a count, the observation that has not happened yet, and whether one forecaster is better than another.
When the observations repeat each other
Every standard error on this site divides by √n, which claims the observations carry independent information. In time order they usually do not: at a lag-one correlation of 0.8 a fifty-point series is worth about six independent observations, its 95% interval covers 47%, and two series that wander are called related three times out of four.
Series that move together
Two random walks regressed on each other are called related three times in four, which is why the time-series field ends in a warning. The exception it names and does not measure is here: when the pair is genuinely tied, the fitted relation converges at rate 1/n rather than the usual 1/√n, and the test that separates the two cases cannot be read against any table that exists.
Three series, and a count
A pair is either tied or it is not, so its whole inference is one test with one answer. Three series can carry none, one or two relations at once, and the quantity being estimated stops being a slope and becomes an integer — read off the gap in a spectrum, against a critical value that depends on how many things are left wandering and on nothing else.
The observation that has not happened
Two fields here stop at estimation. A forecast is the other question — not what the parameter is but what the next observation will be — and the band round it is computed by substituting estimates into a formula derived for the truth. Counted, that 95% band is not 95%, the shortfall grows with the horizon, and the model has to beat two benchmarks that estimate nothing at all.
Comparing two forecasters
The forecast field ranks forecasts by mean squared error and stops. Whether one forecaster is really better is a test, its terms are dependent because neighbouring forecasts overlap, and where one model contains the other it fails completely — declaring the smaller model significantly better while the null it claims to test is true, and more certainly the more data it is given.
A forecast that is a probability
The field before this one ranks point forecasts by squared error. A probability forecast can be held to something stronger and stranger: say 30% often enough and about three in ten of those days should happen, which is a claim anybody can check by counting. Calibration alone passes a forecaster that issues the base rate every time, the diagram that draws it charges a blameless forecaster in proportion to how finely it was binned, and an honest record of a hundred forecasts shows more apparent miscalibration than the amount routinely read as evidence of a problem. A score that is not proper pays a forecaster to answer only 0 or 1, and what that liar gives up in ranking and resolution is computed rather than described.
Where the runs go
Decisions taken before any data exists: which settings to measure at, how to choose a design rather than look one up, what the criterion assumes, and what to do when the answer depends on the answer.
Decided before the data
Every other field here repairs an analysis after the fact. These figures are about the arrangement that makes the repair unnecessary: blocking removes exactly the variance the blocks carry, randomisation buys a known reference distribution rather than balance, and varying every factor at once estimates each of them from every run.
The surface between the corners
A two-level factorial answers which factors matter and is structurally unable to answer what setting is best: every run sits at a corner, every squared term is 1 there, and the column that would estimate curvature is a copy of the intercept. What it takes to see a curve, how far a fitted gradient can be trusted, and why the location of an optimum is a ratio of estimates rather than an estimate.
A design chosen rather than looked up
Every design here so far came from a catalogue and was then measured. Turn the arithmetic around and a design is the answer to an optimisation: maximise a functional of X′X and see what comes back. What comes back is the catalogue's own nine settings, re-weighted — and an exact theorem that says when a search is finished without knowing what it was searching for.
The criterion, and what it assumes
D-optimality gave the catalogue back, which was reassuring, and left two things unsaid. A, D and E are one family and the letter is a choice — the design that wins the first is nearly worst at the last. And for a model whose information depends on its own parameters, a design is optimal only at a guess about the answer, which makes the cost of guessing wrong a quantity with a closed form.
A design that assumes less
A design for a non-linear model is optimal only at a guess about the answer. Two ways out: protect the worst parameter value in a range rather than the average, which is an optimum that sits on a tie rather than a slope; or stop guessing, run part of the experiment, and design the rest at the estimate — where the interesting question turns out to be what that does to the interval afterwards.
Weighting one sample into another
A weight turns the sample that was assigned into the sample a coin would have assigned, and the exchange is exact: integrated over the population, the standardised difference on every covariate goes to machine zero whatever the assignment rule was. What it charges is observations — a treated arm worth 94% of itself where the assignment is nearly a toss-up and 12% of itself where it is nearly decidable — and past that point trimming does not repair the estimate, it replaces the question. Weighting by an estimate of the probability turns out twice as precise as weighting by the probability itself, and weights fitted to balance the covariates directly reach the precision no estimator can beat — while staying exactly balanced, and silent, on every moment they were not told about.
What a sample-size calculation was given
A sample-size calculation takes a standard deviation, an effect and an outcome as given, and each is usually a choice or an estimate. Sized from a pilot of ten's standard deviation, 55.9% of trials have less than their planned 80% power and 11.1% less than half, although the planned sample is right on average; sizing from the pilot's 80% upper confidence limit leaves 19.8% short at 1.65 times the patients. With the effect believed to be half a standard deviation give or take a quarter, a trial of sixty-four per arm succeeds 69.2% of the time, and no trial of any size can succeed more often than the planners believe the effect is positive. Cutting the outcome at a threshold keeps at most 63.7% of its information, and a responder threshold one standard deviation out needs twice the patients.
Who gets which arm
Splitting the units, changing the split while the trial runs, balancing on what was recorded before the outcome, the shape a covariate has to enter by for any of it to be the right criterion, which of those guarantees survive a covariate that is not normal and which survive a different joint law of the ranks, and whether the set of splits a rule admits can be walked at all — before the trial, and afterwards on the statistic it reports.
Splitting the units
The only design decision that costs nothing: the same units, the same measurements, the same analysis, and a different variance. Sample the noisier arm more, the expensive one less, and the shared control by the square root of the number of arms — and then notice that every one of those rules is a function of a quantity nobody has, and measure what happens when it is estimated instead.
Designs that change while they run
The stopping-rule field is a fixed design looked at more than once. Here the design itself is a function of the data — how many units, which arm the next one goes to, which arms survive the interim — and the question stops being what the rule spends and becomes what it leaves behind. The estimate from an arm chosen for being ahead is ahead by more than it should be, and the unbiased estimate is the one that throws away the data the choice was made on.
Balancing on what was recorded first
The adaptive field's rules read outcomes and every one of them broke something. These read baseline covariates, and each holds one imbalance flat while the others grow behind it exactly as a coin's do. What follows is the opposite of the outcome-adaptive story at every step: the unadjusted analysis is too cautious rather than too eager, and the exact test that cost nineteen points of power there costs almost nothing here.
More arms than two
Minimisation balances a trial by making the arms' counts even inside every factor level. With two arms there is one way to measure how uneven two counts are. With three there are several, they are all called minimisation, and they send different patients to different arms — while the ratio a trial was designed to deliver quietly disappears unless the score was told about it.
Balancing what has no levels
Every balancing rule in the two fields before this one reads a level. Age, blood pressure and a baseline score have none, and the first thing that happens to them is that somebody invents some — a choice with a cost available in closed form before any data exists: a median split can see exactly 2/π of a normal covariate, so a rule that balances its two halves perfectly still leaves three fifths of a coin's imbalance. A rule that reads the number instead does not beat that by a factor; it beats it by a rate.
The shape the covariate enters by
Every balancing rule in the two fields before this one optimises one over the variance of the treatment estimate in a model where the covariate enters linearly — which is not an assumption about the analysis, it is what the criterion is. Balancing the mean of a covariate removes exactly 2/π of the imbalance in a median split of it, a quarter of a threshold in the tail, and nothing at all from a quadratic. The ranking of the rules reverses between shapes, and against one of them every rule here is worse than a coin.
Choosing what the rule reads
A balancing rule cannot tell one basis from another with the same span, so choosing what to hand it is choosing a subspace, and everything about it is exact geometry: what a rule removes of a shape is that shape's squared multiple correlation on the span. The maximin basis over a list of shapes is then a finite problem with an exact answer — and over a whole subspace the guarantee is exactly zero rather than small, which is why the list cannot be avoided. Randomising which functions the rule reads is worth twice the best fixed choice.
When the set is too large to walk
Two constructions that were each measured by enumerating everything, at the size where enumerating everything stops being possible. A balancing dictionary over two covariates is an outer product, every inner product in it is still closed form, and a rule holding every main effect of both removes exactly nothing of any interaction. A trial of two hundred units has assignments that cannot be walked — and the acceptance rate turns out not to depend on the size of the trial at all, so the exhaustion a sixteen-unit count runs into is a fact about sixteen units.
When the two are not independent
The dictionary that is an outer product needs the covariates independent, and the count that is a rate times a binomial coefficient needs the draws independent. Neither holds in a real trial. A linearisation turns the missing fourth-order expectation into Mehler's formula applied to products, which makes the whole geometry closed at any correlation — and the result it was built to test does not survive: what a rule removes of a pure interaction is exactly nothing at independence and 4ρ²/(1+ρ²)² everywhere else. A walk on the admissible set is exactly uniform, needs a burn-in, and is dearer than hunting until almost nothing is admissible.
A cut point, at a correlation
The geometry of a balancing dictionary over two dependent covariates is closed for polynomials and was taken to be asymptotic for cut points, because a threshold's Hermite coefficients never terminate. Conditioning on the second variable closes it exactly: every mixed inner product is ρ^j times a one-variable answer and the only two-dimensional object left is an orthant probability, which at the median is (2/π) arcsin ρ. The truncation that was feared falls geometrically in the correlation rather than algebraically in the order — and the interaction guarantee a correlation destroys for powers survives it exactly for median splits, because a two-valued function squares to a constant.
What a dictionary buys and what it costs
A balancing rule is a list of functions, and the list decides two things that pull against each other. What it protects against is decided by parity: an interaction between two odd functions is even, an odd function is orthogonal to an even one at every correlation, and the two things every trial balances — a mean and a median split — are both odd, so their worst case is exactly zero however many of them are held. What it costs is the assignments it leaves, and past a certain thinness those stop being one set: the admissible assignments split into an arrangement and its mirror image, no sequence of admissible single swaps joins them, and the walk that samples them is uniform on half the reference distribution for ever.
A guarantee that needed a symmetry
What a balancing rule can remove of an interaction is decided by parity, and parity is a statement about a symmetry of the law rather than about the covariate. Keep the dependence and change the marginals — a Gaussian copula under a monotone transformation — and the two things every trial balances come apart. A median split is a function of the sign of the latent normal whatever the marginal is, so its exact zero survives every transformation to the last digit. A mean is odd only when the marginal is symmetric, and its zero is gone at a skewness of one. The separating case is a marginal that is heavy-tailed and symmetric, where every zero holds exactly: it is not normality the guarantees needed. And a threshold at a value on the covariate's own scale never had one at all.
What a chain cannot report
A thin admissible set falls into an arrangement and its mirror image, and the walk that samples it is uniform on half the reference distribution for ever while every diagnostic passes. That was found by enumerating fourteen units, and enumeration stops at about twenty-four. The test that does not enumerate is two chains — one from an assignment, one from its complement — compared on a statistic that changes sign under the complement, with a third chain from the same starting point to say whether a large reading is a fact about the set or about the length of the run. It agrees with the enumeration at every tolerance where the answer is known. At two hundred units it reports something else: the walk reaches the whole set at the tolerances a trial would use, and past a point the diagnostic stops agreeing with itself.
The diagnostic after the trial
The test for whether a balanced-assignment walk can reach the whole admissible set is run on a covariate function, before any outcome exists. Run instead on the difference in arm means it is the same test — under the sharp null the outcome is a fixed column — and it is about the statistic the p-value is actually built from. Two findings: an outcome is a probe nobody chose, and on a set that is genuinely split three in ten of them see nothing at all; and the published two-sided p-value is exactly right on half a reference distribution, while a one-sided one is not.
The other half of the dependence
What a balancing rule can remove of an interaction was measured across six marginals with one joint law of the ranks held fixed. Change the copula instead and the two exact zeros come apart for two different reasons. A median split's zero is arithmetic — a centred median split squares to a quarter identically — and holds under every copula there is. A mean's zero needs the copula to be symmetric under reflection as well as the covariate to be symmetric: it is exactly nothing under a Gaussian, t or Frank copula and 7.71% under a Clayton, at the same rank correlation and with a normal covariate throughout.
When to stop, and what it costs
Looking at the data as it arrives, promising a precision rather than a sample size, keeping a procedure from reading what it will report, pacing the blocks the promise is measured in, and making the same promise about a difference instead of a mean.
Stopping rules
A p-value is defined relative to a sampling plan, so two experiments with identical data and different stopping rules have different p-values. That sounds philosophical and it is arithmetical: testing five times at the nominal level rejects a true null 14% of the time.
What the design is asked to guarantee
Two halves of one question that were never separated: which parameters a design is for, and how much precision is enough. Protecting one parameter of a non-linear model over a range of its own values is a worst case of a ratio of determinants, and it is not a special case of either problem it is made of. Letting the experiment stop when it is precise enough is the adaptation with a theorem against it — the rule stops when its own noise estimate is low, so the interval it produces is short.
What the procedure may not read
Two restrictions that were cheap until now. A design for a non-linear model may not read the parameter values, and where the model has two of them the guess is a point in a rectangle: the weights that were exactly 1/√2 become a function, and a design robust in one coordinate turns out to guarantee no more than one built at a single point. A stopping rule may not read the mean it will report — and there is an exact way to obey that, at a price in the width of the interval rather than in the number of observations.
The block size as a schedule
The blinded rule's block size is a dial between an interval's degrees of freedom, a stopping rule's, and the overshoot. Letting it change during the run keeps the coverage exact — the argument never needed the sizes to be equal — and then runs into an identity: the two kinds of degree of freedom add to N − 1 on every run, so a schedule cannot make more of both. What it can do is spend each where it is worth most, which is worth a few per cent and is not the free lunch the dial looked like.
A promise about two arms
The fixed-width interval whose coverage is exact was built for one mean, and almost nothing anybody runs an experiment for is one mean. The construction survives with the harmonic effective size in place of the block size, and four things do not: a unit of that size costs four observations rather than one, the second variance puts Neyman's allocation in reach of a blinded rule, the conservation identity comes up one degree of freedom per block short, and there is an exact estimator under either of two conditions and none under both.
The weights the corner needs
A fixed-width interval about a difference is exact when the arms share a variance or the allocation ratio is constant, and exact under neither when both fail. It is exact there too, with h_b(λ) = (1/m_A + λ/m_B)⁻¹ — the inverse variances, written as a function of the variance ratio alone, which is a contrast and so is readable by a blinded rule. The estimated precision weights that had no theorem behind them turn out to be that rule at an estimated ratio. And an interval's own scale estimate is right in exactly two cases: inverse-variance weights, and equal ones.
What a block may vary
Two questions from two corners of the collection with the same answer. A walk over the admissible assignments that exchanges more than one unit per arm mixes faster and is refused more often, and the trade is exactly computable: the gain is a factor of six where a hunt is fifty times cheaper anyway, and nothing at all where the comparison is actually decided. And a variance ratio that drifts between blocks cannot be estimated inside the block it weights — but it can be modelled across them, once the bias in a log variance estimate is subtracted, because that bias depends on the degrees of freedom and the degrees of freedom alternate with the allocation.
When a fixed width is reached
A fixed-width interval promises a precision rather than a sample size, so the trial stops when it has enough — and what the stopping rule is allowed to read decides whether the interval means anything. A rule that stops when its own interval is short enough is stopping on the spread it is about to quote, at a correlation of 0.938, and every weighting loses three to five points of coverage including the one told every true variance ratio. The width a set of weights will produce is predictable from the within-arm sums of squares alone, which are independent of every difference the interval is about; a rule that stops on the prediction is back at nominal, costs two and a half blocks, and gives up half of the width promise it was keeping.
What makes any of it checkable
A seeded generator, a closed form beside every simulation, and p-values checked for flatness — the discipline the rest of the site is measured against.
Other ways through
a field is one of five
A field says what an essay is about, and it is the coarsest of the five organisations this site carries. The other four cut across it, which is the point of having more than one.
Series group essays by the idea each one makes an argument about, so a series is depth on a single idea rather than breadth across a subject — the first part introduces the idea and the last assumes everything before it. Threads are the motifs that recur where there is no reason to expect them: counting rather than claiming, one run being an anecdote, two routes to every number. Concepts is the index an encyclopedia gets for free and an essay collection has to build — every object named by more than one essay, with the essays that name it. Figures sorts the same material by the generator that drew it, which is the reuse ratio made visible. And search runs over all of it in the browser, with no request to anything.