The thread: Count it, do not claim it
A basis is a subspace
A balancing rule cannot tell one basis from another with the same span, so choosing what to hand it is choosing a subspace — and then what it removes of any outcome shape is a projection, computable exactly, with no trial anywhere in it.
A block size that changes
The blinded rule's exactness never needed the blocks to be the same size. Letting the size be chosen from the contrasts as the run goes on leaves the coverage exactly where it was — and runs straight into an identity that says what a schedule can and cannot buy.
A block weighted inside itself
The triangle every block resample attenuates by is not a fact about blocks. It is the self-convolution of a rectangle, and a block weighted down towards its own ends has a different one — whose leading term is the squared value at the two ends and nothing else about the shape.
A break that was looked for
A two-regime whitening finds its change point by maximising a profile, and then reads a criterion that counts parameters. Under no break there is no parameter to count, because every position describes the same model.
A covariance with no parameter in it
The whitening that repairs a criterion is told the dependence is a first-order autoregression and left to find one number. A real dependence is not one number, and the obvious estimate of it is not a covariance matrix.
A covariate with no levels
Every balancing rule on this site reads a level. Age and blood pressure have none, so somebody cuts them into categories — and a median split can see exactly 2/π of a normal covariate, whatever the rule does with the halves.
A criterion is a prediction of the hold-out
A rolling hold-out spends half the sample measuring what a criterion computes from all of it. Against an oracle that is arithmetic rather than an estimate, the criterion gives up 0.01701 and the hold-out 0.03200 — and the number the hold-out reports for its own winner is optimistic by more than either.
A cut is not a polynomial, and it does not have to be
A threshold's expansion never terminates, which is why a balancing dictionary's geometry was closed for powers and taken to draws for cut points. Conditioning on the second variable closes it for both.
A dependence fitted with the line
Every whitening in this collection reads the dependence off a set of residuals, and residuals are not errors. Fitting the two together recovers most of what that costs, and changes almost nothing about the decision it feeds.
A dependence with a shape
Four ways for errors to repeat, all with the same first lag and nothing else in common. A rule told the errors are a first-order autoregression finds the same number in all four, and is right about one of them.
A dictionary that is a product
Two covariates make what a balancing rule may read an outer product — eight main effects and sixteen interactions — and every inner product in it is still closed form. What a rule holding all eight main effects removes of a pure interaction is not small. It is zero.
A dictionary that is neither
A rule handed two median splits removes none of their interaction; a rule handed two covariates removes none of their product. Those were two results with two explanations, and they are one result with one — and finding it corrected the number underneath both.
A family before a fit
A regression's coefficients and one correlation can be maximised together. Replace the correlation with an estimated covariance and there is nothing left for "jointly" to mean — until a set of covariances is named, and the set turns out not to contain the truth.
A margin that turns over
A skewed covariate's leak grows without limit as the dependence strengthens. A copula's own leak does not — it peaks at a rank correlation of 0.6 and falls. The margin of the table turns over before any cell in it does.
A penalty is a trace
Akaike's 2q is not a count of coefficients. It is the answer a trace collapses to when the rows are independent — and once they are not, the trace is still the right object and is no longer the count.
A proposal that moves more than two units
The walk's autocorrelation is a fact about its step size and not about its acceptance rate. Exchanging three units from each arm mixes nearly twice as fast as exchanging one, and is refused a third more often.
A rate times a size
A sweep reported what it costs to let every candidate choose its own tuning parameter and found it flat across the list. It was reporting a product, and the two things multiplied together do not behave the same way at all.
A table of nested models
A benchmark and eight variants of it, each adding one thing. Every variant is behind before the search begins, by an amount that can be written down before the data exists — and the two most natural ways of reading the table are wrong in opposite directions.
A test rather than a survey
A thin admissible set falls into an arrangement and its mirror image, and the walk that samples it is uniform on half the reference distribution for ever. That was found by enumerating fourteen units, and enumeration stops at twenty-four.
A width promised for a difference
The exact fixed-width interval was built for one mean. Two arms make the target 42.7 units of effective size and each unit costs four observations, so the same promise about a difference costs 169.4 rather than 42.7 — and the theorem survives untouched with the harmonic size in place of the block size.
A width the trial has to stop for
The weighting that covers at 94.9% on twelve blocks covers at 91.5% when the trial stops as soon as its interval is short enough — and so does the rule that is told every block's true variance ratio. The shortfall is the stopping, not the weights.
A zero that is arithmetic
A median split's exact zero was explained by a symmetry of the latent normal. It holds under a Clayton copula, which has no such symmetry, because a centred median split squares to a quarter identically.
A zero that rests on a symmetry
A balancing rule removes exactly none of an interaction between two odd functions, at every correlation. The argument needs the joint sign flip to preserve the law, and no real covariate is symmetric about anything.
An efficiency that is a ratio
A design chosen for a model is not a design chosen for the parameter somebody wanted. Asking for one of two parameters moves the runs, unbalances the weights, and costs the other question exactly 15.07% — at every setting, because it is algebra.
An interval that carries its scale
A percentile interval inherits the resampled distribution's skewness and its scale error together. The standard repair is one extra variance per resample. It was named and not run, so this runs it.
Balanced on the wrong function
A rule that reads a covariate's numbers halves the variance of the treatment estimate, if the covariate enters the outcome as a straight line. If it enters as a threshold the rule is worth a fifth of that, and if it enters as a curve every rule here is worse than a coin.
Balancing what is known in advance
Four allocation rules, three definitions of balance, and no rule that holds more than one of them. Minimisation keeps the worst factor margin near three patients whether the trial has forty or six hundred and forty — and lets the imbalance in the cross-classified cells climb to 86% of a coin's, because the cells are not what it is watching.
Eight groups, one population
Eight hospitals are neither one hospital nor eight unrelated problems. The two obvious answers cost 2.23 and 1.15 in squared error; the estimate between them costs 0.88, and the weight it uses is not a matter of taste.
Not half and half
The same units, the same measurements, the same analysis — and a different variance, decided before anything is measured. When the two arms have different spreads the best split is σ₁ : σ₂, equal allocation costs 2(σ₁²+σ₂²)/(σ₁+σ₂)², and at three to one that is a quarter of the experiment.
Simpson's reversal is a region, not a table
The treatment wins in both groups and loses overall. That is normally shown with one famous table, which cannot answer the two questions a reader has — how often, and how large. Swept, it turns out to occupy 31% of the allocation space.
The charge nobody derived
A band of lags is charged one log-likelihood unit apiece, because that is what a regression coefficient costs. A band's numbers are not regression coefficients, and measuring what they actually cost puts the convention out by a factor of nearly three.
The comparison that was not made
Choosing a whitening's window separately for every candidate costs 0.00401 of regret. The same question about an order was named and left, because the two lists are different lengths. The order's answer is 0.00360, and matching the lists changes almost nothing.
The data that stops early
A subject still event-free when a study ends is not missing and not observed. It is known to exceed something, which is a third state most tools have no slot for — and the two obvious ways of forcing it into one are wrong by 31 and 13 percentage points.
The design for the worst case
A design for a non-linear model is optimal at a guess about the answer. Averaging over a prior repairs that on average; protecting the worst value in a range is a different problem, with a different answer, and it needs a third setting to reach it.
The design that cannot see a curve
A two-level factorial has every run at a corner, where every squared term equals one — so the column that would estimate curvature is a copy of the intercept, and the design has no information about it at all. A few runs at the centre buy one number back, and only one.
The eighth that was not a constant
How often a per-candidate tuning list changes which candidate wins is reported flat at about an eighth across list length. Vary how far apart the candidates are instead and it runs from 17.6% to 1.5%.
The experiments that could have happened
An adaptive trial's allocation is a function of the outcomes it will later be compared against, so the ordinary analysis rejects a true null 9.2% of the time. Hold the outcomes fixed, re-run the rule that assigned them, and count — the same statistic against a reference distribution the trial could actually have drawn from is back at 4.0%.
The fourth moment that was missing
Mehler's formula makes the main effects exact at any correlation and stops there, because the interactions need an expectation of four Hermite functions rather than two. A linearisation turns the four into two, and the whole geometry becomes closed again.
The gap a sample shows
The exact difference between two block windows at a block length of twenty is three tenths of a point. What a hundred and twenty rows report is four and a third, because the autocovariances the window is applied to are attenuated too.
The guess with two numbers in it
Every optimal design for a non-linear model is optimal at a guess. Where the model has one parameter that moves the settings, that guess is a number and everything about it comes out in closed form; where it has two, three constants become functions and one of them becomes zero.
The instrument and the reading
Every comparison between two block windows in this collection is an error in an implied long-run variance. Nobody reads a long-run variance. Read on the 95% point a test uses, the same bootstrap costs half as much again.
The length nobody has
Every comparison of block windows in this collection is made at each window's own best block length. That length has a standard deviation of sixteen across draws and averages twenty-five. No rule is aimed at it.
The observations that repeat each other
Almost every standard error divides by √n, which claims the observations carry independent information. At a lag-one correlation of 0.8 a fifty-point series is worth about six independent observations, and its 95% interval covers 47%.
The part the rule already took
A diagnostic that reports on what a balancing rule was not handed is run through a column that is 92% inside the span the rule balanced — because orthogonality in the population is not orthogonality on fourteen units.
The statistic the p-value is about
The test for whether a balanced-assignment walk reaches its whole set is run on a covariate function chosen before the trial. Run on the difference in arm means it is the same test, and it is about the number the trial publishes.
The variance removed before the data
Arranging forty units in pairs rather than assigning them at random cuts the variance of the estimated effect to a fifth — and the fifth is knowable in advance, because it is exactly the share of the variance the pairs do not carry.
The width a band is measured in
A tapered covariance band spends 84% of its own weights at two lags and 74% at thirty. Every charge in the collection is a straight line through the origin in those weights, so it is too dear at one end and too cheap at the other.
Three arms and three scores
Minimisation balances a trial by keeping the arms' counts even inside every prognostic factor. With two arms there is one way to measure how uneven two counts are. With three there are several, they are all called minimisation, and they send different patients to different arms.
Two effects in one number
How much two searches over one sample share is measured as the net of two things — ground both of them find, and configurations only the joint search reaches. One extra supremum per draw separates them exactly.
Two failures that cancel
A mildly skewed covariate under a lower-tail copula leaks 0.002% of an interaction where each failure alone leaks eight and seven per cent. Turn the copula over and the same pair compounds.
Two searches, one sample
A searched break in a regression manufactures 34.7 of likelihood ratio where a count of coefficients says 11.1. A searched window manufactures 84.0. The two together manufacture 99.4, not 118.7.
Two searches that share nothing
Two searches over independent columns remove shares of the residual sum that add exactly. On the scale a chi-square point is quoted on they look super-additive by a fifth of a unit, and none of it is overlap.
Weights that need only a ratio
A fixed-width interval about a difference is exact under either of two conditions and under neither in the corner. It is exact there too, and the only thing it needs is how much larger one arm's variance is than the other's.
What the 95% refers to
An interval that claims 95% is making a checkable statement about a procedure, not about the interval in front of you. Build every possible sample and count, and the interval taught first turns out to cover 87.6% of the time.
What the correction corrects
Twenty tests of true nulls produce at least one false positive 64% of the time, and the closed form and the count agree. Bonferroni holds it at 5% and Holm holds it at 5% while finding more. Nobody should still be using Bonferroni.
What the other forecast adds
Two forecasters, one series, and two different questions about them. Which is more accurate has an answer that changes with the persistence of the series; whether either is redundant has an answer that never changes at all.
What the plug-in forgets
The shrinkage weight needs a population spread, and the population spread has to be estimated from eight numbers. Empirical Bayes estimates it, substitutes it, and proceeds as though it were known — and the interval that comes out covers 79% rather than the 95% it claims.
What the rule blocks
A balancing rule breaks the admissible set into pieces by refusing exchanges. Which exchanges it refuses is computable from the design and the tolerance alone, before any assignment exists — and it makes a probe.
When the benchmark is a candidate
A specification search with a benchmark nailed down is the case with a closed form. Take the nail out — let the model that would have been reported be one of sixteen, chosen by the same data as its rivals — and the same true null is read three ways, at 2.0%, 7.8% and 76.2%.
Which forecast is better
Two forecasters, one series, and a difference in mean squared error. Whether that difference is real is a hypothesis test, its terms are not independent, and the standard error it needs is not the one a t-test computes.
Three shapes, one limit
A normalised sum has one limit and a normalised maximum has three, indexed by a single number. Twenty blocks put the sign of that number right 97.3% of the time — and naming the family from a light-tailed record gets worse as the record grows, from 83.0% at twenty blocks to 4.8% at five hundred.
The assumption nothing tests
An instrument buys a causal effect with an assumption no sample can check, and the price is set by the same quantity that made the method work. The first stage it needs is 2.7778 times the violation it is assumed not to have, so a direct effect of 0.05 demands a first stage of 0.1389 and least squares wins below it.
What a wrong model estimates
A straight line fitted to a curved truth converges on the tangent at its own design's mean. Two honest studies of one world, fitting the same wrong model, report 2.600000 and 1.600000, and neither is in error.
Coverage from exchangeability alone
A conformal interval's coverage is a fact about the ranks of m+1 numbers, so it can be enumerated before any data arrive — all 40,320 orderings of eight values, agreeing with the closed form to machine precision. What that exactness delivers is not 95%.
An identity in three terms
Reliability minus resolution plus uncertainty is quoted as a rewriting of a probability score. It is an identity to 2.6·10⁻¹⁵ on the one grouping where reliability is the whole score and resolution exactly cancels uncertainty, and it is out by 0.004125 on the coarsest grouping anybody would actually draw.
A score that balances
Weighting each unit by one over its own assignment probability drives the standardised difference between the arms from 0.8310 to 2.8×10⁻¹⁷ — exactly, not nearly. A score fitted without the second covariate leaves that covariate at 0.7057, further apart than doing nothing at all.
Three mechanisms and one dataset
Four rules for which outcomes go missing, each calibrated to lose the same 35% of the rows and each leaning on what it reads with the same coefficient. Three leave the fitted slope exactly where it was, and the one that reads the outcome moves it by 0.163531.
A p-value that is not flat is not a p-value
Under a true null, p-values are uniform. That is stronger than saying the test rejects 5% of the time, it constrains the whole distribution rather than one point of it, and it catches implementation errors that a rejection rate sails past.
The line that one point drew
A single observation among twenty-one reverses the sign of a fitted relationship. Its leverage is known from its x value before the outcome is looked at, so this is a property of the design rather than a surprise in the data.
Two standard deviations of what
The 95.45% inside two standard deviations is a fact about a curve whose centre and width are given. Drawn from ten observations, the same band holds 91.1% on average and less than 95% on 59.9% of samples — and the average is the reading that hides it.
The reversal a coin cannot prevent
Randomisation removes Simpson's reversal in expectation, which is not the same as removing it. A correctly randomised trial of eighty units, on a population where the treatment helps in both groups, reports it losing overall on 3.40% of trials — and stratifying the randomisation takes that to zero at every size.
How many analyses there really were
Bonferroni divides by twenty because twenty analyses were run. Twenty analyses of one dataset are worth 11.37 independent ones at a correlation of 0.6 and 2.58 at 0.95, and the threshold that controls exactly the same error rate is measurable rather than assumed.
R² is a property of the design
One line, one residual spread, five studies that differ only in how far apart they placed their x values. R² runs from 0.021 to 0.849 and the estimated residual spread is 0.993 in every one of them. Nothing about the relationship changed.
Where the two tails disagree
A 95% t interval on an exponential source at 120 observations covers 94.81%, which reads as very nearly right. It misses below the mean on 4.08% of samples and above on 1.11% — one tail 63% too heavy and the other 56% too light, and the total is the statistic that hides it.
A bound written for a coin
The Berry–Esseen theorem guarantees how far a standardised sum can be from the normal, and the guarantee is true. On an exponential source it is 8.62 times the real worst error at every sample size, the worst error sits at the centre rather than in a tail, and at a hundred draws the bound is larger than the 2.5% tail it would be asked to vouch for.
A hole no sample size fills
Wilson's interval is the recommended repair for a proportion, and away from the boundary it wobbles a point or two around 95%. Near zero it has a hole: at an expected count of 0.1765 its coverage is 83.50% at ten trials, 83.79% at a hundred and 83.81% at a thousand, and it never climbs past e to the minus 0.1765, which is 83.82%. The hole is where the interval built on one success stops containing the truth, and it belongs to the count rather than to the sample size.
The spread a pilot supplies
A trial sized for 80% power from a pilot's standard deviation is sized from an estimate that is too small more often than not. With a pilot of ten, 55.9% of the trials it sizes have less than 80% power and 11.1% less than 50%, although the planned sample is right on average. Sizing from the pilot's 80% upper confidence limit instead leaves 19.8% short, at 1.65 times the sample; from its 90% limit, 10.0% short at 2.12 times.
Five times in six
A 95% interval is read as a 95% chance that a replication's estimate will land inside it. With the spread known and a replication of the same size, the chance is 83.42% — five times in six — because both estimates are uncertain. An original that landed two standard errors from the truth captures a replication 48.40% of the time; a replication a tenth the size lands inside 44.54% of the time; and among significant originals from studies with 17% power, 66.94%.
A lag the sample has less of
A sample autocovariance at lag k is an average over n − k products, not n. Count a band's width in the pairs it actually has and the curvature in its charge goes away, on a correction with nothing fitted in it.
A length for each instrument
The block length that is best for an implied variance is 18.92; the one best for the 95% point of the same resamples is 16.05. A rule is a way of guessing a target, and there are two targets.
A model and a count
The share of a unit's exchanges a tolerance box refuses can be modelled from the design or counted over the admissible set. They order the units the same way at a correlation of 0.81 and disagree about the level by 0.027.
A null with a model in it
The distribution to read the winner of a table against cannot be resampled from the data, because the data does not contain the null. It has to be generated from a model — which is the assumption the resampling was chosen to avoid.
A prior on the spread
Integrating over the population spread means putting a prior on it, which sounds like the objection rather than the repair. The prior's effect is measurable, it is invisible where the groups are clearly different, and the reflex choice for a scale parameter turns out not to have a posterior at all.
A probe chosen from the design
The design's own leverage aligns with the separating direction four times better than a random direction in the same subspace. The concentrated direction the argument invites is worse than random.
A probe nobody chose
On a set that is definitively in two pieces, seven of twenty-four outcomes report nothing at all. Every covariate probe reports it. What separates them is not accuracy — it is that one of them can be chosen and the other is what happened.
A search that is already the other
A break search shifts every coefficient after a row, so a step column is one of the directions it can move in. Paired with a dictionary of them it reads exactly one, on every draw, and that fixes the top of the scale.
A split survives what a mean does not
The two things every trial balances come apart on a skewed covariate. A median split is a function of the sign of the latent normal whatever the marginal is; a mean is not, and its exact zero is gone at a skewness of one.
A symmetry that was not enough
A heavy-tailed symmetric covariate has a skewness of zero and leaks exactly nothing under three copulas. Under the two asymmetric ones it doubles the leak, from 7.707% to 14.229%.
A taper and a critical value
Two constructions whose tapers visibly differ give the same critical value, and two that share a taper exactly do not. Adding a construction whose taper is a decision rather than an accident says which half of that is true.
A threshold in the tail
How much of a threshold's imbalance a balanced covariate removes is a correlation, and the correlation is a closed form. At the median it is exactly 2/π — the same 2/π a median split throws away — and two standard deviations out it is an eighth.
A zero that was an assumption
A rule handed every main effect of both covariates removes exactly none of a pure interaction. That is true at machine precision, it is a fact about independence, and it dies as the square of the correlation.
An ordering that depends on the rule
The tapered block beats the rectangular one at the best available block length and at one estimated from the data. At a length written into a protocol, and at the rule of thumb, the rectangle wins — at every sample size measured.
Balancing towards unequal targets
A three-arm trial allocating two to one to one is the ordinary case, and a balancing rule built from raw counts does not know it. It balances the arms towards equality inside every factor level, delivers a third to each arm, and reports that it minimised imbalance.
Bias is not the whole of it
A window that reaches zero at its ends attenuates less and uses less of each block. The block length that minimises its bias is not the one that minimises its error, and comparing two windows at one length compares one of them mis-tuned.
Blinded, and still exact
The one number the exact interval needs is a ratio of within-arm spreads, which is a contrast and contains no mean — so a rule forbidden to look at the effect may compute it, on more degrees of freedom than the interval itself has.
Eight forecasters and one benchmark
A set of forecasters is a multiplicity problem on top of a dependence problem, and the two do not separate. Eight windows of one series carry the multiplicity of two and a half independent comparisons; eight separate problems carry eight.
More data is not monotonically better
Coverage of an interval for a proportion does not improve smoothly as the sample grows. It oscillates, and there are larger samples that cover materially worse than smaller ones — a sample of twenty covers twelve points worse than a sample of nineteen.
One number for a table of candidates
An effective sample size is a real quantity, it is exactly right about one thing, and that thing is a mean. Substituted into Akaike's criterion it changes nothing at all, because the penalty it is meant to fix has no sample size in it.
Pooling a proportion
A proportion cannot be shrunk on its own scale — an estimate would leave the interval, and how much information a count carries depends on where it sits. Move to log-odds and the approximation works, at the price of a group that saw nothing having no estimate at all until the correction supplies one.
Protecting one parameter over a range
A design for a non-linear model is optimal at a guess. A design for one of its parameters over a range of guesses is a worst case of a ratio of two determinants, and it is not a special case of either problem it is made of.
Randomisation is not balance
A third of all ways to split sixteen units leave the two halves more than half a standard deviation apart on a covariate. What randomisation delivers is not balance but a known reference distribution — and it makes a test exact with no assumption about the data's shape at all.
Stationary is not convergent
A walk that exchanges every unit in each arm preserves the uniform distribution exactly and never gets near it. Every doubly stochastic matrix has the same stationary distribution; only some of them have a limit.
Stopping on the arms
The width a trial will report is predictable from quantities the interval is not about. A rule that stops on the prediction covers at 94.5% where one that stops on the interval covers at 91.5, and it costs two blocks and half of the width promise.
The arcsine that closes it, and the error that was overstated
Two median splits of a correlated pair agree with probability ½ + arcsin(ρ)/π, exactly. And the truncation the field was avoiding falls geometrically in the correlation, not algebraically in the order.
The charge that is not a sum
Charging two searches what each costs on its own is conservative, and conservative here means the test never fires. At the largest break measured it declares nothing, on every draw, while a calibrated threshold reaches 29%.
The cost of a unit
Change the constraint from units to money and the allocation rule changes with it — from σᵢ to σᵢ/√cᵢ, which can point the other way. An arm that is noisy and expensive gets fewer units than the same arm would if the money were not the thing running out.
The displacement is a parameter count
A nested variant is behind its benchmark out of sample before anything is searched for. The closed form for how far turns out to have nothing about nesting in it — only two integers and a window length — and it prices a table where no candidate contains any other.
The fit that takes the memory out
A candidate's residuals report less dependence than its errors do, and how much less is arithmetic rather than noise. The rule used for a good reason reads the series that has lost the most.
The interval that forgets it estimated
The forecast band is derived for a model whose parameters are known, and then computed by putting estimates into it. Counted, the 95% interval covers 87.3% six steps ahead on twenty-five observations, and the point forecast inside it returns to the mean a third faster than the series does.
The plug-in and the maximum
A tapered covariance estimate sits five and a half log-likelihood units below the maximum of the likelihood it is substituted into. Four fifths of that is what the optimiser would have found if nothing were missing.
The quarrel that changes the winner
A disagreement about the tuning parameter costs 0.031 when it changes which candidate the table selects and −0.0007 when it does not. The distance between the values disagreed about has nothing to do with it.
The rule that reads the number
Stop categorising and let the rule read the covariate itself. What it should minimise is not an invented distance but the variance of the effect being estimated — and what comes back is not a better constant but a different rate.
The statistic that changes sign
A test for an unreachable half needs a quantity that tells one half from the other. Every symmetric reading of a mirror pair is identical, and a magnitude is the natural thing to reach for.
The symmetry the marginals could not show
A mean's interaction zero needs the covariate to be symmetric and the copula to be symmetric under reflection. Six marginals could only ever test one of those, and the other is broken by the commonest kind of dependence there is.
The test that needs the rule
A randomisation test assumes almost nothing about the data and one thing about the experiment. Tell it a fair coin produced an allocation that an adaptive rule produced — which is what every off-the-shelf permutation routine does — and it rejects 8.0% of true nulls where knowing the rule gives 4.0%.
The test with no table
The statistic that separates a real long-run relation from a spurious one is computed as a t and is not a t. At two hundred observations its 5% point is −3.38 where the t table says −1.65, and reading it against the table calls two unrelated random walks cointegrated 70.5% of the time.
The theorem that says when to stop
A search that maximises the volume of the information has no way of knowing it has finished, because nothing tells it what the maximum is. Kiefer and Wolfowitz's equality does — a design is D-optimal exactly when the worst prediction anywhere in the region equals the number of parameters, which is 6.000000000059 here, gated at machine precision.
The volume a whitening moves
A sieve's whitening has a determinant and this collection's criterion for it never carried one. Shared across a table the term cancels exactly, which is why nothing ever noticed; used per candidate it is worth more than a parameter and the whole comparison turns on it.
The weight that decides
B = se²/(se² + τ²) is not a compromise between two answers. It is exactly the posterior mean's weight, it agrees with a numerical integration to ten digits, and an argument that mentions no population at all arrives at almost the same estimator.
The window that has to be chosen, and the term that was dropped
An estimated covariance has a bandwidth in it, and both ends of the dial are wrong for different reasons. The rule a practitioner would reach for is two thirds worse than the best window there is.
The worst case in two directions
A design that protects a range of one parameter is robust. Protect the range of one parameter while holding the other at a guess and the design is still robust, still has a guarantee, and guarantees no more than a design that protects nothing at all.
The zero that was a crossing
A cell that leaks 0.002% where adding its two halves gives 16.6% is a field's headline. On a finer grid it passes through zero at a rank correlation of 0.38 — two hundredths from where it was measured.
Two degrees of freedom, one total
The block size is a dial, and the two things a fixed-width procedure claims move in opposite directions along it. Divide the width by the square root of the sample size and one of them turns out to depend on the number of blocks and on nothing else.
Two different promises
Bonferroni bounds the chance of any false positive. Benjamini–Hochberg bounds the share of the findings that are false. Both are called correcting for multiple comparisons, and one of them lets the familywise rate reach 20%.
Two factors pointing opposite ways
As the candidates on a table are pulled apart, they quarrel about the tuning parameter three times as often and the quarrel decides the winner thirty times less often. A sweep that reads the first factor has read the one pointing the wrong way.
Two routes to every number
A site about probability that only simulates has one route to each answer and no way to tell a right one from a plausible one. Every important number here is computed twice, by arithmetic that shares nothing, and the two are required to agree.
Two walks and a finding
Regress one random walk on another, independently generated, and the slope is significant 76.7% of the time with a median R² of 0.17. Nothing connects the two series, nothing in the output says so, and more data makes it worse.
What a credible interval covers
A credible interval makes the statement everyone wants and does not claim to have a coverage. It has one anyway, it can be summed over the sample space exactly, and on a reasonable prior it beats the interval taught first.
What a p-value does not say
The same p of 0.04 corresponds to a large effect in ten observations and a negligible one in two thousand. A p-value alone cannot be interpreted, and the number that makes it interpretable is almost never printed beside it.
What a search costs in parameters
An information criterion's penalty is an estimate of the optimism a fit carries. For a break point the optimism can be measured and cannot be counted, and it comes to about two and a half parameters.
What a window leaves free
A Bartlett window's weights sum to exactly half its width, which is a candidate for what the band costs. Varying the weights without varying anything else says the weights are the mechanism; varying the shape at the same weight says they are not the arithmetic.
What a zero is made of
Two disjoint dictionaries of independent columns read an excess of 0.000116 and are made of an overlap of 0.000583 and an interaction of 0.000467. The control the whole scale is anchored on reads zero because two effects cancel.
What studentising costs
Averaged over eight cells the studentised interval is 2.09 times as wide as the percentile one and covers 0.46 points better. At the block lengths the rules choose, the scale it divides by rests on two or three numbers.
What the extra function buys
A rule balancing the mean of each covariate has a worst case of exactly zero. Adding the median split — the other thing every trial balances — leaves it at exactly zero, and one square moves it.
When one model contains the other
The comparison a forecaster most often wants is between a model and the same model with one more term. That is exactly the comparison the standard test cannot make — and it fails by declaring the smaller model significantly better, more confidently the more data it is given.
Where the enumeration stops
A maximin over an eight-function dictionary is a walk over seventy subsets. Over twenty-four it is 735,471 at eight functions, and the exchange algorithm that replaces the walk scores 421. What licenses the second curve is four sizes where both exist and agree, which is a weaker warrant than it looks.
Where the generality runs out
A covariance that changes half way through a sample is not one a window can estimate. One number, a window and an order are worth the same as each other on it — and letting the model change once, at a point nobody can locate, is worth as much again as all three.
Where the minimum is attained
A design that protects a range is finished when its worst case is a tie. That is a checkable property rather than a description, it is why the search cannot climb a derivative, and it is the same corner the criteria field found at the end of the Φₚ family.
Where the two searches cross
The obvious dial between a criterion and a hold-out is how much of the sample to hold out, and moving it never changes the answer. The dial that does is one nobody chooses — how much each row repeats the one before it — and the two rules change places at about 0.81.
Which shapes are worth protecting
Choosing a basis by its worst case is a finite problem with an exact answer. The answer has no tie in it, which a maximin optimum is supposed to have — and the tie comes back, along with twice the guarantee, when the basis is drawn rather than chosen.
Which weights are the inverse variances
There is an exact estimator when the two arms share a variance and another when every block has the same two counts, and between them they cover every trial anybody designs on purpose. In the corner where neither holds, both cover 98.45% instead of 95%, and the only estimator at its level is the one with no theorem behind it.
The maximum converges slowly
The rate at which a normalised maximum reaches its limit law is computable rather than simulable, because the exact law of a maximum is always available. For a normal parent the distance falls like one over the logarithm of the block and is still 0.0091 at a million readings; for an exponential parent, with the same limit, it is 2.707×10⁻⁷.
Dropping the incomplete rows
Push the missingness until the rows that survive have a covariate mean of 0.543905 against a population zero and a variance of 0.5041 against one, and the fitted slope is still exactly right. Where the rule reads the outcome instead, the same sweep takes coverage to 2.42% at eight hundred rows.
Weak, and back where it started
A consistent instrumental estimate at two hundred rows and a concentration parameter of 0.32 is biased by 0.3220 ± 0.0142 against a least-squares inconsistency of 0.3594 — 89.6% of the way back to the problem it was hired to solve. Just identified, it has no mean at all, and that is measured as a rate rather than assumed.
What the split costs
Splitting a sample between fitting and calibrating looks like a trade against the guarantee, and it is not: coverage moves 0.63 points across nine splits and every reading sits on its own promise. The whole cost is 1.38% of width — and at sixty observations the width falls, rises and falls again.
The bread and the filling
The robust standard error is not a safety margin. At one setting of the error variance it is 1.2806 times the model-based one and at another it is 0.8246 times it, and the sign of a single dial decides which.
How many observations a weight leaves
Kish's effective sample size is exact — for an outcome whose mean does not move with the covariates the weights are built from, the studentised variance reads 1.0680 where the formula says one. For the population's own outcome the same reading is 6.769, rising to 52.497.
The curve that survives censoring
Kaplan–Meier recovers the true survival curve to within a fraction of a point at every censoring level from 37% to 71%, where dropping the censored subjects is off by 28 and then by 43. The estimator is a running product and the reason it works is in its denominator.
A tenth as wide, and both of them right
The interval for a mean and the interval for one future observation are both labelled 95%, and at a hundred observations one is 10.05 times the other — exactly the square root of n + 1. Read the narrow one as the wide one and it covers a new value 15.7% of the time.
What naming it in advance costs
Preregistration is argued for as free. Against an effect of two standard errors hiding in one of twenty analyses, naming the right one detects it 51.5% of the time and naming the wrong one detects it 4.7% of the time; correcting all twenty detects it 22.5% wherever it is. The two are worth the same when the chance of having named correctly is 38%.
A degrees of freedom that is not a count
The pooled two-sample test's size runs from 0.55% to 18.91% across forty units split five ways against five variance ratios, with a true null in every cell. Welch's runs from 4.63% to 5.51% — bought with a degrees of freedom that is a function of the data, not an integer, and not a count of anything.
The test is a point somebody chose
A test reported as 90% sensitive and 95% specific is not two properties of a test. It is one property read at a threshold, and the threshold that minimises harm runs from 3.05 standard deviations of the score at a prevalence of one in ten thousand to −0.12 at one in two — 45% of cases detected at one end and 99.9% at the other.
The plot is about the wrong quantity
A t interval needs the sampling distribution of the mean to be normal, not the data. A two-lump source leaves its quantile band on 100% of samples of forty and its interval covers 94.80%; a t on three degrees of freedom leaves it on 57% and covers 95.73%, the best of five sources.
What a guaranteed minimum costs
Clopper–Pearson's interval never covers less than 95%, and at thirty trials it averages 97.34% and is 10.4% wider than Wilson's. Blaker's interval keeps the same guarantee, averages 96.31% and is 4.6% wider. The difference is not waste: Clopper–Pearson guarantees each side separately, holding both below 2.5%, and Blaker guarantees only their sum — so at ten trials and a proportion of 0.15 it misses on one side 5.00% of the time.
Two intervals that overlap
Two 95% intervals that just touch are read as a difference at the edge of significance. With equal standard errors their difference has p = 0.0056, not 0.05; two intervals can overlap by 58.6% of an arm and still differ at exactly 5%; standard-error bars that just touch mark p = 0.157; and when the two estimates are correlated at 0.8, touching intervals conceal a difference of 6.2 standard errors. Read as a test, non-overlap needs 1.66 times the sample for the same power.
Estimates that are too alike
Posterior means give each group its least-error estimate, and as a set they are too alike: with each group's standard error equal to the population's spread, they spread 0.707 as widely as the truth. Beyond two population widths lie 2.28% of the true effects, 7.86% of the groups' own means, and 0.234% of the posterior means — a tenth of the truth. Rescaling the estimates to the right spread counts the tail exactly and costs 17% more squared error; summing each group's posterior chance of being beyond the line counts it without changing any estimate.
A charge that depends on the rule
The break search's charge is 34.7 on its own and 15.4 once a window has been chosen from the same sample. Most of what a break search finds under correlated errors is the correlation, and a whitening has taken it already.
A copula that halves a marginal
Three copulas break nothing on their own and put a factor of two between the same skewed covariate's leaks — 12.118% under a Frank against 23.640% under a t, at the same rank correlation.
A count that has to be estimated
At sixteen units the admissible assignments can be counted by walking all 12,870 of them. At four hundred there are about 2^393.70, and the share admitted is 0.31885 against a closed form of 0.31818 that has no trial size in it at all. The exhaustion a small trial runs into is a fact about small trials.
A defect that is about size
The admitted share of a rerandomisation barely moves with the number of units. The number of admissible neighbours grows like the square of it, and that is what decides whether the walk can go everywhere.
A line that beats two curves
A deferral asked for a curve. Fitted against the same measurements, a straight line in a variable nobody had to fit describes the plateau better than either curve does with a constant more — and for three windows out of four it does not.
A list is not a rule
How often five candidates disagree about a tuning parameter runs from nothing at two values on the list to two draws in five at thirteen. What the disagreement costs does not move at all.
A quantity that loses to a heuristic
Leverage is a heuristic about which units a balancing rule has most to say about. The constraint's active set is the thing the rule actually does. As a probe, the heuristic wins by 4.4 paired standard errors.
A table and a list
A nested ladder of candidates differing by one coefficient was predicted to turn over more often at every list length. It turns over less at every one, and its list changes the winner half as often.
A width that moves and an error that does not
Four charges give four widths a factor of four apart and four errors half a per cent apart. The derived charge wins, significantly, by a quarter of what was on offer — and none of the four is an estimate of anything.
An answer that changes
Eleven of twenty cells cancel and nine compound, at one rank correlation. Sweep the correlation and four of the twenty change sides — all four from compounding to cancelling, all four at the most skewed covariates.
Borrowing towards a line
A group shrunk towards the average of all groups is being compared with groups it has nothing in common with. Fit a group-level predictor and it is shrunk towards what the predictor says a group like it should be — which halves the spread left to borrow against and takes a quarter off the squared error.
Choosing the order
One criterion is consistent and one is not, which is the whole of what gets said about them. At two hundred observations the consistent one is right 95% of the time and the other 70%; at fifty they are both right 54% of the time and wrong in opposite directions, and consistency has not started to mean anything yet.
Choosing whether to break
Charging what the search manufactures takes a rule from splitting a stationary sample on 99% of draws to 16%. It also costs regret, because the two mistakes a rule can make are not the same size.
Correcting the persistence
Least squares estimates how much a series remembers of itself as smaller than it is, at every value it can take, by an amount with a closed form. Subtracting that amount back is one line of arithmetic, and what the line costs is variance.
Counting what is still wandering
The statistic that turns a spectrum into an integer has one name and three distributions. Its 5% point is 8.12, 18.64 or 31.74 depending only on how many series are left wandering under the null being tested — and read against the wrong one of those three, it calls unrelated random walks cointegrated most of the time.
Half a reference distribution
A walk that reaches half its admissible set reports the two-sided p-value exactly right, to the last digit, for ever. A one-sided one it puts on the wrong side of five per cent about once in thirty.
How often it matters
The disagreement rate rises by half across the list and the share of disagreements that decide anything falls by nearly the same factor. Their product — how often the tuning list changes which candidate wins — sits at an eighth and does not move.
Iterating is not maximising
Re-reading a correlation from the generalised residuals and refitting converges in seven steps. What it converges to solves the first-order condition of a sum of squares, and the likelihood has one term more than that.
Measuring a variance rather than a quantile
A resample's implied long-run variance can be computed from the sample with no resampling in it at all. A critical value cannot, and the difference is a factor of three in the draws before any of the resampling is counted.
Nothing in the fit picks the width
A wider band is always a better fit, and it is better by about one unit of log-likelihood a lag — which is the order of what a criterion charges for a parameter. Three defensible rules choose widths a factor of three apart.
One control, many arms
The control appears in every comparison, so it is worth √k treatment arms — and the same sharing makes the k tests correlated at n/(n+n₀), which is the quantity Bonferroni ignores. Both facts come out of one design decision, and it is the size of the control.
R² is not a measure of fit
Adding a predictor with no relationship to anything cannot reduce R², and in expectation raises it by 1/(n − 1). Twenty useless predictors on thirty points give an R² of 0.69 from pure noise.
Stopping when it is precise enough
An experiment that runs until its estimate is precise enough is the natural design and the one with a theorem against it. Its two-stage cousin keeps its promise exactly, for every unknown spread, and pays twice the observations for it.
The analysis after three arms
An unadjusted analysis after a two-arm balancing rule rejects 0.6% of true nulls where it claims 5%. With three arms and a deterministic rule it rejects none at all — and the repair is the same repair, which is a sentence and a column in the model.
The analysis has to know the rule
A trial balanced by minimisation and analysed by comparing the two arms' means rejects a true null 0.6% of the time where it claims 5%, and at full determinism 0.0%. That is not an error anybody complains about — it is a test that has stopped working, paid for by a balance the analysis then refused to use.
The condition that cannot be dropped
The weights may not read the block they weight. Estimate the variance ratio inside each block rather than across the trial and the coverage falls to 83% — on an interval that is at the same time seventy per cent wider.
The cut that is not a quantile
A protocol that says split the covariate at a threshold and one that says split it at the median read the same and are different rules. One has an exact guarantee under every marginal and the other has none under any.
The degrees of freedom in the sums
One arm partitions N − 1 exactly. Two arms give the rule N − 2b and the interval b − 1, which is short by one per block — and the missing ones are in the block sums, which are correlated with the differences at −0.79 and are usable anyway.
The design that stops guessing
Every repair so far protects a guess. The alternative is to run part of the experiment, estimate the parameter from it, and design the rest at the estimate — which recovers most of what a threefold wrong guess costs, and has a best moment to stop guessing that is earlier than anyone expects.
The models that were never in the running
A reference distribution for a set has to assume something about every candidate in it. Assuming that all of them are as good as the benchmark is what makes the reality check honest, and it is what sixteen hopeless candidates use to destroy it.
The ordering reverses again
One field found two of four rules changing sign between two readings of one resampling. Turn the same resamples into a studentised interval instead of a percentile one and all four change sign.
The plus one and the round number
A sampled randomisation test counts the observed allocation as one of its own reference draws, and the correction is invisible at B = 19, 39, 59 and 999 — every value anybody uses. At B = 20 the version without it is an 8.00% test where the corrected one is 3.80%, and the convention protecting everybody is a preference for round numbers minus one.
The prior the data estimates
A hierarchical model needs a population spread, and it does not ask for one. It reads τ off the distance between the group means — biased six per cent low, exactly zero on 53% of datasets where the groups are identical — and the prior stops being a belief.
The repair that was exact and made it worse
A penalty computed from the trace is exactly the optimism it estimates, and selecting with it gives up a fifth more than not correcting anything. The row count entered the criterion twice, and a penalty is the second place.
The reversal that was the instrument's
On an implied variance the rectangle wins at a protocol length and at the rule of thumb. On the 95% point a test reads, and on the coverage an interval delivers, the taper wins at all four rules.
The rule that cannot see the mean
A sequential rule stops when its own estimate of the spread is small, which is more often on the samples whose spread came out low — so the interval afterwards is short. There is a way to keep updating the estimate and stop being able to see the mean at all.
The set a dictionary leaves
A rule constrained on six functions at a loose tolerance leaves a set as thin as one constrained on three at a tight one. Both sampling methods cross over at the same thinness, and the tolerance where that happens moves by a factor of three.
The triangle that was not the multiplier's
A resampling that leaves each residual on its own row can keep only what the residuals have, times a triangle. A construction that moves every one of them has the same triangle — and the one in this collection's own table has a different taper entirely.
The weight that is a vector
Two forecasts have a best combination and one number describes it. Eight have a best combination too, and the vector describing it puts nothing at all on the forecast with the smallest mean squared error.
The window a whitening wants
Every law here is best whitened by a window several times longer than its own memory, including the one whose memory ends at the fourth lag. The three ways of choosing it from the sample all land in the same place, and it is the wrong one.
The zero that survives a cut
A rule holding both main effects removes half of a pure interaction between correlated powers and exactly none between correlated median splits. The guarantee that a correlation destroyed was never about interactions.
Three functions of one number
A rule that balances the covariate is exposed to every shape the outcome might have. A rule that balances three functions of it costs two points of variance against the shape the first was built for and takes the worst case from a coin's to about half of it.
Three levels, and the ring where the design says the same thing
A central composite design puts its axial runs at ±α, and α is not a matter of taste. At F to the quarter the prediction variance depends only on how far a point is from the centre and not at all on which direction it lies in — a property with no simulation in it, exact or absent.
Three quarters of the way to one search
The pair that started this reads 0.762 on a scale whose one is containment. And the pair that shares nothing but its response reads −0.306, so the sign the earlier field found does not transport at all.
Twenty intervals and one expected miss
The 95% belongs to the procedure, not to the interval in front of you. Twenty intervals from twenty samples make that visible in a way no definition does, and the one that misses is not a mistake.
Two defects and one resampling
Four resamplings, each the repair for one defect and wrong about the other. Put both defects in the same world and the statistic's 5% point is 3.8028, where the best of the four reaches 2.8326 — until a multiplier that stays on its own row and shares a sign with its neighbours reaches 2.9988.
Walking the admissible set
A rerandomisation test hunts for admissible assignments and throws away the rest. A walk visits them instead — and it is exactly uniform only because it stands still when a proposal fails, which is the step that looks like waste.
What a chosen probe finds
On a chain of eight hundred draws the probe the earlier fields use misses 44% of the sets that are split. Its own residual off the rule's span misses 12%, for one least-squares fit.
What a schedule actually buys
Big blocks early and small blocks late is the right instinct and it does not take both ends of the trade, because there are not two ends to take. What it does take is the overshoot — about four per cent of the observations — and a steadier stopping point.
What choosing the length costs
The gap between two block windows at the best available length is 2.12 points. What the best rule a practitioner could run gives up against that same length is 7.26. The argument is a third of the size of the thing it is inside.
What differencing costs
Differencing takes the false-positive rate between two unrelated walks from 76.7% to 4.9%, and takes a genuine relationship's R² from 0.91 to 0.33. Applied to a series that did not need it, it doubles the variance and installs a correlation of −0.5 that the data never had.
What normal actually looks like
A single quantile plot of forty normal points wanders enough to look suspicious. Twenty of them, all genuinely normal, show what the noise looks like — and any single panel a reader would have rejected is in there.
What the balanced trial is worth
A rule that reads the covariate removes three quarters of the imbalance. An analysis that does not know it happened prices the imbalance anyway, rejects one true null in two hundred instead of one in twenty, and finds a real effect less often than a coin-tossed trial does.
When the spread estimates to zero
The usual estimate of a population spread is a difference of two positive quantities, clamped at zero. On a third of eight-group datasets with a real spread in them the difference comes out negative, the estimate is exactly zero, and every group is pooled completely on data that said no such thing.
Where the gain is, and where the decision is
A bigger proposal is worth a factor of six at a loose tolerance and nothing at a tight one. The tolerances where it helps are the ones where a hunt costs two evaluations a draw, and the crossing barely moves.
Where the guarantee is exactly zero
An experimenter who declines to name the shapes, and asks instead to be protected against anything in a class, is asking for a number that is not small but zero. Bounding the class is unavoidable, and the two ways of doing it choose different bases.
Which tail the cut sits in
The same copula and its reflection have the same rank correlation, the same Kendall tau and the same marginals. A balancing rule holding a threshold at a dose leaves 5.33% under one and 33.36% under the other.
Adjusting for everything
"Control for every covariate that was measured" leaves a larger bias than controlling for nothing on 65.5% of four thousand randomly drawn structures and a smaller one on 33.8%. Its squared error is 4.110 times that of using no covariate at all, and half of it sits in its worst tenth of structures.
The threshold is a dial
A peaks-over-threshold analysis has one knob, and raising it buys accuracy with exceedances. For a normal parent the error is smallest at the 0.925 quantile and 80.6% of it is still bias there — and both diagnostics practitioners use to set the knob lose to a fixed 0.90 rule, one by a factor of 1.590 and one by 11.881.
One imputation is not an observation
Three ways of filling a missing outcome, under a mechanism that makes dropping the rows beyond reproach. Filling with the observed mean covers 13.85%, filling with a fitted value covers 80.85%, adding noise covers 85.78%, and the thing all three were meant to improve on covers 95.93%.
Robust is not free
A robust standard error's promise is asymptotic and its use is not. Its 95% interval covers 88.73% at twenty rows, and under mild heteroskedasticity it is the worse of the two intervals until a hundred.
Calibrated and useless
Six forecasters that are calibrated to 2·10⁻³³ run from resolution exactly 0 to 0.092758, and three forecasters with reliabilities from 0 to 0.013025 have areas under the ROC curve identical to every bit a double carries. Each measure is exactly blind to what the other one sees.
The region with no comparison
A trimmed interval covers the average effect over everybody 90.8% of the time at six hundred rows and 41.0% at nine thousand six hundred, while covering the average effect over the units it kept 94.3% and 96.0% throughout. An interval that gets worse as the sample grows is an interval about something else.
The interval at the end of the curve
The interval most software prints around a survival curve covers 89.7% at five years, where 3.3 of forty subjects are still being watched and where the curve is actually read. The same variance carried on a log–log scale covers 94.8% there — and the failure was never the width.
A coverage table with its own error
Twenty cells estimating the coverage of an interval that is exactly 95%, at a thousand replications each, read from 93.9% to 96.5% — and a table like that flags at least one of its correct cells on 69.9% of honest runs. Ten times the replications does not repair it: at ten thousand the same table still flags one 63.3% of the time.
What the first stage does not know
A single weak instrument does not make the conventional interval undercover — it makes it cover 99.1% at a width of 7.320. Where the promise actually breaks is many instruments — coverage falls from 97.2% to 51.5% while the median width falls from 1.454 to 0.583.
The effect a stopped trial reports
An O'Brien–Fleming trial at 88.45% power holds its error rate exactly and reports an effect 9.6% too large on average. The 11.39% of trials that stop at the second look report 1.83 times the truth, the ones that cross at the last look report 0.80 times it, and pooling every trial by its size gives the truth back to the last digit.
The trials that stopped early
A fixed-width trial that stops when its own interval is short enough covers 91.45% — an average of 78.2% among the 22.3% of runs that stop within eight blocks and 96% to 99% among those that run longer. Widening every interval by 17.1% brings the average to 95% and leaves the early stops at 85.6%, while 92.8% of runs now report an interval wider than the width they promised. Even doubling every interval leaves the early stops short.
The summary that was meant to work
Distance correlation is zero if and only if two variables are independent, which is exactly the guarantee a correlation coefficient lacks. Run on the four datasets that share a correlation, it spreads them by 0.10 — and Spearman, which guarantees nothing, spreads them by 0.49.
The prevalence the test has to estimate
Every predictive value takes a prevalence as given, and the prevalence is usually estimated from the same test's positive rate. At a true prevalence of one in a thousand that rate reads 5.09% — fifty times the truth — and the correction that inverts it is unbiased, 18% more variable, and negative on 48.6% of samples of a thousand.
Ninety-three observations, and nothing assumed
The interval between the smallest and largest of a sample holds a share of the population whose distribution does not depend on the population — Beta(n − 1, 2), for anything continuous. Buying the 95/95 that normality buys at ten observations costs 93 of them, and that number is the exchange rate between an assumption and data.
The skewness of a difference
Welch's test holds its size to within half a point when both groups are normal. Give both groups the same skewed population and it still balances at twenty and twenty — and at eight and thirty-two it rejects low on 7.16% of samples and high on 0.66%. One number decides which: the skewness of the difference of the two means, which ranks twenty-five cells by their imbalance with a correlation of 0.997.
The coin that makes it exact
Every interval for a proportion either covers less than 95% somewhere or more than 95% on average, because a count is discrete. One construction covers exactly 95% at every proportion: it adds a uniform random draw to the count. At thirty trials it is 0.9% wider than Wilson's interval and narrower than both exact ones — and two analysts with the same data report different intervals, and one study in forty that sees nothing reports an empty one.
An outcome cut in two
Replacing a measured outcome with whether it crossed a threshold keeps 63.7% of the information when the cut is at the mean, 34.2% at the top tenth and 13.1% two standard deviations out. A trial that needs 63 patients per arm on the measured outcome needs 102 cut at the mean and 185 cut at one and a half standard deviations. The responder rates that result read as a share of patients who respond — 50.0% against 69.1% — when every patient moved by the same amount; and a cut chosen after looking turns a 5% test into a 17.7% one.
A distribution drawn from the null
Between nested models the ordinary comparison statistic has a null distribution centred at minus one and a 95% point of a quarter. A correction to its mean repairs the centre and leaves the shape; simulating the null repairs both.
A ratio that changes between blocks
A wrong weight costs width and a random weight costs level. The rule aimed at the quantity that actually varies is the only one that misses its own coverage, and the rule that models it across blocks recovers the whole of what knowing it is worth.
A schedule that reads the mean
The block sizes may be anything at all provided they are functions of the contrasts. Two natural schedules break that, in opposite directions — and the most natural mistake of the three is not a schedule at all but a stopping rule, at 86.87% coverage and fewer observations.
A second break on a flat profile
Searching a hundred and twenty rows for one change point manufactures five units of likelihood. Searching for a second manufactures four more, on a series that has at most one — and on a profile whose whole range is under seven.
A window for every candidate
The window and the order a whitening needs are chosen once, from the fullest candidate, on an argument that was made about an estimated covariance. A tuning parameter is not a covariance, and the two cost different amounts.
Allocating on a guess
Every allocation rule in this field is a function of quantities the experiment is being run to find out. Fed a pilot's estimate of them, the rule that minimises the variance makes the experiment worse than not bothering — until the arms differ by about a factor of two, which is further than anyone would guess.
Balancing a skewed covariate
The worst case of the rule every trial runs goes from exactly zero to somewhere between a quarter of a per cent and two and a half. Which is small, and is a number that cannot be stated without the covariate's distribution in it.
Balancing more than one number
The criterion generalises to several covariates without a word changing, which makes the question what it is worth rather than whether it can be done. Each one added takes a share of the assignment's freedom, and the imbalance left in every one of them rises.
Before the trial and after
The same diagnostic run at two moments answers two different questions. Before, a positive verdict changes the design. After, it changes which number gets reported — and only for the numbers the defect can reach.
Counting it exactly does not help
If a modelled active set lost because the model was crude, the exact one would win. It is computed at a cost no trial can pay, and it is worse — so the approximation was never what was costing the probe.
Draws that repeat each other
A hunt costs 1/p evaluations per independent draw. A walk costs one per step and yields an effective draw every τ steps. Both are counted in the same unit, and the walk is dearer at every tolerance a trial is designed at.
Errors generated from a fitted model
The one construction that is not bounded by the residuals, because a model extrapolates past the lags it was told about and a truncated sample sequence cannot. It is nearly exact where the only defect is dependence, and it pays for it where there are two.
Guessing one arm in three
A balancing rule is guessable because it is balancing. With three arms the next assignment is worked out less often than with two — and by more, relative to what a guesser gets for nothing, and the damage they can do is almost unchanged.
How long a block a multiplier shares
Sharing a sign over more rows keeps more of the dependence and leaves fewer independent signs to build a distribution from. The bias falls from 1.6885 to 0.8479 and the spread rises from 1.3073 to 2.1716, and the rejection rate walks straight through its nominal level on the way from 11.3% to 1.3%.
The analysis and the shape
An unadjusted analysis after a rule that read the covariate is too cautious — by a third against a linear outcome, by nothing at all against a quadratic. And an adjustment for the wrong function recovers almost none of the precision the right one would.
The check before the standard error
One number decides whether every interval in an analysis is trustworthy, and the check for it flags a lag-one correlation of 0.5 nine times in ten — and one of 0.2 only one time in five, where the interval already covers 88.6% instead of 95%.
The corner the test is calibrated at
"No candidate is better than the benchmark" is not a null but a face of a region, and a reality check is calibrated at one corner of it. Fill the table with candidates that are hopeless rather than equal and the test finds a genuine improvement 0.0% of the time.
The cost of differencing a pair
Differencing two cointegrated series makes every standard error honest and throws away the one thing known about where they are going. The error-correction model forecasts better by exactly what a closed form says — and at four hundred observations it is better on four series in five and worse on average.
The design that has to be integers
The optimal design is a set of real weights and an experiment is a set of runs, so the theory's answer is never available. Thirteen runs reach 99.77% of it and fourteen reach 99.44% — adding a run makes the design worse per run, and the search that finds it does not always find the same one.
The design that hedges
A locally optimal design is right at one value of the unknown and 23.9% efficient at the edge of a sixteenfold range. Averaging the criterion over a prior instead buys the worst case back to 56.3% — and buys it by adding support points, at spreads the arithmetic decides rather than the experimenter — a third setting at a factor of 3.36 and a fourth at 8.86.
The diagnostic at two hundred
Pointed at a trial size no enumeration reaches, the test gives three answers rather than one — and past a certain thinness it stops agreeing with itself, which is the honest reading and the one nothing could give before.
The error no window repairs
Every block window's best estimate of a long-run variance is wrong by about forty per cent at a hundred and twenty rows, and the largest part of that is not a bias at all. Choosing the window moves a twentieth of it.
The estimate after the choice
An arm chosen for being ahead is ahead by more than it should be, and the trial then publishes the average of the stage that chose it and the stage that did not. The unbiased estimate is the one built from a third of the data — and it is the least accurate of the three.
The interval after a stop it chose
A rule that stops when the estimated precision is good enough stops on the samples whose estimate was small. Its interval covers 90% and claims 95%, and a fresh sample of the same random size covers 95.4%.
The interval after the choice
Estimating the coefficients of a known model costs a 95% forecast interval about two points of coverage. Choosing which coefficients to estimate, from the same forty observations, costs another four and a half — so the step nobody records in the output is the more expensive of the two.
The interval that integrates
A credible interval for one group in a hierarchy has to average over every value the population spread might take. That averaging is what makes it cover — 95.2% against the plug-in's 78.8% — and it costs 31% more width, a heavier tail, and a mixture rather than a normal.
The optimum is a ratio, and its interval is sometimes the whole line
The best setting is −b₁/2b₂: a ratio of two estimates whose denominator is a curvature the design can often barely see. The delta method reports a finite interval every time and covers 68.8% where the curvature is weak; Fieller's set covers 95% and says so by being unbounded.
The order the tail is drawn at
A fitted autoregression reproduces the sample exactly at the lags it was fitted on, so everything it says past them is extrapolation — and the order is the dial that decides how much of it there is.
The repair that moves the wrong number
Correcting the bias in a persistence parameter is one line of arithmetic that works. Feeding the corrected estimate into a forecast repairs the number everybody looks at, makes the forecast worse by squared error at moderate persistence, and improves the interval for a reason that has nothing to do with bias.
The residuals are not the errors
A fit removes the part of the errors lying in its own column space, and a persistent design's column space is itself slow — so what is left behind is smoother than what went in, at every lag, by an amount that grows with the lag.
The shortest interval is the one that misses
Four intervals for the same data, with their widths and their coverage measured together. The narrowest is the one that fails its stated level, which is exactly why it looks the most appealing.
The walk that cannot cross
A thin enough admissible set is not one set. It splits into an assignment and its mirror image, no sequence of admissible single swaps joins them, and the walk that samples it is uniform on half the reference distribution for ever.
The zero that survives both
A median split's interaction leak is under 10⁻¹⁶ at all thirty combinations of copula and marginal. It is the only guarantee in the collection that neither half of the dependence can touch.
Twenty residual plots
Judging whether a residual plot looks wrong requires knowing what a correct one looks like, and almost nobody has seen twenty of those. Here they are, from a model that is exactly right, at the sample size that matters.
Two levels at once
A third level of grouping adds no new arithmetic and produces one number — the design effect — that decides how many independent observations a clustered study is worth. It is the same quantity the time-series field computes for autocorrelated data, arrived at from a completely different picture.
What a better charge buys
Four charges derived from the same measurements pick band widths within six per cent of each other and deliver errors within two per cent of the gap any of them leaves. The scale a charge is levied on decides the width; the shape of the charge decides nothing.
What a design chosen from the data costs
Two fields on this site measured what happens when a rule reads the data, and the error rate broke both times. A design that reads the data to decide where to put its runs breaks nothing — and the control that proves it also finds what the real shortfall is.
What a reference distribution costs to sample
A randomisation test on a trial too large to enumerate has to sample its reference distribution, at 1/p attempts per draw and a p-value resolved to 1/(B + 1). Six constraints cost 9,878 attempts per thousand draws, and a thousand draws resolve p to 9.99·10⁻⁴ and not one digit finer.
What a two-arm rule may not pool
A spread computed "within the block" without the arm label carries a share of the effect, so the trial runs 173 observations at a null and 282 at an effect of 1.5. The stopping rule is reading the thing it exists to measure, and the phrase that produced it is one word long.
What fitting them together buys
Maximising over the coefficients and the covariance together beats the two-step under one of four dependences and ties under the other three. It is the one the band family contains, and the likelihood said so before any coefficient was compared.
What the blindfold costs
The exactly-covering rule pays for it in the width of the interval, and the block size is a dial between two costs that run in opposite directions. And on an interval whose width was fixed in advance, the same repair buys nothing at all.
What the interval covers
Eight rules and windows, and not one of them reaches its promised 95%. The range is 80.8% to 91.0%, and the choice between two block windows is a choice inside a shortfall that is four times larger.
When borrowing goes wrong
Partial pooling wins on the total and can lose badly on one group. Placed six population widths out, the group that was never from the population is estimated six times worse than by its own mean — and nothing in the output says so.
When every null is true
A reality check assumes that every candidate in the set is exactly as good as the benchmark, which is a configuration nobody's data is ever in. Test a combination against its own parts and that configuration is not assumed — it is what the arithmetic makes true.
When the constraints run out
Every function added to a basis is a constraint the assignment has to satisfy with the same units. At sixteen units and a stated tolerance the admissible assignments run 3,874, then 1,006, then 314, then none — and the count is exact, because the assignment space is finite.
Which series does the moving
“y adjusts towards x” and “x adjusts towards y” are different mechanisms with identical long-run relations, and a single-equation model cannot tell them apart because it only writes one equation. Writing all of them recovers a vector — and a gap that closes at 25% a step where one equation alone reports 15%.
The other dial
The table is swept along the strength of the dependence and never along the shape of the covariate. Swept along the shape at a fixed correlation, the same two copulas cross, the same way — and the near-zero cell turns out to be a minimum in both directions at once.
The count or the length
A block length and a block count are one number read two ways at one sample size. Read at three, the studentised interval's width penalty tracks the count — with an R² of 0.9911 against a closed form that has no length in it — and its coverage tracks the length.
A step that is not a ratio
Run the separation sweep on a tuning list of integers rather than a geometric ladder and the two factors still point opposite ways. The invariant does not survive: along a row of integers the probability moves by 2.163 where along the geometric ladder it moves by 1.208.
A level with no data in it
The largest of fifty block maxima is a 51-block event by its own plotting position, so a hundred-block level is read 1.96 times past the longest event the record contains — and it lands above the largest reading on 52.4% of records. The estimate stays nearly unbiased out there; what grows is its error, sixfold from ten blocks to a thousand.
The variance between imputations
Pooling several filled datasets covers 94.10% at two imputations and reaches its promise at five, where a single fill covered 85.78%. The correction everybody quotes is the smaller of the two doing the work — 1.00 ± 0.22 points against 1.55 ± 0.28.
Whose effect it is
With a perfectly valid instrument and no violation of anything, the estimate converges on 1.1000 where the population average effect is 0.5000. The gap is exactly θ(1 − p_c), the always-takers and never-takers cancel out of both halves of the ratio, and five per cent defiers move the answer to 1.2667.
Three corrections and a leverage
On an even design of twenty rows the four robust corrections read 0.8603, 0.9559, 1.0000 and 1.1647 of the truth and the choice barely matters. Add one point at x = 8 and they read 0.3191, 0.3419, 1.0000 and 5.1127.
Either model, but not neither
The augmented estimator's bias is −0.0085, −0.0083 and −0.0016 wherever one nuisance model is right, against components off by 0.8064 and 0.8190. One step past the overlap sweep it is the least biased estimator on the table at 0.0857 and the worst on it at 1.9265.
The same draws for both methods
Two intervals computed on the same simulated datasets give a difference in coverage whose variance can be 4.891 times smaller than on separate datasets — or, for a pair that covers different samples, 1.164 times larger. Which one a comparison gets is an exact sum over the counts each interval covers, and a standard error that ignores the sharing covers 100.00% for one pair and 93.07% for the other.
A score that rewards lying
An absolute-error score pays a forecaster exactly ⅛ of a point to replace a true quarter with a zero, and over two hundred records a liar beats a truthful forecaster on 200 of 200. A skill score against the forecaster's own average buys 0.012633 of reported skill for 0.002035 of real score.
False discoveries that arrive together
Correlate twenty tests and Benjamini–Hochberg still holds its false discovery rate — 1.66% at a correlation of 0.9 with ten real effects, against 2.55% when the tests are independent. What changes is how the errors come. A family of true nulls reports anything 2.34% of the time instead of 5.08%, and when it does, it reports 16.56 false findings out of twenty.
The bias that lands in the slope
The bias in a log variance estimate depends on nothing but its degrees of freedom, so it goes into the intercept — unless the degrees of freedom alternate with the design, which is exactly what a block-randomised trial makes them do.
What a multiplier cannot keep
Two reasons were named for the quarter a blocked resampling falls short, and taking either away makes the gap larger. What is left is a bound — a multiplier can only take dependence out, and the residuals' own is already below the errors'.
Where the bootstrap lies
Resampling is the most generally useful trick in the subject and it has a failure mode that is easy to state: it cannot see past the data. For a statistic that lives at the edge of the sample, coverage collapses from 95% to almost nothing.
A set of pairs, not a vector
The active set is a graph on the units, and every probe built from it so far has been its degree. Read as a graph it recovers 0.1326 of the alignment the summary lost — and draws level with leverage rather than passing it.
What the correction assumes
A correction with nothing fitted in it repairs one window of four. The reason is that its size is set by where a window puts its weight and the curvature it must repair is set by something else — and for one window at one sample size the two happen to agree.
Two intervals for one return level
Two 95% intervals read off the same fits of the same records, against a level known in closed form. The symmetric one covers 80.3% at twenty-five blocks and reaches only 89.0% at two hundred — and 99.24% of its misses are the interval sitting entirely below the truth, which is not the endpoint anybody expects to fail.
An imputation model the analysis does not contain
A model that fills the gaps without a covariate the analysis fits attenuates that covariate's coefficient by exactly the missing share, 0.4 to 0.26, and moves the one it did carry by exactly γρf, 0.6 to 0.642. The reverse case is supposed to inflate the interval, and at four strengths of the extra knowledge it does not.
Two instruments that disagree
The overidentification test keeps its size at 5.0% and reaches 86.4% power against a violation carried by one instrument. Against the same error carried by both in proportion to their first stages it rejects on 4.6% of draws — its own size — while the estimate is wrong by 0.3000, which is 94.2% of the confounding the instruments were brought in to remove.
When the order matters
Three ways of breaking exchangeability cost 4.93, 11.07 and 1.07 points of coverage, and the ordering by cost is the reverse of the ordering by how soon a test would have caught them. The departure practitioners check for is the cheapest one.
Right for the wrong reason
A robust standard error costs no coverage where the risk is absent — 95.52% against 95.06% at twenty rows. It costs a 6.89% wider interval and a variance estimate 2.572 times as variable, and the pre-test that would avoid paying recovers 15.9% of what the insurance is worth.
The p-value a replication gets
Under a true null a p-value is flat. Under a real effect its distribution is closed form and wide — a study with 80% power returns anything from 4.4×10⁻⁵ to 0.13 in eight runs of ten — and the chance that an exact replication of a p = 0.05 result is significant again is exactly one half, under both of the models people use without naming them.
Two points that hide each other
One far observation among twenty-one has a Cook's distance of 24.1. Put a second beside it and the two read 0.966 and 0.772, neither crossing 1, while together they reverse the slope and deleting both moves the fit by 53.3.
The measurement that got them enrolled
Enrol the top tenth of one screening reading and give them nothing, and they fall by 0.702 standard deviations at follow-up. Measured from a fresh reading taken after enrolment they fall by nothing. Averaging ten screening readings still leaves 0.101, and it takes twenty-one to get under 0.05.
One minus Kaplan–Meier is not a risk
With two ways for observation to end, one minus Kaplan–Meier for one cause reads 0.6318 at t = 5 where the chance of actually having had that event is 0.3670. Added across the two causes, the complements reach 1.4088 — more than the whole cohort. Nothing is estimated badly: the complement estimates, correctly, the risk in a world where the other cause does not exist.
The estimated weight is the better one
The propensity is known exactly here, so it can be weighted by — and estimating it from the same data and weighting by that gives a variance ratio of 0.4769 on paired draws. The reason is a projection: the draw's own imbalance explains 56.33% of the true-weight variance and 0.05% of the estimated-weight one.
Where the derivative is zero
The delta method reads a standard error off a tangent line, and at a flat point the tangent says the spread is zero. The interval built on it for a squared mean covers 99.991% there and 85.978% one and a half standard errors away, with nearly every miss on the same side — and the law it should have used is a χ², not a normal.
Estimating how many nulls are true
Benjamini–Hochberg at 5% delivers 2.55% when half of twenty nulls are false, because it cannot tell how many are. Storey's estimate of that share, read off the p-values above one half, spends the rest and finds 81.93% of the real effects instead of 74.70% on independent tests. Correlated at 0.9, the same procedure reports a finding in 19.29% of families in which every null is true.
The interval with no resampling in it
Replace 1.96 in a normal interval on the block-means variance with Student's t on one fewer degrees of freedom than there are whole blocks, and resample nothing. Across twenty-four cells it covers at least as often as the studentised bootstrap interval at every one, by 0.42 to 10.42 points; it is narrower wherever seven blocks or fewer are left; and at fifteen blocks of 32 it covers 95.0%, which no resampled interval on the grid reaches.
The shortest interval, and the one that does not move
Two 95% intervals come out of every posterior and they are not the same set. The shorter one is shorter by 4.86% on average and 22.41% at its best, it covers 86.72% where the other covers 95.68%, and it is not even the shortest once the parameter is written a different way.
The fewest groups that can borrow
At three groups the estimator that shrinks towards its own data's mean returns the group means untouched, on every dataset, because its constant is J − 3. At two it expands instead of shrinking. And the number of groups at which partial pooling starts to be worth doing is five, or two, or never — it depends on how far apart the groups are.
Correcting the forecast instead
The complaint against the usual repair is that a correction aimed at the persistence lands on the wrong quantity. Aiming it at the decay factor the forecast actually uses fixes exactly that — the error stops compounding with the horizon, 69.7% becomes 9.5% at twelve steps — and the forecast still gets worse.
How many places a design goes
Carathéodory's bound puts an optimal design's support between six and twenty-one settings, and every design in this field that can fit the model visits exactly nine. The count is not a choice anybody makes, it decides how many degrees of freedom are left to check the model with, and the first spare setting costs six points of efficiency to get back.
The sign the curvature has
A fitted surface reports a maximum, a minimum or a saddle, and the report is a comparison of two estimated eigenvalues against zero. At a true second eigenvalue of −0.25 the fit calls a genuine maximum a saddle on 26.4% of studies, and at +0.25 it calls a genuine saddle a maximum on 25.1%.
Two groupings that cross
Pupils belong to a school and to a neighbourhood, and neither is nested in the other. There is then no design effect: the overall mean is worth 8.5 independent observations out of 240, a row difference 11.1 and a column difference 26.6, and which grouping matters depends on the question rather than on the study.
The arm whose variance is its answer
With a binary outcome the allocation rule is a function of the proportions the trial exists to estimate. It costs at most 4.36% of variance to ignore it anywhere between a tenth and nine tenths, because √(p(1−p)) stays within a factor of two of its peak across 98% of the unit interval.
The cliff that is a slope
A regression between two independent series is called significant 4.9% of the time at no persistence, 52.4% at a lag-one correlation of 0.9, and 83.4% at a unit root. The rule the field offers asks whether the last of those holds, and at 0.9 the unit-root test correctly refuses one 87.2% of the time.
The rank is a decision
The sequential procedure's 5% bounds one of its two errors. Over-counting reads between 4.2% and 7.2% at every sample length from fifty observations to three hundred; under-counting reads 69.5% at fifty and 0.0% at three hundred, and nothing in the procedure bounds it.
The correction for not knowing the spread
The t distribution exists because the standard deviation is estimated rather than known. At eight observations, using the normal instead makes every interval 12% too short — and the coverage that follows can be measured rather than argued about.
A charge that reads the draw
Three charges built to read the sample track the best band width on their own draw at −0.012, −0.019 and −0.041, deliver more error than the fixed rule they are calibrated to, and pick a width half again as variable. The statistic moves; the answer does not.
The clustering the tail has
Every threshold method counts exceedances as though they were independent pieces of information, and in a dependent series they arrive in clusters. Ignoring that overstates a return level by the reciprocal of the extremal index — ×3.527 counted where the mean cluster holds four — and leaves a reported standard error 2.151 times too small.
The count that is not the rows
Three hundred rows in five clusters of sixty carry 6.9000 times the variance an independent-rows calculation reports, and the interval that counts rows covers 53.42%. The same five unequal sizes laid out two ways give design effects of 9.3158 and 5.4652.
A ratio whose interval has to be the whole line
The delta interval for a ratio of two means covers 95.61% when the denominator is eight standard errors from zero and 1.10% at a ten-thousandth of one, and ten times as wide it still covers only 3.48%. Linearising is not the fault. Gleser and Hwang proved that every interval that is always finite fails the same way, so an interval that keeps its promise has to be the whole line some of the time.
A robust loss and a far x
One far row drags least squares to a slope of −0.389. Huber's loss, the standard robust line, reaches only 0.171, and carried further out the same row gets its full weight back. Least trimmed squares reads 0.420 at every distance, and at the normal model keeps 7.13% of least squares' efficiency to do it.
Leaving each row out of its own first stage
Spread a fixed first-stage strength over thirty-two instruments and two-stage least squares covers 51.5%. Build each row's fitted treatment from a first stage that never saw that row and the same draws cover 98.7% — through an interval 5.99 times as wide, around an estimate that misses by more than the whole effect on 34.7% of draws. At eight times the strength the same repair covers 95.3% and costs a width factor of 1.66.
The liar with two answers
The forecaster an absolute-error score pays for says only 0 or 1, and on the ROC square it is a single point: its area is (TPR + TNR)/2 = 0.7684, against the honest forecaster's 0.8683, and it falls below the honest one on 200 of 200 counted records. No relabelling of its two answers returns what it threw away — the best recovers a Brier score short of the honest one by exactly the 0.022154 of resolution lost — and below a signal correlation of 0.7332 the same score prefers saying no every time to an honest forecast.
Two ways to combine p-values
Fisher's and Stouffer's combinations are both exactly right when every null is true, for the single reason that each p-value is flat. Under a real effect they disagree about which evidence counts: with Stouffer held at 50% power across ten studies, Fisher is the more powerful while the signal sits in six or fewer of them and the less powerful from seven.
Where the borrowing goes
Pooling cuts the total squared error across eight groups by 56%. Two of the eight take 61% of that reduction, the four best-measured groups share 11% between them, and the largest group gets 1.5% of what the smallest does. The headline is a fact about the groups nobody was asking about.
The design that refuses the corners
Box–Behnken runs three factors in fifteen runs and puts none of them at a corner, which is what makes it usable where a corner cannot be run. It predicts the corner 1.84 times worse than the seventeen-run design that goes there, and 1.31 times worse at the middle of a face, and all three numbers are matrix computations with no simulation in them.
Augmenting a design that has already run
The equivalence theorem still certifies when some runs are already spent, and one number in it changes: the bound is no longer p but (p − λ·tr(M⁻¹M_fixed))/(1 − λ). It equals p again exactly when the runs already made can still be absorbed into the design that would have been chosen — so the certificate says whether the experiment is still recoverable.
When the best setting is outside the region
On a flat surface at twice the noise the fitted optimum lands outside the experimental region on 24.9% of studies and more than three coded units out on 11.8%. The answer is a ridge — the best setting at each radius, with a closed form — and the two obvious rules for using it turn out to be within four per cent of each other.
A level with two units
A variance estimated from two units is a scaled chi-square on one degree of freedom. Its interquartile range spans a factor of thirteen, its ten-to-ninety range a factor of a hundred and seventy-one, and it comes out exactly zero on 26.7% of studies — so the design effect it decides runs from 1.00 to 7.01 against a truth of 4.69.
An interval for something else
An interval for the odds is free — put the endpoints through the odds and the coverage does not move, exactly, for any interval at all. The method everyone uses instead computes a new standard error on the new scale, and at twenty trials that costs four points of coverage, produces negative odds, and has no value at all when nothing was observed.
The repair that keeps the question
A regression between two independent trending series is significant 82.9% of the time on random walks and 100.0% on trend-stationary ones. Subtracting a fitted line leaves 74.2% and 33.5%; differencing leaves 5.0% and 5.2% and throws away the trend the study was about.
The weight that has to be estimated
A likelihood ratio sixteen times too large costs 5.5% of interval width and no coverage at all; one a thirtieth of the right size covers 67.90%. The estimate from a batch of five unlabelled covariates covers 95.10% against an exact repair's 95.30%, and the binomial says why.
A statistic that is exact twice
Dividing the difference in means by its own separate-variance standard error before permuting takes the rejection rate under a true weak null from 20.47% to 6.07%, keeps the exactness under the sharp null at 4.07%, and costs 0.8 points of power against a real effect. At an even split it changes nothing at all, in every draw.
Intervals for the findings
Benjamini–Hochberg's findings usually go out each with its ordinary 95% interval. With ten real effects of two standard errors among twenty tests, 11.59% of those intervals miss their effect, every miss on the far side, and the interval around the most prominent finding covers 72.36% of the time — 2.38% when the effects are one standard error. Intervals widened for the number of findings hold the share that miss under 5%.
The smallest of three combinations
Reporting whichever of Fisher's, Stouffer's and Tippett's combinations is smallest is a test of its own, and on ten studies of nothing it rejects 9.66% of the time — not 5%, and nowhere near the 15% the three sizes add to, because the statistics are correlated at up to 0.903. Read at 2.448% each it is exact, and then it trails the best single combination by at most 7.45 points and leads the worst by at least 10.30.
A forecaster that rounds
An honest probability issued in tenths loses 0.0033 of ROC area and 0.000708 of resolution — the variance its bands average away, and 89.5% of the 0.000792 it adds to the Brier score. Two hundred records of two thousand forecasts show that loss on 189; it takes about 3,300 forecasts to put it two standard errors from zero. And 3.207 in every thousand forecasts in tenths are a 0% on an event that happened, which a logarithmic score charges without limit.
A horizon chosen after looking
A difference in restricted mean survival read at whichever of eleven horizons looks most convincing rejects 11.24% of trials in which the treatment does nothing, against 4.70% at a horizon fixed in advance. The correlation of the differences across horizons is closed, and the Gaussian process it defines prices the choice at a critical value of 2.317 — which brings the counted size back to 4.99% and keeps 96.92% of the power that a horizon nobody could have known to fix would have had.
The start an efficient robust line inherits
The MM-estimator carries a trimmed fit on through a redescending loss, and it does what it promises on one far row: slope 0.479 at every distance, the row at weight exactly zero, and 87.2% of least squares' efficiency at twenty rows. What it cannot do is choose. At eight far rows of twenty the exact trimmed fit picks the wrong half on 111 datasets; the efficient step repairs none of them, spoils none of the other 89, and ends nearer the wrong line than the start did.
A standard error that knows about the instruments
Limited-information maximum likelihood came out least biased when a concentration parameter of 8 was spread over thirty-two instruments, and its conventional interval covered 79.0%. Bekker's many-instrument standard error covers 93.8% on the same draws, at 63% of the jackknife's width — and it gets there with a median standard error of 0.561 against a true spread of 0.797, because it is large on the draws that need it. At eight times the strength it covers 94.9% at 91% of the jackknife's width, and nothing measured here beats it.
The slope of a density nobody can see
Tweedie's formula corrects a reading by the slope of the readings' own log-density, and a study has its readings. Estimated from a thousand of them, the correction for the top one per cent beats the correlation's linear rule on 84.0% to 98.0% of studies from heavy-tailed populations and on 75.0% to 81.5% from a bounded one — and costs an error of 0.09 to 0.12 where the population is normal and the rule was already exact. At 250 readings the log-spline loses to the rule it replaces, and at 16,000 the same log-spline gets worse on a power tail.
The run length a declustering chooses
The runs estimator of an extremal index carries a constant nobody derives. Where a cluster is a run of neighbouring exceedances the constant barely matters; where a cluster's members fall six steps apart, the estimate is 0.9069 at a run length of six and 0.3649 at seven against an index of 0.40, and a run length of four removes under a tenth of the overstatement declustering exists to remove. A rule that reads the run length off the data has the smallest worst error of the three.
When the prior is confident and wrong
A prior worth thirty-five observations, centred in the wrong place, produces a 95% interval that covers nothing at all — and reports a width 5% narrower than an honest one. It takes seventeen thousand observations to repair, not thirty-five, and the worst study to run is the one whose sample size equals the prior's weight, exactly.
One population, or two
Group effects from two clusters rather than one bell, with the same total spread. The analysis recovers the same population spread, uses the same weight for every group, and reports nothing unusual — while 46% of its estimates land in a region holding 6.6% of the truths.
What the interval is short by
The forecast interval covers 88.42% where it claims 95%. Correcting the persistence recovers 2.66 points, correcting the innovation variance 0.56, propagating the persistence's own standard error 0.40 — and all three together recover 4.20 of the 6.58, leaving a residual none of the standard repairs reaches.
The run that confirms it
The setting a response-surface analysis recommends was chosen because the fitted surface was highest there, so the height the fit predicts at it is a maximum over a random field. At twice the noise the fit predicts 0.858 more than is there — 0.72 of the prediction's own standard error — and the gap is not noise, it is the selection.
How slow a return a sample can see
At two hundred observations the test finds a gap that halves in five steps four times in five, one that halves in eight 37.3% of the time, and one that halves in fifty 4.95% of the time — which is the rate at which it finds pairs with no mechanism at all. The boundary moves with the sample, not with its square root.
An interval that covers and says nothing
A procedure returning the whole line 95% of the time and the empty set otherwise has coverage exactly 95% at every parameter value. Two real intervals at forty observations have expected widths of 0.2418 and 0.2417 and worst-case coverages of 55.31% and 92.21%.
What a two-unit study should report
The fixed-effect interval covers the mean of the two sites in hand 96.37% of the time and the population mean 54.77%. The random-effects interval covers the population mean 94.96% — exactly its level, from one degree of freedom — and is 11.6 times wider.