Error rate — where it appears
Named by 47 essays across 24 fields — each of them below, with the objects they name alongside it.
Choosing n after looking
Re-estimating the sample size from an interim is the one adaptation with a defence, and the defence is exactly what it costs: an analyst kept blind to the arms measures a spread that contains the effect, so the design overshoots by 1 + Δ²/4σ². Re-estimating the effect instead breaks the error rate.
The experiments that could have happened
An adaptive trial's allocation is a function of the outcomes it will later be compared against, so the ordinary analysis rejects a true null 9.2% of the time. Hold the outcomes fixed, re-run the rule that assigned them, and count — the same statistic against a reference distribution the trial could actually have drawn from is back at 4.0%.
What the other forecast adds
Two forecasters, one series, and two different questions about them. Which is more accurate has an answer that changes with the persistence of the series; whether either is redundant has an answer that never changes at all.
When the benchmark is a candidate
A specification search with a benchmark nailed down is the case with a closed form. Take the nail out — let the model that would have been reported be one of sixteen, chosen by the same data as its rivals — and the same true null is read three ways, at 2.0%, 7.8% and 76.2%.
When the looking happens
A p-value is defined relative to a sampling plan, so the same data means different things under different stopping rules. Testing five times at the nominal level rejects a true null 14% of the time, and no observation in the dataset changed.
Which forecast is better
Two forecasters, one series, and a difference in mean squared error. Whether that difference is real is a hypothesis test, its terms are not independent, and the standard error it needs is not the one a t-test computes.
A taper and a critical value
Two constructions whose tapers visibly differ give the same critical value, and two that share a taper exactly do not. Adding a construction whose taper is a decision rather than an accident says which half of that is true.
Eight forecasters and one benchmark
A set of forecasters is a multiplicity problem on top of a dependence problem, and the two do not separate. Eight windows of one series carry the multiplicity of two and a half independent comparisons; eight separate problems carry eight.
Randomisation is not balance
A third of all ways to split sixteen units leave the two halves more than half a standard deviation apart on a covariate. What randomisation delivers is not balance but a known reference distribution — and it makes a test exact with no assumption about the data's shape at all.
Randomising towards the winner
Allocating more patients to the arm that is doing better is the humane thing to want and it buys nothing statistically: at a fixed total it costs thirty points of power. And because the allocation is a function of the outcomes, the ordinary test on it rejects a true null 7.8% of the time before any time trend is applied — and 58% after one.
Spending the error rate
The repair for interim testing is to spend 5% across the looks rather than at each one. The boundaries are solvable rather than quotable, and a trial that can stop early uses 298 observations where a fixed design uses 400 — at a cost of half a point of power.
The test that needs the rule
A randomisation test assumes almost nothing about the data and one thing about the experiment. Tell it a fair coin produced an allocation that an adaptive rule produced — which is what every off-the-shelf permutation routine does — and it rejects 8.0% of true nulls where knowing the rule gives 4.0%.
The test with no table
The statistic that separates a real long-run relation from a spurious one is computed as a t and is not a t. At two hundred observations its 5% point is −3.38 where the t table says −1.65, and reading it against the table calls two unrelated random walks cointegrated 70.5% of the time.
Two different promises
Bonferroni bounds the chance of any false positive. Benjamini–Hochberg bounds the share of the findings that are false. Both are called correcting for multiple comparisons, and one of them lets the familywise rate reach 20%.
When one model contains the other
The comparison a forecaster most often wants is between a model and the same model with one more term. That is exactly the comparison the standard test cannot make — and it fails by declaring the smaller model significantly better, more confidently the more data it is given.
A degrees of freedom that is not a count
The pooled two-sample test's size runs from 0.55% to 18.91% across forty units split five ways against five variance ratios, with a true null in every cell. Welch's runs from 4.63% to 5.51% — bought with a degrees of freedom that is a function of the data, not an integer, and not a count of anything.
Choosing whether to break
Charging what the search manufactures takes a rule from splitting a stationary sample on 99% of draws to 16%. It also costs regret, because the two mistakes a rule can make are not the same size.
Counting what is still wandering
The statistic that turns a spectrum into an integer has one name and three distributions. Its 5% point is 8.12, 18.64 or 31.74 depending only on how many series are left wandering under the null being tested — and read against the wrong one of those three, it calls unrelated random walks cointegrated most of the time.
Dropping the losers
Carrying the best of eight arms forward and testing it at 1.96 rejects a true null 10.3% of the time — the hypothesis was chosen by looking at the data, so the statistic is a maximum wearing a single comparison's clothes. The value that holds the rate is 2.313, and it has to be solved for.
Residuals that keep their own variance
A reference distribution for a search has to be generated from a fitted model, and the generator draws residuals. Four ways of drawing them keep four different things — and the one this site has reached for three times repairs nothing at all here.
The analysis after three arms
An unadjusted analysis after a two-arm balancing rule rejects 0.6% of true nulls where it claims 5%. With three arms and a deterministic rule it rejects none at all — and the repair is the same repair, which is a sentence and a column in the model.
The analysis has to know the rule
A trial balanced by minimisation and analysed by comparing the two arms' means rejects a true null 0.6% of the time where it claims 5%, and at full determinism 0.0%. That is not an error anybody complains about — it is a test that has stopped working, paid for by a balance the analysis then refused to use.
The models that were never in the running
A reference distribution for a set has to assume something about every candidate in it. Assuming that all of them are as good as the benchmark is what makes the reality check honest, and it is what sixteen hopeless candidates use to destroy it.
The plus one and the round number
A sampled randomisation test counts the observed allocation as one of its own reference draws, and the correction is invisible at B = 19, 39, 59 and 999 — every value anybody uses. At B = 20 the version without it is an 8.00% test where the corrected one is 3.80%, and the convention protecting everybody is a preference for round numbers minus one.
The triangle that was not the multiplier's
A resampling that leaves each residual on its own row can keep only what the residuals have, times a triangle. A construction that moves every one of them has the same triangle — and the one in this collection's own table has a different taper entirely.
Two defects and one resampling
Four resamplings, each the repair for one defect and wrong about the other. Put both defects in the same world and the statistic's 5% point is 3.8028, where the best of the four reaches 2.8326 — until a multiplier that stays on its own row and shares a sign with its neighbours reaches 2.9988.
What the balanced trial is worth
A rule that reads the covariate removes three quarters of the imbalance. An analysis that does not know it happened prices the imbalance anyway, rejects one true null in two hundred instead of one in twenty, and finds a real effect less often than a coin-tossed trial does.
The skewness of a difference
Welch's test holds its size to within half a point when both groups are normal. Give both groups the same skewed population and it still balances at twenty and twenty — and at eight and thirty-two it rejects low on 7.16% of samples and high on 0.66%. One number decides which: the skewness of the difference of the two means, which ranks twenty-five cells by their imbalance with a correlation of 0.997.
Errors generated from a fitted model
The one construction that is not bounded by the residuals, because a model extrapolates past the lags it was told about and a truncated sample sequence cannot. It is nearly exact where the only defect is dependence, and it pays for it where there are two.
Guessing one arm in three
A balancing rule is guessable because it is balancing. With three arms the next assignment is worked out less often than with two — and by more, relative to what a guesser gets for nothing, and the damage they can do is almost unchanged.
How long a block a multiplier shares
Sharing a sign over more rows keeps more of the dependence and leaves fewer independent signs to build a distribution from. The bias falls from 1.6885 to 0.8479 and the spread rises from 1.3073 to 2.1716, and the rejection rate walks straight through its nominal level on the way from 11.3% to 1.3%.
The analysis and the shape
An unadjusted analysis after a rule that read the covariate is too cautious — by a third against a linear outcome, by nothing at all against a quadratic. And an adjustment for the wrong function recovers almost none of the precision the right one would.
The corner the test is calibrated at
"No candidate is better than the benchmark" is not a null but a face of a region, and a reality check is calibrated at one corner of it. Fill the table with candidates that are hopeless rather than equal and the test finds a genuine improvement 0.0% of the time.
The estimate after the choice
An arm chosen for being ahead is ahead by more than it should be, and the trial then publishes the average of the stage that chose it and the stage that did not. The unbiased estimate is the one built from a third of the data — and it is the least accurate of the three.
Twenty residual plots
Judging whether a residual plot looks wrong requires knowing what a correct one looks like, and almost nobody has seen twenty of those. Here they are, from a model that is exactly right, at the sample size that matters.
What a reference distribution costs to sample
A randomisation test on a trial too large to enumerate has to sample its reference distribution, at 1/p attempts per draw and a p-value resolved to 1/(B + 1). Six constraints cost 9,878 attempts per thousand draws, and a thousand draws resolve p to 9.99·10⁻⁴ and not one digit finer.
What the exactness buys
Against a z test calibrated to reject exactly 5% of true nulls on this design, the randomisation test loses nineteen points of power. What it buys is that the calibration needs the success rate — which moves the critical value from 1.668 to 2.718 and is the quantity the trial was run to find out.
When every null is true
A reality check assumes that every candidate in the set is exactly as good as the benchmark, which is a configuration nobody's data is ever in. Test a combination against its own parts and that configuration is not assumed — it is what the arithmetic makes true.
False discoveries that arrive together
Correlate twenty tests and Benjamini–Hochberg still holds its false discovery rate — 1.66% at a correlation of 0.9 with ten real effects, against 2.55% when the tests are independent. What changes is how the errors come. A family of true nulls reports anything 2.34% of the time instead of 5.08%, and when it does, it reports 16.56 false findings out of twenty.
The miscalibration a perfect forecaster shows
A forecaster whose true reliability is exactly zero shows a calibration error of 0.1252 on fifty forecasts and 0.0090 on ten thousand. Every one of 1,200 blameless hundred-forecast records exceeds the 0.02 routinely read as evidence of a problem, and the mean does not fall under it until 1,976 forecasts.
A simulation that stops when it looks settled
A simulation of an interval that covers exactly 95%, checked every 250 replications for a significant departure and stopped when it finds one, flags that correct interval on 29.54% of runs. Stopped instead as soon as its estimate reaches 95%, it reports an interval that covers 94% as meeting its level on 37.21% of runs. Stopped when the estimate stops moving, it reports the right number — and has quietly chosen to run about fifteen hundred replications.
A boundary for giving up
Adding "stop if z is below zero" to an O'Brien–Fleming trial costs 5.20 points of power at the effect it was designed for and halves the observations a trial with no effect uses. Stopping when conditional power at the observed trend falls under 10% costs 13.23 points and stops 21.28% of trials with a real effect. Making that rule binding lowers the benefit boundary from 2.040 to 1.901, and a binding rule that is then ignored rejects a true null 3.523% of the time instead of 2.5%.
Estimating how many nulls are true
Benjamini–Hochberg at 5% delivers 2.55% when half of twenty nulls are false, because it cannot tell how many are. Storey's estimate of that share, read off the p-values above one half, spends the rest and finds 81.93% of the real effects instead of 74.70% on independent tests. Correlated at 0.9, the same procedure reports a finding in 19.29% of families in which every null is true.
The null the exactness is for
A permutation test is exact under the hypothesis that the treatment changed nothing for anybody. Under the hypothesis it changed nothing on average, with a quarter of the units treated and the effect varying between them, it rejects a true null 22.93% of the time.
The rank is a decision
The sequential procedure's 5% bounds one of its two errors. Over-counting reads between 4.2% and 7.2% at every sample length from fifty observations to three hundred; under-counting reads 69.5% at fifty and 0.0% at three hundred, and nothing in the procedure bounds it.
A look the trend asked for
Under an O'Brien–Fleming-type spending function, every schedule of looks fixed in advance spends exactly 5.0000%. A committee that adds a look at three quarters of the trial whenever the interim z is 1.5 or more spends 5.2323% — 5.315% counted over a hundred thousand trials — and the most a committee choosing among six schedules could spend is 5.4390%.
An order that spends the error rate
Test twenty hypotheses in a declared order, each at the full 5% and each only if every one before it was rejected, and the first is found 85.3% of the time where Holm finds it 52.5%. The tenth is found 20.4% of the time, the product of the powers before it. Move one true null to the head of the list and every real effect behind it is found no more than 4.3% of the time.
Named alongside it
The objects these essays reach for when they reach for this one.
Monte CarloReference distributionStatistical powerNull hypothesisCritical valuep-valueRandomisation testAdaptive designBlock bootstrapMultiple comparisonsBenchmark forecastExperimental design