The thread: Which rate is being controlled
The comparison that was not made
Choosing a whitening's window separately for every candidate costs 0.00401 of regret. The same question about an order was named and left, because the two lists are different lengths. The order's answer is 0.00360, and matching the lists changes almost nothing.
The family behind the letters
A, D and E are not three ideas. They are three points of one family with a single dial, and running the dial from one end to the other doubles the smallest eigenvalue of the information matrix while closing the gap above it fifty-three-fold — which is the family driving its own last member to the place where it stops being differentiable.
What the correction corrects
Twenty tests of true nulls produce at least one false positive 64% of the time, and the closed form and the count agree. Bonferroni holds it at 5% and Holm holds it at 5% while finding more. Nobody should still be using Bonferroni.
When the benchmark is a candidate
A specification search with a benchmark nailed down is the case with a closed form. Take the nail out — let the model that would have been reported be one of sixteen, chosen by the same data as its rivals — and the same true null is read three ways, at 2.0%, 7.8% and 76.2%.
When the looking happens
A p-value is defined relative to a sampling plan, so the same data means different things under different stopping rules. Testing five times at the nominal level rejects a true null 14% of the time, and no observation in the dataset changed.
How many analyses there really were
Bonferroni divides by twenty because twenty analyses were run. Twenty analyses of one dataset are worth 11.37 independent ones at a correlation of 0.6 and 2.58 at 0.95, and the threshold that controls exactly the same error rate is measurable rather than assumed.
The band the eye was standing in for
The confidence band software draws on a quantile plot holds each point at 95%, and a genuinely normal sample of forty has forty chances to leave it — so 45.0% of them do. The band a reader is actually using is that one widened by a factor of 1.502, and nothing draws it.
A null with a model in it
The distribution to read the winner of a table against cannot be resampled from the data, because the data does not contain the null. It has to be generated from a model — which is the assumption the resampling was chosen to avoid.
An ordering that depends on the rule
The tapered block beats the rectangular one at the best available block length and at one estimated from the data. At a length written into a protocol, and at the rule of thumb, the rectangle wins — at every sample size measured.
Eight forecasters and one benchmark
A set of forecasters is a multiplicity problem on top of a dependence problem, and the two do not separate. Eight windows of one series carry the multiplicity of two and a half independent comparisons; eight separate problems carry eight.
Spending the error rate
The repair for interim testing is to spend 5% across the looks rather than at each one. The boundaries are solvable rather than quotable, and a trial that can stop early uses 298 observations where a fixed design uses 400 — at a cost of half a point of power.
The charge that is not a sum
Charging two searches what each costs on its own is conservative, and conservative here means the test never fires. At the largest break measured it declares nothing, on every draw, while a calibrated threshold reaches 29%.
The symmetry the marginals could not show
A mean's interaction zero needs the covariate to be symmetric and the copula to be symmetric under reflection. Six marginals could only ever test one of those, and the other is broken by the commonest kind of dependence there is.
The two terms anybody wanted
D-optimality estimates all six parameters of a quadratic as precisely as possible. Nobody wants that. An experimenter looking for a maximum wants the two curvature terms, and the design that gives them is not the D-optimal one — it is a quarter of the runs at the centre, exactly, and the D-optimal design is 75.3% efficient for the question that was actually asked.
The volume a whitening moves
A sieve's whitening has a determinant and this collection's criterion for it never carried one. Shared across a table the term cancels exactly, which is why nothing ever noticed; used per candidate it is worth more than a parameter and the whole comparison turns on it.
Two different promises
Bonferroni bounds the chance of any false positive. Benjamini–Hochberg bounds the share of the findings that are false. Both are called correcting for multiple comparisons, and one of them lets the familywise rate reach 20%.
What a credible interval covers
A credible interval makes the statement everyone wants and does not claim to have a coverage. It has one anyway, it can be summed over the sample space exactly, and on a reasonable prior it beats the interval taught first.
A degrees of freedom that is not a count
The pooled two-sample test's size runs from 0.55% to 18.91% across forty units split five ways against five variance ratios, with a true null in every cell. Welch's runs from 4.63% to 5.51% — bought with a degrees of freedom that is a function of the data, not an integer, and not a count of anything.
Dropping the losers
Carrying the best of eight arms forward and testing it at 1.96 rejects a true null 10.3% of the time — the hypothesis was chosen by looking at the data, so the statistic is a maximum wearing a single comparison's clothes. The value that holds the rate is 2.313, and it has to be solved for.
Four letters and two camps
D, A, G and I are four ways of turning one matrix into one number, and they do not agree. The design that wins on D is the worst thing here on I. And the same two designs swap places entirely when the region changes from a square to a disc — on all four criteria at once.
Nothing in the fit picks the width
A wider band is always a better fit, and it is better by about one unit of log-likelihood a lag — which is the order of what a criterion charges for a parameter. Three defensible rules choose widths a factor of three apart.
One control, many arms
The control appears in every comparison, so it is worth √k treatment arms — and the same sharing makes the k tests correlated at n/(n+n₀), which is the quantity Bonferroni ignores. Both facts come out of one design decision, and it is the size of the control.
The analysis has to know the rule
A trial balanced by minimisation and analysed by comparing the two arms' means rejects a true null 0.6% of the time where it claims 5%, and at full determinism 0.0%. That is not an error anybody complains about — it is a test that has stopped working, paid for by a balance the analysis then refused to use.
The models that were never in the running
A reference distribution for a set has to assume something about every candidate in it. Assuming that all of them are as good as the benchmark is what makes the reality check honest, and it is what sixteen hopeless candidates use to destroy it.
The price of control
Every correction is paid for in power, and the exchange rate can be measured. Holm buys familywise control for 33 percentage points of power; Benjamini–Hochberg buys a weaker guarantee for 10. Neither is free and neither is a matter of taste.
The rule that cannot see the mean
A sequential rule stops when its own estimate of the spread is small, which is more often on the samples whose spread came out low — so the interval afterwards is short. There is a way to keep updating the estimate and stop being able to see the mean at all.
Two defects and one resampling
Four resamplings, each the repair for one defect and wrong about the other. Put both defects in the same world and the statistic's 5% point is 3.8028, where the best of the four reaches 2.8326 — until a multiplier that stays on its own row and shares a sign with its neighbours reaches 2.9988.
What choosing the length costs
The gap between two block windows at the best available length is 2.12 points. What the best rule a practitioner could run gives up against that same length is 7.26. The argument is a third of the size of the thing it is inside.
What the balanced trial is worth
A rule that reads the covariate removes three quarters of the imbalance. An analysis that does not know it happened prices the imbalance anyway, rejects one true null in two hundred instead of one in twenty, and finds a real effect less often than a coin-tossed trial does.
Which tail the cut sits in
The same copula and its reflection have the same rank correlation, the same Kendall tau and the same marginals. A balancing rule holding a threshold at a dose leaves 5.33% under one and 33.36% under the other.
Marginal is not conditional
One exactly valid interval covers 100.00% of a quiet group and 90.66% of a noisy one, and the floor is arithmetic rather than a measurement — a group of share π is guaranteed only 1 − α/π, which is zero when the group is as rare as the miss rate.
Robust is not free
A robust standard error's promise is asymptotic and its use is not. Its 95% interval covers 88.73% at twenty rows, and under mild heteroskedasticity it is the worse of the two intervals until a hundred.
A coverage table with its own error
Twenty cells estimating the coverage of an interval that is exactly 95%, at a thousand replications each, read from 93.9% to 96.5% — and a table like that flags at least one of its correct cells on 69.9% of honest runs. Ten times the replications does not repair it: at ten thousand the same table still flags one 63.3% of the time.
What the first stage does not know
A single weak instrument does not make the conventional interval undercover — it makes it cover 99.1% at a width of 7.320. Where the promise actually breaks is many instruments — coverage falls from 97.2% to 51.5% while the median width falls from 1.454 to 0.583.
How long a block a multiplier shares
Sharing a sign over more rows keeps more of the dependence and leaves fewer independent signs to build a distribution from. The bias falls from 1.6885 to 0.8479 and the spread rises from 1.3073 to 2.1716, and the rejection rate walks straight through its nominal level on the way from 11.3% to 1.3%.
The analysis and the shape
An unadjusted analysis after a rule that read the covariate is too cautious — by a third against a linear outcome, by nothing at all against a quadratic. And an adjustment for the wrong function recovers almost none of the precision the right one would.
The corner the test is calibrated at
"No candidate is better than the benchmark" is not a null but a face of a region, and a reality check is calibrated at one corner of it. Fill the table with candidates that are hopeless rather than equal and the test finds a genuine improvement 0.0% of the time.
What a reference distribution costs to sample
A randomisation test on a trial too large to enumerate has to sample its reference distribution, at 1/p attempts per draw and a p-value resolved to 1/(B + 1). Six constraints cost 9,878 attempts per thousand draws, and a thousand draws resolve p to 9.99·10⁻⁴ and not one digit finer.
What a two-arm rule may not pool
A spread computed "within the block" without the arm label carries a share of the effect, so the trial runs 173 observations at a null and 282 at an effect of 1.5. The stopping rule is reading the thing it exists to measure, and the phrase that produced it is one word long.
What fitting them together buys
Maximising over the coefficients and the covariance together beats the two-step under one of four dependences and ties under the other three. It is the one the band family contains, and the likelihood said so before any coefficient was compared.
What the blindfold costs
The exactly-covering rule pays for it in the width of the interval, and the block size is a dial between two costs that run in opposite directions. And on an interval whose width was fixed in advance, the same repair buys nothing at all.
When every null is true
A reality check assumes that every candidate in the set is exactly as good as the benchmark, which is a configuration nobody's data is ever in. Test a combination against its own parts and that configuration is not assumed — it is what the arithmetic makes true.
The outcomes a trial could have stopped with
A trial that stops at its second look with z = 3.3 has a two-sided p-value of 0.000969, 0.000987, 0.00187 or 0.0421, depending on how the outcomes it could have stopped with are ordered. One of the four orderings does not change when the looks the trial never reached are replanned, and the same one gives a trial that ran to the end with z = 6 a p-value of 0.0256.
False discoveries that arrive together
Correlate twenty tests and Benjamini–Hochberg still holds its false discovery rate — 1.66% at a correlation of 0.9 with ten real effects, against 2.55% when the tests are independent. What changes is how the errors come. A family of true nulls reports anything 2.34% of the time instead of 5.08%, and when it does, it reports 16.56 false findings out of twenty.
Two instruments that disagree
The overidentification test keeps its size at 5.0% and reaches 86.4% power against a violation carried by one instrument. Against the same error carried by both in proportion to their first stages it rejects on 4.6% of draws — its own size — while the estimate is wrong by 0.3000, which is 94.2% of the confounding the instruments were brought in to remove.
When the order matters
Three ways of breaking exchangeability cost 4.93, 11.07 and 1.07 points of coverage, and the ordering by cost is the reverse of the ordering by how soon a test would have caught them. The departure practitioners check for is the cheapest one.
A simulation that stops when it looks settled
A simulation of an interval that covers exactly 95%, checked every 250 replications for a significant departure and stopped when it finds one, flags that correct interval on 29.54% of runs. Stopped instead as soon as its estimate reaches 95%, it reports an interval that covers 94% as meeting its level on 37.21% of runs. Stopped when the estimate stops moving, it reports the right number — and has quietly chosen to run about fifteen hundred replications.
A boundary for giving up
Adding "stop if z is below zero" to an O'Brien–Fleming trial costs 5.20 points of power at the effect it was designed for and halves the observations a trial with no effect uses. Stopping when conditional power at the observed trend falls under 10% costs 13.23 points and stops 21.28% of trials with a real effect. Making that rule binding lowers the benefit boundary from 2.040 to 1.901, and a binding rule that is then ignored rejects a true null 3.523% of the time instead of 2.5%.
Estimating how many nulls are true
Benjamini–Hochberg at 5% delivers 2.55% when half of twenty nulls are false, because it cannot tell how many are. Storey's estimate of that share, read off the p-values above one half, spends the rest and finds 81.93% of the real effects instead of 74.70% on independent tests. Correlated at 0.9, the same procedure reports a finding in 19.29% of families in which every null is true.
The null the exactness is for
A permutation test is exact under the hypothesis that the treatment changed nothing for anybody. Under the hypothesis it changed nothing on average, with a quarter of the units treated and the effect varying between them, it rejects a true null 22.93% of the time.
The rank is a decision
The sequential procedure's 5% bounds one of its two errors. Over-counting reads between 4.2% and 7.2% at every sample length from fifty observations to three hundred; under-counting reads 69.5% at fifty and 0.0% at three hundred, and nothing in the procedure bounds it.
The count that is not the rows
Three hundred rows in five clusters of sixty carry 6.9000 times the variance an independent-rows calculation reports, and the interval that counts rows covers 53.42%. The same five unequal sizes laid out two ways give design effects of 9.3158 and 5.4652.
A look the trend asked for
Under an O'Brien–Fleming-type spending function, every schedule of looks fixed in advance spends exactly 5.0000%. A committee that adds a look at three quarters of the trial whenever the interim z is 1.5 or more spends 5.2323% — 5.315% counted over a hundred thousand trials — and the most a committee choosing among six schedules could spend is 5.4390%.
An order that spends the error rate
Test twenty hypotheses in a declared order, each at the full 5% and each only if every one before it was rejected, and the first is found 85.3% of the time where Holm finds it 52.5%. The tenth is found 20.4% of the time, the product of the powers before it. Move one true null to the head of the list and every real effect behind it is found no more than 4.3% of the time.
Which mistake about the rank costs
On a system with two relations, imposing none costs 29.2% of squared forecast error and imposing three costs 2.5%. The expensive mistake is under-counting, which is the error the procedure's 5% does not bound — so the guarantee protects the cheap side.
Intervals for the findings
Benjamini–Hochberg's findings usually go out each with its ordinary 95% interval. With ten real effects of two standard errors among twenty tests, 11.59% of those intervals miss their effect, every miss on the far side, and the interval around the most prominent finding covers 72.36% of the time — 2.38% when the effects are one standard error. Intervals widened for the number of findings hold the share that miss under 5%.
A detector built for the ordering
The best of three checks for a drifting scale fires at half the growth factor the standard one needs — 2.12 against 4.31 — and still leaves 6.50 points of coverage gone before it does, against 0.51 for serial correlation. The reversal was not a property of the test.
The reference the sandwich is read against
The cluster-robust interval covers 75.05% at five clusters and 93.58% at eighty. The same estimate read against a t on G − 2 covers 87.95% at five, and the estimator is unchanged — three hundred rows grouped into five clusters cover 74.28% where the same three hundred grouped into seventy-five cover 94.63%.