Theme

The thread: The second number

A p-value cannot be interpreted alone. The same 0.04 corresponds to a large effect at ten observations and a negligible one at two thousand, and it means nothing at all until you know how many analyses were available to produce it.
One margin rises; the other turns over. The two halves of the table's margin across the sweep: a lower-tail copula's own leak with a symmetric covariate, and a covariate skewed at 0.95 under a Gaussian copula. The marginal's leak rises at every step, from 3.727% to 36.056%. The copula's does not: it rises to 9.064% at a Spearman of 0.6 and falls to 8.219% by 0.7. It has to turn over, because at a rank correlation of one the two variables are a deterministic function of each other and there is no interaction left for a split to leak. So the margin of the table turns over before any cell in it does. The same table at seven correlations

A margin that turns over

A skewed covariate's leak grows without limit as the dependence strengthens. A copula's own leak does not — it peaks at a rank correlation of 0.6 and falls. The margin of the table turns over before any cell in it does.

One factor moves and the other does not. The two factors of the same average, each drawn against its own largest value so that they share an axis. The rate at which the five candidates disagree about the tuning parameter rises from 28.6% at 4 values on the list to 43.3% at 8, a factor of 1.52. What a disagreement costs, given that there was one, is 0.00975 ± 0.00224 and 0.00848 ± 0.00113 at the same two points — 0.5 standard errors apart, and the paired comparison on the draws that disagree under both lists puts it the other way. The guess this field was written to test was that a longer list makes disagreements commoner and each one smaller. The first half is right and there is no second half. The rate and the size of a disagreement

A rate times a size

A sweep reported what it costs to let every candidate choose its own tuning parameter and found it flat across the list. It was reporting a product, and the two things multiplied together do not behave the same way at all.

What each variant loses before anything has been searched for. The mean loss differential of each of the eight variants against the benchmark, over 600 tables of 60 origins, with every fit given 71 rows. The series is an AR(1) and every variant adds a lag whose coefficient is zero, so in population the two forecasts are the same forecast and the difference drawn here is estimation noise and nothing else. The marked line is σ²(q₁ − q₀)/n = -0.01408, which is an expression in how many coefficients each model has and how many rows it was fitted on — it knows nothing about the series, the persistence or which lag the variant added, and every bar is within a fifth of it. This is the amount a reference distribution recentred at each column's own sample mean believes the candidates are already behind by. Searching among fitted models

A table of nested models

A benchmark and eight variants of it, each adding one thing. Every variant is behind before the search begins, by an amount that can be written down before the data exists — and the two most natural ways of reading the table are wrong in opposite directions.

The construction survives a difference of two weighted means. Coverage of δ̂ ± t√(S_D²/H) on b − 1 degrees of freedom, over 900 runs at a requirement of 0.3, where δ̂ is the block differences weighted by h_b = (1/m_A + 1/m_B)⁻¹ and H is their total. The theorem the one-mean field rests on goes through with h_b in place of the block size, and the reason is that the weights a weighted least squares decomposition needs are the inverse variances — which is exactly what h_b is. The stopping rule reads only within-arm within-block contrasts, so it is a function of nothing the interval reports, whatever it does with the block sizes. Each bar is within 2.9% of the level it claims. A promise about two arms

A width promised for a difference

The exact fixed-width interval was built for one mean. Two arms make the target 42.7 units of effective size and each unit costs four observations, so the same promise about a difference costs 169.4 rather than 42.7 — and the theorem survives untouched with the harmonic size in place of the block size.

Three intervals, one shortfall. What each of three intervals actually covers, at four rules and two block windows, over 300 samples of 120 rows. All three are built from the same resamples on the same draws, so a difference between them is a difference in what is done with the resampled series. Not one of the twenty-four cells reaches the ninety-five per cent it promises. The studentised interval runs from 75.7% to 92.3%, the percentile interval — the earlier field's — from 80.0% to 89.7%, and a normal interval on the same scale from 81.7% to 89.0%. The standard repair for a percentile interval's shortfall does not repair it. The interval, studentised

An interval that carries its scale

A percentile interval inherits the resampled distribution's skewness and its scale error together. The standard repair is one extra variance per resample. It was named and not run, so this runs it.

The area under the window is what the band actually costs. The three windows' weight sequences at a width of 30 lags, drawn against the lag as a share of the window. A truncated window applies a weight of one to every lag inside it and zero outside, which is why its sum is the width and why every conventional charge is right for it — and it is a covariance matrix on almost no sample, so it cannot be used. The Bartlett window falls linearly to zero and its weights sum to exactly 15.000000000000004, which is half the width, at every width: Σ(1 − k/(L+1)) over k = 1 … L is L − L/2. The Parzen window sums to 11.13 here, three eighths of the width, and it gets there by holding a weight near one over the first few lags and then falling faster. A plug-in estimate multiplied by a weight below one is a shrunk estimate, and a shrunk estimate is worth less than a free one — which is the whole of why a charge levied per lag is a charge for parameters the window has already spent. A charge for a covariance's own dimension

The charge nobody derived

A band of lags is charged one log-likelihood unit apiece, because that is what a regression coefficient costs. A band's numbers are not regression coefficients, and measuring what they actually cost puts the convention out by a factor of nearly three.

The eighth was not a constant. The probability that a per-candidate tuning list changes which candidate the table selects, at each of 5 separations between the candidates, over 800 draws apiece. The earlier field reports this flat at about an eighth across list length, on a table and a world it never varies. Vary how much the omitted coefficients are worth — one multiplier, with the table, the list, the law and the sample size all held — and it runs from 17.6% to 1.5%, a factor of 11.75. The world in which every candidate is true is the world in which the tuning list decides most; the world in which one candidate dominates is the world in which it decides nothing. What decides whether a tuning list decides

The eighth that was not a constant

How often a per-candidate tuning list changes which candidate wins is reported flat at about an eighth across list length. Vary how far apart the candidates are instead and it runs from 17.6% to 1.5%.

A quantile is the dearer reading, everywhere. The error each rule and window delivers on the two error readings, over 400 draws. The lower pair of lines is the implied long-run variance — the instrument the earlier field uses — and the upper pair is the 95% point of the standardised resampled mean, read against the finite-sample truth of 3.889 found by simulating the law directly. The quantile costs more at every one of the eight cells: at the plug-in rule it is 59.1% against 45.3% for the taper. That is not a defect in the bootstrap; a quantile is a statement about the shape of a distribution as well as its scale, and a fixed number of resamples estimates a tail worse than a variance. What matters for the comparison is that the two orderings between the windows are not the same, which the margins figure is about. The block length read on a quantile

The instrument and the reading

Every comparison between two block windows in this collection is an error in an implied long-run variance. Nobody reads a long-run variance. Read on the 95% point a test uses, the same bootstrap costs half as much again.

The curvature is in the denominator. The optimism a Bartlett band of each width actually costs, divided by that width, on two ways of measuring the width, over 2000 draws at 120 rows. Measured in the weights the band spends — Σ w(k), which is what the earlier field levies its charges on — the reading falls from 0.9528 at two lags to 0.7486 at thirty, so a charge proportional to the summed weights is too dear at one end and too cheap at the other. Measured in the pairs the band uses — Σ w(k)(1 − k/n), because a lag of k is an average over n − k products — the same readings are flat from 4 lags up, at 0.0084 of χ² per width against 0.2359. The correction has no fitted parameter in it: it is a function of the window, the width and the sample size. A charge that is not a straight line

The width a band is measured in

A tapered covariance band spends 84% of its own weights at two lags and 74% at thirty. Every charge in the collection is a straight line through the origin in those weights, so it is too dear at one end and too cheap at the other.

What each rung is made of. Each pair of searches, over 300 draws, split into the two effects its excess is the difference of. The overlap is what the second search loses by having the first already run at its own answer; the interaction is what the joint search finds by moving the first off it. They subtract to the excess exactly, on every draw, because the pinned supremum cancels. Two disjoint dictionaries of independent columns read an excess of 0.000011 and are made of 0.000514 and 0.000503. A break paired with a dictionary of step columns has an interaction of exactly 0 and is all overlap. And a break paired with an independent column has an overlap of -0.004395 against an interaction of 0.002364, which is what puts its excess below zero. Overlap and complementarity, separated

Two effects in one number

How much two searches over one sample share is measured as the net of two things — ground both of them find, and configurations only the joint search reaches. One extra supremum per draw separates them exactly.

What a mean split leaves, with both halves varying. The share of a mean split's interaction that survives the rule balancing it, at every copula and every marginal, matched at a Spearman correlation of 0.40. The three radially symmetric copulas leave exactly nothing with a symmetric covariate and rise steeply with the skew. The two asymmetric ones start at 7.707% and go opposite ways: the lower-tail copula falls to 0.002% at a skewness of 0.95 — the two failures cancel almost exactly, and a guarantee both fields report as broken is restored — while the upper-tail one climbs to 40.288%. And the heavy-tailed symmetric covariate, which leaks exactly nothing on its own, doubles what the asymmetric copulas leak: 14.229% against 7.707%. Both halves of the dependence at once

Two failures that cancel

A mildly skewed covariate under a lower-tail copula leaks 0.002% of an interaction where each failure alone leaks eight and seven per cent. Turn the copula over and the same pair compounds.

Two searches find some of the same luck. What each search reports on a sample with no break in it, and what the two report together, on four dependences. The dashed line is the sum of the two — what a rule charging each search separately would levy — and the two together always come in below it: 19.30, 15.44, 14.24, 26.64 short, on 100%, 99%, 99%, 100% of draws. The shortfall is not a rounding. Under AR(1) at 0.8 it is 19.30 of the 34.70 the break search manufactures on its own, which is more than half of it. Two searches over one sample are looking at the same noise, and the second one has less left to find. Two searches over one sample

Two searches, one sample

A searched break in a regression manufactures 34.7 of likelihood ratio where a count of coefficients says 11.1. A searched window manufactures 84.0. The two together manufacture 99.4, not 118.7.

What the rule blocks is not where it splits. How much of the separating direction each probe carries, over 100 designs of 14 units whose admissible set is enumerated and split into two pieces. The deferral this field answers proposed the constraint's active set — which exchanges the tolerance box actually blocks — as a better probe than the design's own leverage, on the ground that leverage is a heuristic and the active set is the quantity. Modelled from the design and the tolerance, it reads 0.4272 against leverage's 0.5395, at 4.43 paired standard errors the wrong way. Counted exactly over the enumerated set — at a cost no trial can pay — it reads 0.3854, worse again. Both beat a random direction at 0.2622, so they are probes; neither beats the two the earlier field already had. A probe from what the rule blocks

What the rule blocks

A balancing rule breaks the admissible set into pieces by refusing exchanges. Which exchanges it refuses is computable from the design and the tolerance alone, before any assignment exists — and it makes a probe.

What x-bar plus or minus 2 sample standard deviations holds, at n = 10. The content of the band is a random variable. Across 20,000 normal samples of 10 it averages 91.1%, its fifth percentile is 74.7%, and it falls short of 95% on 59.9% of samples. The band that is drawn to show where 95% of the data lies. The interval that holds observations, not a mean

Two standard deviations of what

The 95.45% inside two standard deviations is a fact about a curve whose centre and width are given. Drawn from ten observations, the same band holds 91.1% on average and less than 95% on 59.9% of samples — and the average is the reading that hides it.

Which side a 95% t interval misses on, exponential source. Both tails should be 2.5%. At 8 observations the interval falls short of the mean on 9.75% of samples and overshoots on 0.31%. At 500 they are 3.31% and 2.05%, and the total is 5.36% — which a coverage table reports as very nearly right. Shape, and what it does to a two-sample test

Where the two tails disagree

A 95% t interval on an exponential source at 120 observations covers 94.81%, which reads as very nearly right. It misses below the mean on 4.08% of samples and above on 1.11% — one tail 63% too heavy and the other 56% too light, and the total is the statistic that hides it.

What a second positive is worth, prevalence 0.10%. A 90% sensitive, 95% specific test. One positive gives 1.77%. Two independent positives give 24.49%, which is what multiplying the likelihood ratios says. At a correlation of 0.1 between the tests' errors it is 10.16%, and at 0.5 it is 3.16%. Two tests, a threshold, and the rate they are read against

The second test that is not a second opinion

Two positives from a 90/95 test on a one-in-a-thousand condition give a 24.49% chance of disease if the tests are independent. At a correlation of 0.1 between their errors it is 10.16%, and at 0.5 it is 3.16% — barely more than the 1.77% one positive was worth.

Which groups partial pooling serves, standard error 1 population width. Pooling's expected squared error for a group, divided by its own mean's, against how far the group truly sits from the centre. It is ×0.25 at the centre and crosses ×1 at 1.732 population widths, beyond which 8.33% of a normal population lies; capping the shift at one standard error holds every group under ×2. What partial pooling does to one group, to the set, and to a ranking

A group from the population's own tail

Partial pooling halves the total squared error when a group's own standard error equals the spread between groups. Every group whose true effect sits more than 1.73 population widths from the centre — 8.33% of a perfectly normal population — does worse than it would have with its own mean, and its loss grows without bound. Among eight groups with the spread estimated, the most extreme is worse off in 61.6% of datasets. Capping the shift at one standard error keeps the total at 0.528 of the unpooled error and holds every group under twice it.

The curvature is in the denominator. The optimism a Bartlett band of each width actually costs, divided by that width, on two ways of measuring the width, over 2000 draws at 120 rows. Measured in the weights the band spends — Σ w(k), which is what the earlier field levies its charges on — the reading falls from 0.9528 at two lags to 0.7486 at thirty, so a charge proportional to the summed weights is too dear at one end and too cheap at the other. Measured in the pairs the band uses — Σ w(k)(1 − k/n), because a lag of k is an average over n − k products — the same readings are flat from 4 lags up, at 0.0084 of χ² per width against 0.2359. The correction has no fitted parameter in it: it is a function of the window, the width and the sample size. A charge that is not a straight line

A lag the sample has less of

A sample autocovariance at lag k is an average over n − k products, not n. Count a band's width in the pairs it actually has and the curvature in its charge goes away, on a correction with nothing fitted in it.

Two instruments, two block lengths. The block length that would actually have been best on each draw, for each of the two error readings, averaged over 400 samples of 120 rows. For the rectangular window the implied long-run variance wants 18.92 and the 95% point wants 16.05; for the tapered window, 21.82 against 17.74. The quantile wants a shorter block under both windows — a ratio of 0.848 and 0.813. That is the mechanism the whole field turns on: a rule for choosing a block length is a way of guessing a target, and the two instruments do not have the same target. A rule tuned to one is systematically long for the other, and the two windows do not pay the same price for being long. The block length read on a quantile

A length for each instrument

The block length that is best for an implied variance is 18.92; the one best for the 95% point of the same resamples is 16.05. A rule is a way of guessing a target, and there are two targets.

A model of the active set, and the active set. Each of the 14 units of one design, at the share of its exchanges the tolerance box blocks — computed from the design's columns and the tolerance under a uniform position in the box, against counted over all 116 admissible assignments. The diagonal is where the two would agree. Over 192 designs they agree about the ordering of the units at a correlation of 0.8141 ± 0.0112, negative on 0.5% of them, and disagree about the level: 0.8442 counted against 0.8170 modelled, a gap of 0.0272 ± 0.0051. An admissible assignment does not sit uniformly in its box, and this is the size of that. A probe from what the rule blocks

A model and a count

The share of a unit's exchanges a tolerance box refuses can be modelled from the design or counted over the admissible set. They order the units the same way at a correlation of 0.81 and disagree about the level by 0.027.

Three charges, and only one of them is a test. What each of three thresholds does to the same decision, under AR(1) at 0.8, against the size of a genuine break in the mean at row 60. A chi-square on the 5 coefficients a split adds — 11.07 — declares a break on 73.6% of samples that have none: it is not a test at all. The break search's own 95% point, 69.6, carried into a rule that also chooses its window, fires on 0.0% of null samples and on 0.0% of samples with the largest break measured — the natural way of combining two published corrections does not lose a little power, it switches the test off. The calibrated charge, 27.2, holds 5.6% at no break and reaches 29.2% at the largest. Two searches over one sample

The charge that is not a sum

Charging two searches what each costs on its own is conservative, and conservative here means the test never fires. At the largest break measured it declares nothing, on every draw, while a calibrated threshold reaches 29%.

What a disagreement costs, split on whether it decided anything. The regret from choosing the tuning parameter per candidate, on the draws where the candidates disagreed, split on whether the disagreement changed which candidate the table selects. Over 1200 draws at each list length: when the winner changes the regret is 0.02215, 0.03029, 0.03145; when it does not it is -0.00243, -0.00069, -0.00065 — negative, and small enough that it is inside two standard errors of nothing at every length. The whole of the cost lives in the first column, and the second column is not merely small but slightly the wrong sign: when the table's answer is unaffected, letting each candidate use its own window is a very slightly better rule than making them share one. So a disagreement about the tuning parameter is not a cost. A disagreement that changes the winner is. The rate and the size of a disagreement

The quarrel that changes the winner

A disagreement about the tuning parameter costs 0.031 when it changes which candidate the table selects and −0.0007 when it does not. The distance between the values disagreed about has nothing to do with it.

No block size is best at both things the procedure claims. Two claims and one dial. The honest interval's half-width falls as the blocks get smaller, because the interval's degrees of freedom are the number of blocks: 0.2602 at blocks of two against 0.2933 at blocks of sixteen. The fixed-width claim — that the mean is within 0.25 of the truth — gets more reliable as they get larger, because the sample size is less variable: 92.40% against 95.00%. Both are computed from the same runs, and the second is reproduced to within a tenth of a point by E[2Φ(d√N/σ) − 1], which needs the sample-size distribution and nothing else. The schedules sit at the bottom left: as narrow as the smallest fixed block and as few observations, with the rule's spread estimate on half as many degrees of freedom again. The block size as a schedule

Two degrees of freedom, one total

The block size is a dial, and the two things a fixed-width procedure claims move in opposite directions along it. Divide the width by the square root of the sample size and one of them turns out to depend on the number of blocks and on nothing else.

Two factors, opposite directions. The two factors the cost of a per-candidate tuning parameter is a product of, as the candidates are pulled apart, over 800 draws at each of 5 separations. How often the candidates disagree about the tuning parameter rises from 31.8% to 88.8%; the share of those disagreements that change which candidate the table selects falls from 51.6% to 1.7%. So the setting where the candidates quarrel most about the tuning parameter is the setting where the quarrel matters least, and a sweep that reads the rate and stops has read the factor pointing the wrong way. What decides whether a tuning list decides

Two factors pointing opposite ways

As the candidates on a table are pulled apart, they quarrel about the tuning parameter three times as often and the quarrel decides the winner thirty times less often. A sweep that reads the first factor has read the one pointing the wrong way.

Power at an effect of 0.5 standard deviations. The curve is the non-central t on 2n − 2 degrees of freedom with δ = d√(n/2); the dots are 4,000 experiments run at each size. Reaching 80% power needs 64 per arm. Tests, and the second number

What a p-value does not say

The same p of 0.04 corresponds to a large effect in ten observations and a negligible one in two thousand. A p-value alone cannot be interpreted, and the number that makes it interpretable is almost never printed beside it.

What a positive test means, sensitivity 90%, specificity 95%. At a prevalence of one in a thousand, 98 of every hundred positives are false. At one in 10, 33 are. The test has not changed. Reversals that are not errors

What a positive test is worth

A test that is 90% sensitive and 95% specific sounds accurate. For a condition affecting one person in a thousand, 98% of its positive results are wrong, and a worse test on a commoner condition beats a better test on a rare one.

What studentising costs. How much wider the studentised interval is than the percentile one, cell by cell, over 300 draws, with what each cell gains in coverage beside it. Averaged over the eight cells the interval is 2.09 times as wide and covers 0.46 points better. At the two rules that choose short blocks the two intervals are within a fifth of each other; at the oracle's length, where a resample holds two or three whole blocks, the studentised interval is 4.37 and 5.04 times as wide. A repair that doubles the width and buys half a point is not one a reader could not have had by widening the interval it replaced. The interval, studentised

What studentising costs

Averaged over eight cells the studentised interval is 2.09 times as wide as the percentile one and covers 0.46 points better. At the block lengths the rules choose, the scale it divides by rests on two or three numbers.

The split decides the width. The width of the interval against the share of 200 observations spent on fitting rather than on calibrating, over 3000 draws. Spending more on the fit shrinks the residuals; spending more on calibration builds the interval at a less extreme order statistic. The two meet at 0.5, where the width is 4.0416 against 4.1603 at 0.1 and 4.3820 at 0.9. Full conformal, which spends the same 200 points on both jobs, is 3.9865 — so the whole cost of splitting is 1.38%. Coverage without a distribution

What the split costs

Splitting a sample between fitting and calibrating looks like a trade against the guarantee, and it is not: coverage moves 0.63 points across nine splits and every reading sits on its own promise. The whole cost is 1.38% of width — and at sixty observations the width falls, rises and falls again.

Three bands called 95%, at n = 20. Half-widths in sample standard deviations: 0.468 for the mean, 2.14 for one future observation, 2.75 to hold 95% of the population. The first two differ by exactly the square root of n + 1, which is 4.58 here. The interval that holds observations, not a mean

A tenth as wide, and both of them right

The interval for a mean and the interval for one future observation are both labelled 95%, and at a hundred observations one is 10.05 times the other — exactly the square root of n + 1. Read the narrow one as the wide one and it covers a new value 15.7% of the time.

Two studies with an R-squared of 0.85. The left study's points sit 0.50 from the line and the right study's 2.00 — a factor of 4.0. Both report an R-squared of 0.85 and, at the same sample size, the same standard error for the slope. The design was chosen to make it so, and it can always be chosen. What a summary of a scatter is a property of

The t statistic wearing different clothes

For a simple regression, t² = (n − 2)R²/(1 − R²), exactly, on every dataset — checked to sixteen significant figures over five hundred fits. So a paper reporting R² and a p-value has reported one number twice, and two studies with the same R² have points four times further from the line.

What the plot says, and what the interval does, at n = 40. For each source: how often a quantile plot of the data leaves its pointwise band, and how often the 95% t interval for the mean misses. The two-lump source leaves the band on 100% of samples and its interval covers 94.80%; the t on three degrees of freedom leaves it on 57% and covers 95.73%, the best of the five. What a diagnostic plot is showing

The plot is about the wrong quantity

A t interval needs the sampling distribution of the mean to be normal, not the data. A two-lump source leaves its quantile band on 100% of samples of forty and its interval covers 94.80%; a t on three degrees of freedom leaves it on 57% and covers 95.73%, the best of five sources.

Worst and average coverage of six 95% intervals for a proportion, 30 trials. The worst coverage over every proportion beside the average over a uniform one, with the average expected width. Clopper–Pearson: worst 95.05%, average 97.34%, width 0.299. Blaker: worst 95.00%, average 96.31%, width 0.283. Wilson: worst 83.71%, average 95.24%, width 0.271. A proportion's interval near the boundary, and the coin

What a guaranteed minimum costs

Clopper–Pearson's interval never covers less than 95%, and at thirty trials it averages 97.34% and is 10.4% wider than Wilson's. Blaker's interval keeps the same guarantee, averages 96.31% and is 4.6% wider. The difference is not waste: Clopper–Pearson guarantees each side separately, holding both below 2.5%, and Blaker guarantees only their sum — so at ten trials and a proportion of 0.15 it misses on one side 5.00% of the time.

Two 95% intervals 2.772 standard errors of the difference apart, standard errors in the ratio 1. The intervals are separated, and the test of the difference gives p = 0.0056. Two 95% intervals with equal standard errors just touch at p = 0.0056. An interval read beside something else

Two intervals that overlap

Two 95% intervals that just touch are read as a difference at the edge of significance. With equal standard errors their difference has p = 0.0056, not 0.05; two intervals can overlap by 58.6% of an arm and still differ at exactly 5%; standard-error bars that just touch mark p = 0.157; and when the two estimates are correlated at 0.8, touching intervals conceal a difference of 6.2 standard errors. Read as a test, non-overlap needs 1.66 times the sample for the same power.

How the truth, the raw means, the posterior means and the constrained estimates spread, standard error 1. Beyond two population widths above the centre lie 2.28% of the true values, 7.86% of the raw means, 0.234% of the posterior means, and 2.28% of the constrained estimates. What partial pooling does to one group, to the set, and to a ranking

Estimates that are too alike

Posterior means give each group its least-error estimate, and as a set they are too alike: with each group's standard error equal to the population's spread, they spread 0.707 as widely as the truth. Beyond two population widths lie 2.28% of the true effects, 7.86% of the groups' own means, and 0.234% of the posterior means — a tenth of the truth. Rescaling the estimates to the right spread counts the tail exactly and costs 17% more squared error; summing each group's posterior chance of being beyond the line counts it without changing any estimate.

The chance a trial succeeds against its size, when the expected effect of 0.5 is uncertain by four amounts. With the effect known, 80% is reached at 63 per arm. With the effect uncertain by 0.25 standard deviations it takes 113; by 0.5, 1268; by 0.75, no sample size at all, because the chance can never exceed the 74.8% prior probability that the effect is positive. What a sample-size calculation was given

The chance a trial succeeds

A trial of sixty-four per arm has 80% power at an effect of half a standard deviation. If the effect is only believed to be about half a standard deviation, give or take a quarter, the chance the trial reaches significance is 69.2%; give or take a half, 61.4%. Reaching 80% then takes 113 per arm, or 1,268 — and when the belief is uncertain by three quarters of a standard deviation no number of patients reaches 80%, because the chance can never exceed the 74.8% probability that the effect is positive at all.

What a search costs is not a property of that search. The likelihood ratio a searched break in the regression reports, two ways, on every law. On its own — the whole rule being a split of the sample, at no whitening — it averages 34.70 under AR(1) at 0.8, against the 11.07 a chi-square on the five coefficients a split adds would use as a threshold. Inside a rule that also chooses a window from a list of eight, the same search adds only 15.40 — less than half. Most of what a break search finds under correlated errors is the correlation, and a whitening chosen from the same sample has taken it already. A charge measured for one search, carried into a rule that makes two, is not conservative in some harmless direction: it is measuring a different quantity. Two searches over one sample

A charge that depends on the rule

The break search's charge is 34.7 on its own and 15.4 once a window has been chosen from the same sample. Most of what a break search finds under correlated errors is the correlation, and a whitening has taken it already.

Three copulas that break nothing, and a factor of two between them. The three radially symmetric copulas, at a matched Spearman correlation of 0.40, against the covariate's marginal. All three leave exactly nothing with a symmetric covariate — that is the guarantee, and it holds to twenty decimal places. What they do to a skewed covariate is not the same at all: at a skewness of 2.26 a Frank copula leaves 12.118% where a Gaussian leaves 21.539% and a t on four degrees of freedom leaves 23.640%. A factor of 2.0 between two copulas that are both symmetric, both matched on rank correlation, and both harmless on their own. So the copula matters to the marginal's leak without breaking any symmetry of its own, which is a milder version of the same finding and applies to every trial rather than to the asymmetric ones. Both halves of the dependence at once

A copula that halves a marginal

Three copulas break nothing on their own and put a factor of two between the same skewed covariate's leaks — 12.118% under a Frank against 23.640% under a t, at the same rank correlation.

A line in the right width beats two curves. How far each candidate charge sits from the measured optimism across the plateau, in units of each width's own standard error, over 2000 draws. The straight line through the origin in the band's summed weights — which is what the earlier field levies — misses by 0.2382 per width. The same straight line in the pairs the band actually uses, Σ w(k)(1 − k/n), misses by 0.0095. A fitted power law misses by 0.0293 and a fitted decaying rate by 0.0172, both on one fitted constant more. The deferral this field answers asked for a curve; the answer is a line, in a variable with nothing fitted in it. A charge that is not a straight line

A line that beats two curves

A deferral asked for a curve. Fitted against the same measurements, a straight line in a variable nobody had to fit describes the plateau better than either curve does with a constant more — and for three windows out of four it does not.

What each probe can see. How far apart the two components of the admissible set are on each probe, over the spread inside a component, on 100 designs whose set is enumerated and split. It is the population quantity a chain is trying to report. The separating direction itself reads 10.5646; the projected fourth power 5.0800, the design's own leverage 3.8362, the modelled active set 1.9529, the counted active set 2.0170 and a random direction in the same subspace 0.9422. The two active-set probes beat the random direction and lose to both of the earlier field's, which is the field's answer to the question that opened it. A probe from what the rule blocks

A quantity that loses to a heuristic

Leverage is a heuristic about which units a balancing rule has most to say about. The constraint's active set is the thing the rule actually does. As a probe, the heuristic wins by 4.4 paired standard errors.

The split depends on the order. How much two searches share, measured both ways round, over 300 draws. Pinning the first search at its own answer and searching the second gives one overlap; pinning the second and searching the first gives another. A break paired with an independent column reads -0.004395 one way and 0.002163 the other, at 8.70 paired standard errors and on opposite sides of zero. The excess the two components subtract to is the same in both orders by construction, so what changes is only how it is attributed. There is no order-free way to say which of two searches found ground both can reach, and the two orders bracket it. Overlap and complementarity, separated

A split that depends on the order

Run the second search first and pin that instead, and the same draw gives a different overlap and a different interaction — with the same difference. And one pair has no second order at all.

Flat along a row, apart between them. The probability that a per-candidate tuning list changes the winner, at three list lengths on three candidate tables, over 800 draws in each of the nine cells. Along a row — the reading the earlier field takes — it moves by a factor of at most 1.21, so that field's invariant survives on every table. Down a column it moves by up to 1.98. The list length is the dial that does not move this number and the table is one that does, and the earlier field varied only the first. What decides whether a tuning list decides

A table and a list

A nested ladder of candidates differing by one coefficient was predicted to turn over more often at every list length. It turns over less at every one, and its list changes the winner half as often.

Cut the charge and the width follows it. The band width each charge picks, averaged over 400 draws of 120 rows under AR(1) at 0.8, with the standard deviation across draws beside it. Schwarz's charge — half a log n a lag, which is 2.39 here — picks 3.67. Akaike's picks 6.02. Charging the numbers the window actually leaves free, which is half the width, picks 10.12; charging what the optimism measures, 0.767 of that, picks 14.15. A charge and the width it buys are very nearly reciprocal, which is what a likelihood rising at a fixed rate a lag implies and is why the four answers span a factor of 3.86. The width that was actually best on the draw averages 13.90 and moves by 10.30 from draw to draw — three times as much as any rule's answer does. A charge for a covariance's own dimension

A width that moves and an error that does not

Four charges give four widths a factor of four apart and four errors half a per cent apart. The derived charge wins, significantly, by a quarter of what was on offer — and none of the four is an estimate of anything.

Four cells change their answer. The four cells of the twenty whose excess changes sign as the dependence strengthens, over 7 recalibrations. Above the line the two failures compound — the cell leaks more than adding the copula's own leak and the marginal's — and below it they cancel. All four start above and end below, and all four are at the two most skewed covariates: skew 0.90 under heavy-tailed, skew 0.95 under heavy-tailed, skew 0.90 under upper tail, skew 0.95 under upper tail. Whether two failures of a dependence compound or cancel is therefore not a property of the pair. It is a property of the pair at a strength of dependence, and a fifth of the table changes its answer inside the range measured here. The same table at seven correlations

An answer that changes

Eleven of twenty cells cancel and nine compound, at one rank correlation. Sweep the correlation and four of the twenty change sides — all four from compounding to cancelling, all four at the most skewed covariates.

The quantity that does not depend on the list. The probability that letting each candidate choose its own tuning parameter changes which candidate the table selects — the product of the two moving shares — against the length of the list, over 1200 draws apiece. It is 14.2%, 11.9%, 12.3%: a spread of 2.2% across a list length that moves the disagreement rate by a factor of 1.52. This is the invariant the whole field turns on. Everything downstream of the winner — the coefficients, the regret, whatever a reader is going to quote — is a function of whether the winner changed, and how often that happens is not something the list controls. A longer list changes how often the candidates quarrel and not how often the quarrel matters. The rate and the size of a disagreement

How often it matters

The disagreement rate rises by half across the list and the share of disagreements that decide anything falls by nearly the same factor. Their product — how often the tuning list changes which candidate wins — sits at an eighth and does not move.

Two arms leave one degree of freedom per block unaccounted for. Each point is one run. The one-mean field's identity is (b − 1) + (N − b) = N − 1, and every schedule moves along that line rather than off it. Two arms give the rule N − 2b and the interval b − 1, which come to N − b − 1 — short of the N − 2 two arms leave by exactly one per block, since a block's arm counts absorb one degree of freedom each and only one of the two directions carries the difference. The hollow points add what the block sums are worth, b − 1 more, and land on the total. The missing degrees of freedom are not lost; they are in a place the interval has to be shown it may read. A promise about two arms

The degrees of freedom in the sums

One arm partitions N − 1 exactly. Two arms give the rule N − 2b and the interval b − 1, which is short by one per block — and the missing ones are in the block sums, which are correlated with the differences at −0.79 and are usable anyway.

Four rules of four change sign. The margin between the two block windows in points of coverage, under each of four rules, on each of three intervals built from the same resamples, over 300 draws. Positive is the tapered window covering better. On the percentile interval the taper wins at all four rules, by 5.33, 1.67, 5.00 and 4.00 points, which is the earlier field's own reading. On the studentised interval the rectangle wins at all four, by 4.33, 7.00, 4.33 and 2.67. And a normal interval, which uses no resampling at all, puts the two within a third of a point at every rule — so the disagreement is manufactured entirely by what is done with the resamples. The interval, studentised

The ordering reverses again

One field found two of four rules changing sign between two readings of one resampling. Turn the same resamples into a studentised interval instead of a percentile one and all four change sign.

Power to find a real effect of 3 standard errors, 10 of 20 real. no correction finds 85.1%, Bonferroni finds 49.1%, Holm finds 52.5%, Benjamini–Hochberg finds 74.9%. The uncorrected procedure finds the most and controls nothing. Corrections, and what each controls

The price of control

Every correction is paid for in power, and the exchange rate can be measured. Holm buys familywise control for 33 percentage points of power; Benjamini–Hochberg buys a weaker guarantee for 10. Neither is free and neither is a matter of taste.

The reversal is a property of the instrument. The margin between the two block windows under each of four rules, on three readings of the same resampled means, signed so that a positive bar is the tapered window winning. On the implied long-run variance the taper wins at the best available block length and at one estimated from the sample and loses at a length written into a protocol and at the rule of thumb — which is the reversal the earlier field's whole argument turns on, at 1.48 and 4.52 points. On the 95% point a test actually reads, the taper wins at all four, by 6.13 to 7.08 points. On the coverage the interval actually delivers, the taper wins at all four again, by 2.50 to 5.75 percentage points. Two of the four rules change sign between the first reading and the other two, and the two that change are exactly the two the earlier field's recommendation is about. The block length read on a quantile

The reversal that was the instrument's

On an implied variance the rectangle wins at a protocol length and at the rule of thumb. On the 95% point a test reads, and on the coverage an interval delivers, the taper wins at all four rules.

12,000 studies of a real effect of 0.3, n = 16. Power is 21%. The studies that reached significance report a mean effect of 0.621 — 2.07 times the truth. Every one of them is honest; the selection did the inflating. Tests, and the second number

The winner's curse

Filter honest studies down to the ones that reached significance and the effects they report are systematically too large. At low power the inflation is a factor of two, nobody has done anything wrong, and the selection did all of it.

The ladder is the same ladder under every law. Each pair's overlap under each of the four laws, over 400 draws apiece. The scale is fixed at both ends by construction: two searches over disjoint sets of independent columns read -0.001, 0.013, -0.002, 0.004, and a break search paired with a step column it already contains reads exactly one under every law. Between them the pair that reads one residual series twice runs 0.763, 0.805, 0.577, 0.752 — lowest under long memory, where a whitening has most to do and the break search has least left to find that the whitening has not taken. The rung that moves most is the pair of step dictionaries, from -0.148 under a moving average to 0.419 under a break; and the pair that is negative is negative under all four. Two searches over different features

Three quarters of the way to one search

The pair that started this reads 0.762 on a scale whose one is containment. And the pair that shares nothing but its response reads −0.306, so the sign the earlier field found does not transport at all.

The overshoot is the last block size and nothing else. A run stops at a multiple of its own block sizes and cannot land between them, so it ends past its own target by about half a block. Fixed sizes overshoot by 1.5, 2.2, 3.1, 4.7, 8.5 observations as the size goes 2, 3, 5, 8, 16. Every schedule here ends in blocks of two and every one of them lands where blocks of two land — 1.62, 1.32, 1.37 against 1.48 — while having spent most of the run inside blocks four and eight times larger. That is the one thing on this page a schedule genuinely takes from both ends. The block size as a schedule

What a schedule actually buys

Big blocks early and small blocks late is the right instinct and it does not take both ends of the trade, because there are not two ends to take. What it does take is the overshoot — about four per cent of the observations — and a steadier stopping point.

A filled value is not an observation. What a 95% interval for the slope actually covers after each way of handling 35.0% missing outcomes, counted over 4000 studies of 200 rows. Dropping the incomplete rows covers 95.93%. Filling with the observed mean covers 13.85%, because the estimate itself has moved. Filling with a fitted value covers 80.85% against a closed prediction of 79.73%: the estimate is right and the reported standard error is short by a factor of 0.6567 against a predicted 0.6500, because the residual sum of squares is divided by the whole sample's degrees of freedom. Adding residual noise recovers the spread and covers 85.78% against a predicted 84.62%, since the interval still ignores the variance of having imputed at all. The value that is not there

One imputation is not an observation

Three ways of filling a missing outcome, under a mechanism that makes dropping the rows beyond reproach. Filling with the observed mean covers 13.85%, filling with a fitted value covers 80.85%, adding noise covers 85.78%, and the thing all three were meant to improve on covers 95.93%.

What the family-wise correction does to the effect it lets through. At two standard errors the estimate that clears an uncorrected 5% threshold averages 1.35 times the truth, and the one that clears the family-wise threshold averages 1.69 times it. The correction fixes the error rate by demanding a larger estimate, and a larger estimate is a more selected one. The analyses that were available and not run

The correction that makes the estimate worse

Correcting for twenty analyses repairs the p-value by demanding a larger statistic, and a larger statistic is a more selected one. At two standard errors the surviving estimate averages 1.35 times the truth before the correction and 1.69 times it after — so the honest error rate is bought with a more inflated effect.

Estimating a prevalence of 0.10% from 1,000 tests. The positive rate reads 5.09%, which is 50.9 times the truth. The Rogan-Gladen correction averages 0.100% — unbiased — with a standard deviation of 0.820 points against the positive rate's 0.695, and it comes out negative on 48.6% of samples. Two tests, a threshold, and the rate they are read against

The prevalence the test has to estimate

Every predictive value takes a prevalence as given, and the prevalence is usually estimated from the same test's positive rate. At a true prevalence of one in a thousand that rate reads 5.09% — fifty times the truth — and the correction that inverts it is unbiased, 18% more variable, and negative on 48.6% of samples of a thousand.

Welch's test on skewed groups: the low and high rejection rates in every cell, with equal means throughout. Each cell should read 2.5 / 2.5. Two identical exponentials at 20 and 20 read 2.04 / 2.22; the worst cell, a wide exponential against a normal at 8 and 32, reads 9.79 / 0.47. Shape, and what it does to a two-sample test

The skewness of a difference

Welch's test holds its size to within half a point when both groups are normal. Give both groups the same skewed population and it still balances at twenty and twenty — and at eight and thirty-two it rejects low on 7.16% of samples and high on 0.66%. One number decides which: the skewness of the difference of the two means, which ranks twenty-five cells by their imbalance with a correlation of 0.997.

Two studies of the same effect, z statistics with mean 1.96: where one is significant and the other is not. Five hundred pairs. With the true effect identical in both, exactly one of the two is significant in 50.0% of pairs; among those, the difference between the two is significant in 9.7%. The dashed lines mark 1.96 on each axis; the diagonal lines mark a significant difference. An interval read beside something else

Significant in one, not in the other

Two studies of exactly the same effect, each with 50% power, disagree about significance half the time — and when they do, the test of the difference between them is significant in 9.75% of cases. A p of 0.01 beside a p of 0.20 is a difference with p = 0.36. Among four subgroups sharing one effect, at least one significant and one not happens 87.5% of the time, and the test that would tell a real difference apart needs four times the sample the effect itself needed.

How many of a league table's top ten are small groups, ranked four ways. A hundred groups with sizes from 4 to 400, of which 36% have twenty units or fewer. Small groups make up 36.3% of the true top ten, 62.1% of the top ten by raw means, 13.4% by posterior means and 22.9% by the posterior chance of being in the top ten. The three rankings recover 4.43, 5.38 and 5.47 of the true top ten. What partial pooling does to one group, to the set, and to a ranking

A league table of a hundred

A hundred groups with sizes from 4 to 400, and a top ten to publish. Ranked by their own means, small groups fill 62.1% of the top ten against their 36.3% share of the true top ten. Ranked by posterior means they fill 13.4%. The ranking built from each group's chance of being in the top ten recovers 5.47 of the true ten, the best of three and barely half; and the group ranked first could hold any rank from 1 to 31.

A wrong weight costs width; a random weight costs level. Five weightings on a trial whose variance ratio drifts by a factor of 20.1 between the first block and the last, over 4000 runs. The rule that knows every λ_b covers at 95.1% and sets the width. One ratio for the whole trial is wrong for every block and costs nothing in level — 94.8% — while being 20% wider; equal weights are calibrated by an identity and 22% wider. The ratio estimated inside each block is the only rule aimed at the quantity that actually varies, and it is the only one that misses the level, at 92.0%: a weight computed from a handful of degrees of freedom is mostly noise, and noise in a weight is not a wrong weight. Modelling the drift across blocks recovers the oracle's width at 94.8%. What a block may vary

A ratio that changes between blocks

A wrong weight costs width and a random weight costs level. The rule aimed at the quantity that actually varies is the only one that misses its own coverage, and the rule that models it across blocks recovers the whole of what knowing it is worth.

Power at an effect of 0.5 standard deviations. The curve is the non-central t on 2n − 2 degrees of freedom with δ = d√(n/2); the dots are 4,000 experiments run at each size. Reaching 80% power needs 64 per arm. Decided before the data

How many subjects

Sixty-four per arm for 80% power at half a standard deviation — a power figure that could only be simulated, with nothing to disagree with, until the non-central t was written. Two routes now, agreeing to within the simulation's own error.

The false-positive rate against the number of analyses, on pure noise. The data has no effect in it. Each analysis is correct and each p-value is honest. With 20 correlated outcomes available, something reaches p below 0.05 58% of the time. Tests, and the second number

Twenty analyses of nothing

Twenty honest, correct analyses of data with no effect in it find something significant 57% of the time. Nobody p-hacked, every individual p-value is right, and the reported one is the smallest of twenty.

The scale moves the width; the curve does not. The band width each charge picks, averaged over 150 draws of 120 rows. The two conventions — a unit a lag and half a log n a lag — pick 6.08 and 3.65 lags. The four charges derived from the measured optimism pick 15.05, 15.60, 15.47 and 15.44, against a best width on the draw of 14.13. So the scale a charge is levied on moves the width by a factor of 4.27 and the shape of the charge moves it by 3.6%. None of the six is an estimate of the draw's own best width: the correlations are -0.006, -0.003, 0.017, -0.001, 0.016, -0.004. A charge that is not a straight line

What a better charge buys

Four charges derived from the same measurements pick band widths within six per cent of each other and deliver errors within two per cent of the gap any of them leaves. The scale a charge is levied on decides the width; the shape of the charge decides nothing.

What the interval actually covers. The coverage of the two-sided interval each rule and window builds, over 400 samples of 120 rows, against the 95% it promises. Not one of the eight reaches it: the best is 91.0% and the worst is 80.8%, on a promise of 95%. So the first thing this instrument says is that the choice between the two windows is a choice inside a range that is already four to fourteen points short, which neither of the other two readings can express at all. The second is the ordering: the tapered window covers better under every one of the four rules, by 5.00, 2.50, 5.75 and 4.00 points — including at a length written into a protocol and at the rule of thumb, where the implied variance says the rectangle wins. The block length read on a quantile

What the interval covers

Eight rules and windows, and not one of them reaches its promised 95%. The range is 80.8% to 91.0%, and the choice between two block windows is a choice inside a shortfall that is four times larger.

Eight encompassing nulls, all of them true at once. The forecast under test is the variance-minimising combination of the eight candidates, and the first-order condition that defines its weights is cov(e_c, e_j) = var(e_c) for every j — so every one of the eight nulls is exactly true simultaneously and every rejection counted here is false. 800 draws of 60 origins. The largest of the eight statistics rejects 25.0% of the time at a nominal 5% and one comparison stated in advance rejects 3.1%, which makes the set worth about 9.1 independent comparisons — nearly the eight it has. Bonferroni, which is far inside its level on a search over accuracy comparisons of the same eight forecasters, is at 4.8% here. Searching among fitted models

When every null is true

A reality check assumes that every candidate in the set is exactly as good as the benchmark, which is a configuration nobody's data is ever in. Test a combination against its own parts and that configuration is not assumed — it is what the arithmetic makes true.

The table swept along the covariate instead. What a rule holding a mean of each covariate fails to remove of their interaction, at each of 4 copulas, as the covariate is skewed further and the rank correlation is held at 0.4. The sweep runs from a symmetric covariate at g = 0 to a skewness of 11.16, and the three settings the earlier table names — g = 0.3, 0.6 and 0.9 — are on it, where this sweep reproduces that table to the last digit. Every row rises and then falls: the lower-tail copula from 7.707% through 0.0015% and back to 5.587%, the upper-tail one to a maximum of 36.213%. So the quantity a trial is exposed to is not monotone in how skewed its covariate is. The same table at seven correlations

The other dial

The table is swept along the strength of the dependence and never along the shape of the covariate. Swept along the shape at a fixed correlation, the same two copulas cross, the same way — and the near-zero cell turns out to be a minimum in both directions at once.

One penalty, read along two dials. How much wider the studentised interval is than the percentile one, at every sample size and every block length on the grid, with the number of whole blocks each cell leaves written beneath. Read across a row and the block length changes; read down a column and the sample size does. The penalty is nearly a function of the block count alone: the cells at 15 blocks read 1.16, 1.20, 1.17, 1.15, 1.13, 1.10 across three sample sizes and three block lengths, while the cells at one block length read anything from 1.10 to 3.95. The largest penalty on the grid is 3.95, at the cell with 3 whole blocks in it. The interval, studentised

The count or the length

A block length and a block count are one number read two ways at one sample size. Read at three, the studentised interval's width penalty tracks the count — with an R² of 0.9911 against a closed form that has no length in it — and its coverage tracks the length.

The same dial, on a list that steps by one. The probability that a per-candidate tuning list changes which candidate the table selects, at each of 5 separations, for both tuning parameters at a matched list length of 8. The sieve order runs from 13.1% to 2.6%, a factor of 5.00; the whitening window, from 16.4% to 1.5%, a factor of 11.75. What is held is the number of options, the table, the law, the sample size and the seeds; what cannot be held is the size of a step, since an integer step and a geometric step are different amounts of change. The dial moves both, and it moves them by 2.35 times as much on one as on the other. What decides whether a tuning list decides

A step that is not a ratio

Run the separation sweep on a tuning list of integers rather than a geometric ladder and the two factors still point opposite ways. The invariant does not survive: along a row of integers the probability moves by 2.163 where along the geometric ladder it moves by 1.208.

Six scores, one coverage. What each nonconformity score's interval covers, over 2000 draws with 200 calibration points, against the 95.0249% the rank argument promises. The column runs from 94.30% to 94.90%, a spread of 0.60% against a standard error of a difference of 0.69% — one number, six times. That includes a score aimed five units off the fit and a score that never reads the response at all, because the rank argument does not read the score either: it needs the scores exchangeable and nothing else. Every decision a modeller makes has to show up somewhere else, and the next two readings are where. Coverage without a distribution

The score is the modelling

Six nonconformity scores on the same draws cover within 0.60 points of each other, against a standard error of a difference of 0.69 — one number six times. Their widths run over a factor of 2.361 and their adaptivity over a factor of 8.377.

Which samples Wilson and Clopper–Pearson each cover, n = 50, p = 0.2. Each bar is the probability of one count, shaded by which interval built on that count contains 0.2. Both cover 95.1% of samples, only Wilson 0.0%, only Clopper–Pearson 1.6%, neither 3.3%. The correlation between their hits is 0.810, so on shared draws the variance of their difference is 4.891 times smaller than on independent ones. What makes it checkable

The same draws for both methods

Two intervals computed on the same simulated datasets give a difference in coverage whose variance can be 4.891 times smaller than on separate datasets — or, for a pair that covers different samples, 1.164 times larger. Which one a comparison gets is an exact sum over the counts each interval covers, and a standard error that ignores the sharing covers 100.00% for one pair and 93.07% for the other.

Four 95% intervals for the mean of 30 exponential observations, split by the side they miss on. Both bars should read 2.5%. The t interval misses below the mean on 6.38% of samples. Widened until its total is exactly 5%, it misses below on 4.69% and above on 0.30%. Hall's transformation misses below on 3.31% and above on 1.96%. Shape, and what it does to a two-sample test

The side a bound is read from

On thirty exponential observations the upper limit of a 95% t interval is exceeded by the true mean 6.38% of the time, against the 2.5% a safety margin set from it assumes. Widen the interval until its total coverage is exactly 95% and the upper limit is still exceeded 4.69% of the time. A symmetric repair fixes the number that is reported and not the one that is used; Hall's transformation, which bends the interval, takes the same rate to 3.31%.

The factor that holds all of the next m observations with 95% probability, from a sample of 10. From 10 observations, the band for one future value has factor 2.371, for ten 3.716, for a hundred 4.942 and for a thousand 6.008. The 95%-content tolerance factor is 3.382 and is passed by m = 10; the Bonferroni stretch of the prediction factor reaches 7.567 at a thousand. The interval that holds observations, not a mean

All of the next ten

A warranty, a batch release or a monitoring rule promises something about every one of the next ten observations, not about one. From a sample of ten, the band that holds all ten with 95% probability reaches 3.716 sample standard deviations either side of the mean — already wider than the 3.382 of a tolerance interval for 95% of the population — and it keeps widening: 4.942 for a hundred, 6.008 for a thousand, with no ceiling. A 95% prediction interval, read as the answer, holds all ten 67.9% of the time: more than 0.95 to the tenth power, because the ten succeed and fail together.

The pairing recovers most of it and passes nothing. How much of the separating direction each probe carries, over 100 designs of 14 units whose admissible set is enumerated and split in two. The four rows the pairing adds are the dominant direction of what each blocking matrix keeps past its degrees, and the cut that direction's signs induce. Counted, they read 0.5266 and 0.5258 against the counted per-unit share's 0.3854 — most of the gap between that share and the design's own leverage at 0.5395, closed. Modelled, they read 0.4274 and 0.4954 against 0.4272. Nothing built from the active set passes leverage, and the projected fourth power is still ahead of all of them at 0.6583. A probe from what the rule blocks

A set of pairs, not a vector

The active set is a graph on the units, and every probe built from it so far has been its degree. Read as a graph it recovers 0.1326 of the alignment the summary lost — and draws level with leverage rather than passing it.

Largest where least is needed. What the pairs correction supplies against what each window's measured profile needs, across this field's plateau, over 2000 draws at 120 rows. Both are stated as the multiplicative rise the charge per unit of width has to take between four lags and thirty. What the correction supplies is arithmetic — (1 − μ(4)/n)/(1 − μ(30)/n), where μ is the mean lag of the weight the band adds — and it runs 1.0795, 1.0580, 1.0456, 1.0539 for the four windows. What the measurement needs runs 1.1076, 1.2928, 1.2296, 1.6550. The two orderings are opposite: the plain Bartlett window has the longest mean lag, so it gets the biggest correction, and the flattest profile, so it needs the smallest. They coincide to 0.9746 of each other, and nowhere else does the correction account for more than 85.0% of the fall. A charge that is not a straight line

What the correction assumes

A correction with nothing fitted in it repairs one window of four. The reason is that its size is set by where a window puts its weight and the curvature it must repair is set by something else — and for one window at one sample size the two happen to agree.

How long an honest forecaster looks broken for. The calibration error shown by a forecaster with no miscalibration in it at all, at five record lengths, drawn against one over the square root of the length so that the closed form is a straight line through the origin. It is 0.1252 at fifty forecasts and 0.0090 at ten thousand, against a closed form of √(2K/πn) times the mean root bin variance which gives 0.1257 and 0.0089. The threshold drawn across it is 0.02, a figure routinely read as evidence that something is wrong; the mean falls under it at 1976 forecasts and the 95th percentile at about 4111. Below that, an honest forecaster and a miscalibrated one are being told apart by a statistic that is mostly the sample size. A forecast that is a probability

The miscalibration a perfect forecaster shows

A forecaster whose true reliability is exactly zero shows a calibration error of 0.1252 on fifty forecasts and 0.0090 on ten thousand. Every one of 1,200 blameless hundred-forecast records exceeds the 0.02 routinely read as evidence of a problem, and the mean does not fall under it until 1,976 forecasts.

The p-value of a study with 80% power, twenty thousand times. Twenty thousand two-sided z-tests, each on 25 observations whose true mean is 0.5603 standard deviations from the null, a noncentrality of 2.802. The bars are the counted share of p-values in bins a quarter of a power of ten wide, with the leftmost bin holding everything smaller; the line is the closed form. The middle eighty per cent of the p-values runs from 4.4×10⁻⁵ to 0.13, 3.46 orders of magnitude, the median is 0.0051, and 80.0% fall below 0.05, which is what the power means. Tests, and the second number

The p-value a replication gets

Under a true null a p-value is flat. Under a real effect its distribution is closed form and wide — a study with 80% power returns anything from 4.4×10⁻⁵ to 0.13 in eight runs of ten — and the chance that an exact replication of a p = 0.05 result is significant again is exactly one half, under both of the models people use without naming them.

The risk of one cause, estimated two ways. Two causes of an ending event with constant hazards 0.2 (the one of interest) and 0.3 (the competitor), random dropout at 0.1 and follow-up to 6. The lower line is the cumulative incidence, (0.2/0.5)(1 − e^(−0.5t)), the chance of actually having had this event by t; the dots on it are the Aalen–Johansen estimate over 2000 studies of 300, 0.3670 at t = 5 against 0.3672. The upper line is 1 − e^(−0.2t), and the dots on it are one minus Kaplan–Meier with the competing event treated as censoring: 0.6318 at t = 5 against 0.6321. The second is larger by a factor of 1.722 at t = 5, and it is not an error of estimation. It estimates, correctly, the risk in a population where the competing cause does not exist. When the data stops early

One minus Kaplan–Meier is not a risk

With two ways for observation to end, one minus Kaplan–Meier for one cause reads 0.6318 at t = 5 where the chance of actually having had that event is 0.3670. Added across the two causes, the complements reach 1.4088 — more than the whole cohort. Nothing is estimated badly: the complement estimates, correctly, the risk in a world where the other cause does not exist.

Forty trials at a true effect of 0.16, under the rule "power at the trend < 10%". The upper line is the benefit boundary (4.56, 3.23, 2.63, 2.28, 2.04); the lower line is where the rule stops a trial for futility (0.40 at 80, 0.66 at 160, 0.95 at 240, 1.31 at 320). Of forty trials with a real effect, 29 cross for benefit and 11 are stopped for futility. Stopping rules

A boundary for giving up

Adding "stop if z is below zero" to an O'Brien–Fleming trial costs 5.20 points of power at the effect it was designed for and halves the observations a trial with no effect uses. Stopping when conditional power at the observed trend falls under 10% costs 13.23 points and stops 21.28% of trials with a real effect. Making that rule binding lowers the benefit boundary from 2.040 to 1.901, and a binding rule that is then ignored rejects a true null 3.523% of the time instead of 2.5%.

Storey's estimate of the share of true nulls over twenty thousand families, independent and correlated at 0.6. The true share is 0.5. Independent tests: mean 0.610, spread 0.160, below half the truth in 0.92% of families. Correlated at 0.6: mean 0.609, spread 0.240, below half the truth in 9.33%. Corrections, and what each controls

Estimating how many nulls are true

Benjamini–Hochberg at 5% delivers 2.55% when half of twenty nulls are false, because it cannot tell how many are. Storey's estimate of that share, read off the p-values above one half, spends the rest and finds 81.93% of the real effects instead of 74.70% on independent tests. Correlated at 0.9, the same procedure reports a finding in 19.29% of families in which every null is true.

Reading the draw changes what is charged, not what is tracked. The correlation between the band width each rule picks and the best band width on the same draw, over 400 draws. The three fixed charges read -0.069, -0.066, 0.012. The three that read the sample read -0.012, -0.019, -0.041. None of the six is distinguishable from nothing. The statistic the first plug-in reads does vary — the draw's own summed squared autocorrelation runs from 2.06 to 10.19 with a mean of 4.36 — so the failure is not that the charge stopped moving. It is that what it moves with carries no information about which width this draw wanted. A charge that is not a straight line

A charge that reads the draw

Three charges built to read the sample track the best band width on their own draw at −0.012, −0.019 and −0.041, deliver more error than the fixed rule they are calibrated to, and pick a width half again as variable. The statistic moves; the answer does not.

One treatment, a different hazard ratio at every follow-up. The hazard ratio a Cox model converges to, found as the root of its expected score by numerical integration, as the trial runs longer; dropout at 0.1 throughout. The proportional treatment reads 0.5 at every τ. The waning treatment reads 0.5000 at τ = 1, 0.6362 at τ = 3 and 0.7890 at τ = 8 — the same two arms, the same effect in the same first year, and a number that drifts towards one as later, effect-free events are added to the average. The dots are the mean of 400 Cox fits with 400 subjects an arm: 0.5014 at τ = 1, 0.7020 at τ = 2, 0.7635 at τ = 3, 0.8034 at τ = 5, 0.8208 at τ = 8. The crossing treatment reads 0.3429 at τ = 1, exactly 1 at τ = 3 by construction, and 1.0611 at τ = 8: beneficial, null or harmful according to when the trial stopped. When the data stops early

The hazard ratio the follow-up chose

A treatment that halves the hazard for one year and then does nothing has a Cox hazard ratio of 0.5000 if the trial stops at one year, 0.7617 at three and 0.8194 at eight. Nothing about the treatment differs between those numbers. When hazards are not proportional the hazard ratio is an average, and the length of follow-up and the dropout rate choose its weights.

Three intervals as one strength is spread thinner, at a concentration of 8. Coverage of three nominal 95.0% intervals on the same 1000 draws of 200 rows at each count, when a total concentration parameter of 8 is spread over 1 to 32 instruments. Two-stage least squares covers 97.2%, 96.4%, 94.0%, 86.7%, 73.2%, 51.5%. Building each row's fitted treatment from a first stage that never saw that row covers 97.1%, 97.2%, 98.3%, 97.8%, 97.9%, 98.7%. Limited-information maximum likelihood with its conventional standard error covers 97.2%, 96.7%, 95.7%, 90.9%, 85.2%, 79.0%. At one instrument the likelihood estimator is two-stage least squares exactly, which is why the first readings of those two agree to the last draw. A variable that moves one thing only

Leaving each row out of its own first stage

Spread a fixed first-stage strength over thirty-two instruments and two-stage least squares covers 51.5%. Build each row's fitted treatment from a first stage that never saw that row and the same draws cover 98.7% — through an interval 5.99 times as wide, around an estimate that misses by more than the whole effect on 34.7% of draws. At eight times the strength the same repair covers 95.3% and costs a width factor of 1.66.

One curve, and two forecasters that are each a single point. The ROC curve of an honest forecaster whose signal has correlation 1 with the latent state — every threshold on its probability, from the bivariate normal — and two forecasters that only ever say 0 or 1. The one an absolute-error score pays for says 1 wherever the honest probability exceeds a half: it has a true-positive rate of 0.6742 and a false-positive rate of 0.1375, and its "curve" is the two straight segments through that point, with area (TPR + TNR)/2 = 0.7684. The honest curve's area is 0.8683. Thresholding at the base rate of 0.3744 instead of at a half gives the largest area any two-valued forecaster can have here, 0.7818, and it is still 0.0866 short. A forecast that is a probability

The liar with two answers

The forecaster an absolute-error score pays for says only 0 or 1, and on the ROC square it is a single point: its area is (TPR + TNR)/2 = 0.7684, against the honest forecaster's 0.8683, and it falls below the honest one on 200 of 200 counted records. No relabelling of its two answers returns what it threw away — the best recovers a Brier score short of the honest one by exactly the 0.022154 of resolution lost — and below a signal correlation of 0.7332 the same score prefers saying no every time to an honest forecast.

What each group gains from being pooled, τ = 1. Eight groups whose sizes span a factor of 13.3. The smallest gains 2.076 of squared error, which is 36.8% of the total reduction; the largest gains 0.030, which is 0.5%. 2 of the eight account for half of everything pooling buys. Groups that borrow

Where the borrowing goes

Pooling cuts the total squared error across eight groups by 56%. Two of the eight take 61% of that reduction, the four best-measured groups share 11% between them, and the largest group gets 1.5% of what the smallest does. The headline is a fact about the groups nobody was asking about.

What each wrong count costs, 4 steps ahead. Squared forecast error 4 steps ahead at each imposed rank, relative to the correctly specified fit, at 200 observations. With 1 genuine relations, imposing 0 costs 13.3% and imposing 2 costs 4.8%. With 2 genuine relations, imposing 1 costs 15.6% and imposing 3 costs 2.5%. Under-counting is the more expensive mistake in both systems, and it is the one the procedure's level does not bound. Three series, and a count

Which mistake about the rank costs

On a system with two relations, imposing none costs 29.2% of squared forecast error and imposing three costs 2.5%. The expensive mistake is under-counting, which is the error the procedure's 5% does not bound — so the guarantee protects the cheap side.

Three contrasts on one dataset, three different splits. The variance-minimising allocation for each of three ways of reporting the same two-arm comparison, against the first arm's proportion, with the second at 0.1. A risk difference wants the arm with the larger p(1 − p) to get more units; a log odds ratio wants it to get fewer, and the two curves are exact reflections of each other in the half line. A log risk ratio wants something else again. At a first-arm proportion of 0.6 they ask for 62.0%, 21.4% and 38.0% of the units. A trial reporting more than one of them cannot be optimal for either. Splitting the units

Two contrasts, one split

A risk difference wants 62.0% of the units in the first arm, a log risk ratio wants 21.4% and a log odds ratio wants 38.0% — on one dataset, with one pair of proportions. The difference's rule and the odds ratio's are exact reflections of each other, so no split can be near-optimal for both.

Every finding Benjamini–Hochberg made in thirty families, with its interval, at real effects of 2. 85 findings, sorted by their estimate. 12 of their ordinary 95% intervals miss the true effect, every one of them on the far side; 6 of the wider false-coverage-rate intervals miss. Corrections, and what each controls

Intervals for the findings

Benjamini–Hochberg's findings usually go out each with its ordinary 95% interval. With ten real effects of two standard errors among twenty tests, 11.59% of those intervals miss their effect, every miss on the far side, and the interval around the most prominent finding covers 72.36% of the time — 2.38% when the effects are one standard error. Intervals widened for the number of findings hold the share that miss under 5%.

The honest curve, and the same forecaster in three coarse vocabularies. The ROC curve of the honest probability at full signal, area 0.8683, beside the same forecaster rounded to the nearest whole number, area 0.7684 — the two-valued liar — to the nearest half, area 0.8205, and to the nearest tenth, area 0.8650. A vocabulary of v values is v points on the square joined by straight segments, and tied reports count half. A forecast that is a probability

A forecaster that rounds

An honest probability issued in tenths loses 0.0033 of ROC area and 0.000708 of resolution — the variance its bands average away, and 89.5% of the 0.000792 it adds to the Brier score. Two hundred records of two thousand forecasts show that loss on 189; it takes about 3,300 forecasts to put it two standard errors from zero. And 3.207 in every thousand forecasts in tenths are a 0% on an event that happened, which a logarithmic score charges without limit.

Three intervals as one strength is spread thinner, at a concentration of 8. Coverage of four nominal 95.0% intervals on the same 1000 draws of 200 rows at each count, when a total concentration parameter of 8 is spread over 1 to 32 instruments. Two-stage least squares covers 97.2%, 96.4%, 94.0%, 86.7%, 73.2%, 51.5%. Building each row's fitted treatment from a first stage that never saw that row covers 97.1%, 97.2%, 98.3%, 97.8%, 97.9%, 98.7%. Limited-information maximum likelihood with its conventional standard error covers 97.2%, 96.7%, 95.7%, 90.9%, 85.2%, 79.0%. The same estimate with Bekker's many-instrument standard error covers 97.2%, 97.2%, 97.2%, 95.0%, 94.3%, 93.8%. At one instrument the likelihood estimator is two-stage least squares exactly, which is why the first readings of those two agree to the last draw. A variable that moves one thing only

A standard error that knows about the instruments

Limited-information maximum likelihood came out least biased when a concentration parameter of 8 was spread over thirty-two instruments, and its conventional interval covered 79.0%. Bekker's many-instrument standard error covers 93.8% on the same draws, at 63% of the jackknife's width — and it gets there with a median standard error of 0.561 against a true spread of 0.797, because it is large on the draws that need it. At eight times the strength it covers 94.9% at 91% of the jackknife's width, and nothing measured here beats it.

What the forecast interval is short by, φ = 0.85, 6 steps ahead. The plug-in interval covers 88.42% against a claimed 95%. Correcting the variance recovers 0.56 points, propagating the persistence's own standard error recovers 0.40, correcting the persistence recovers 2.66, and all three together recover 4.20 — leaving 2.38 points unaccounted for. Comparing two forecasters

What the interval is short by

The forecast interval covers 88.42% where it claims 95%. Correcting the persistence recovers 2.66 points, correcting the innovation variance 0.56, propagating the persistence's own standard error 0.40 — and all three together recover 4.20 of the 6.58, leaving a residual none of the standard repairs reaches.

How fast a gap has to close before a sample can see it close. The power of the test against the half-life of a disagreement, at 100, 200, 400 observations, each read against its own simulated critical value. Every pair in every reading is genuinely tied together, so a non-rejection is a miss. At 200 observations a gap that halves in 3 steps is found 99.9% of the time and one that halves in 12 steps is found 15.3% of the time — and by 35 steps the reading is 6.1%, which is the test's own size. Beyond that the curves are flat because there is nothing left to detect with. When the observations repeat each other

How slow a return a sample can see

At two hundred observations the test finds a gap that halves in five steps four times in five, one that halves in eight 37.3% of the time, and one that halves in fifty 4.95% of the time — which is the rate at which it finds pairs with no mechanism at all. The boundary moves with the sample, not with its square root.

One estimator, three answers, and only the reference changes. Coverage of the cluster-robust 95% interval for the slope against the number of clusters, at 30 rows in each. The estimator is identical in all three curves; what differs is the number it is compared against. At 5 clusters it covers 75.05% against a normal, 85.30% against a t on 4 degrees of freedom and 87.95% against a t on 3. At 80 clusters the three agree to within a point. The correction costs nothing: the same standard error, a different table. A standard error for a model that is wrong

The reference the sandwich is read against

The cluster-robust interval covers 75.05% at five clusters and 93.58% at eighty. The same estimate read against a t on G − 2 covers 87.95% at five, and the estimator is unchanged — three hundred rows grouped into five clusters cover 74.28% where the same three hundred grouped into seventy-five cover 94.63%.

A proxy removes less than its reliability, always. The share of the confounding bias removed by adjusting for a proxy, against how well the proxy measures the confounder. The diagonal is the answer a reader would guess — a covariate that is 80% signal removes 80% of the problem. The curve is what the arithmetic gives: the reliability, times one minus the squared correlation between the treatment and the confounder, divided by one minus the product of those two. That squared correlation is 0.4475. A reliability of 0.8 removes 68.85% and one of 0.6 removes 45.32%. The two agree only at the ends, and the gap is widest where most applied covariates sit. What conditioning on a variable does

Adjusting for a shadow

A covariate that is 80% signal removes 68.85% of the confounding, not 80% — the share is λ(1 − ρ²)/(1 − λρ²) and it is below the reliability everywhere. The residual bias is 0.1084 against an effect of 0.5, and at 25,600 rows it is 17.6 standard errors wide.

Three promises, and no procedure keeps all three. Average coverage and worst-case coverage for four 95% intervals for a proportion at n = 40, computed exactly. Their expected widths are 0.2418, 0.2417, 0.2472, 0.2641 in the same order. The textbook interval and the score interval have the same expected width to four digits — 0.2418 and 0.2417 — and worst-case coverages of 55.31% and 92.21%. The exact interval never breaks its promise and is 9.3% wider than the score interval to do it. Each of the three columns orders the four procedures differently. Intervals, counted

An interval that covers and says nothing

A procedure returning the whole line 95% of the time and the empty set otherwise has coverage exactly 95% at every parameter value. Two real intervals at forty observations have expected widths of 0.2418 and 0.2417 and worst-case coverages of 55.31% and 92.21%.

Two companions on one simulation, two hundredfold apart. How many times as many draws each companion is worth, on the same 4,000 simulated samples of 40 observations. The coverage of the interval is estimated with the observed count as its companion, whose expectation is 12 exactly; they correlate at 0.2665 and the companion is worth 1.08 times the draws. The expected width is estimated with p̂(1 − p̂) as its companion, whose expectation is 0.20475 exactly; they correlate at 0.9977 because the width is a monotone function of it, and the companion is worth 214 times the draws — 856 thousand simulated samples' worth of precision from four thousand. What makes it checkable

The check worth more than the check

The same exactly known companion that verifies a simulation can sharpen it. On one set of four thousand draws, one companion is worth 1.08 times the draws and another is worth 214 times them, and the factor is 1 − ρ² with nothing else in it.

Each interval covers one question and not the other. Coverage of each interval for the overall mean, scored against both estimands, over 20,000 two-site studies of 10 observations apiece. The fixed-effect interval covers the mean of the two sites in hand 96.37% of the time and the population mean 54.77%. The random-effects interval covers the population mean 94.96% — exactly its level, from one degree of freedom — and over-covers the two sites in hand at 98.25%. Both are correct; they are answers to different questions printed in the same place. Hierarchy past one number

What a two-unit study should report

The fixed-effect interval covers the mean of the two sites in hand 96.37% of the time and the population mean 54.77%. The random-effects interval covers the population mean 94.96% — exactly its level, from one degree of freedom — and is 11.6 times wider.

All themes