Concept

Sample size — where it appears

How many observations an experiment takes, which is either decided in advance from a guess or decided by the data from a rule. Deciding it from the data costs nothing unless the rule read the same quantity the interval afterwards reports.

Named by 63 essays across 31 fields — each of them below, with the objects they name alongside it.

One experiment, with the blocks getting smaller as the target comes into range. A single run at a requirement of 0.25, with the block sizes 5, 5, 11, 25, 11, 8, 3, 2 and a total of 70 observations in 8 blocks. The rule stops when the observations in hand reach z²σ̂²/d², with σ̂² pooled from the within-block contrasts — an estimate that moves as the run goes on, so the target moves too. Early blocks are large because the target is far away and cannot be overshot; late ones are small because a block is the granularity of the answer. The interval afterwards is built from the 8 block means and from nothing the rule looked at, and it has 7 degrees of freedom against the rule's 62.

A block size that changes

The blinded rule's exactness never needed the blocks to be the same size. Letting the size be chosen from the contrasts as the run goes on leaves the coverage exactly where it was — and runs straight into an identity that says what a schedule can and cannot buy.

pace · Stopping
What the interim sees, at an effect of 1. The same 60 observations, estimated two ways. Keeping the arms separate gives 0.995, which is σ. Pooling them without separating the arms — the price of staying blind to the comparison — gives 1.114, against the identity √(1 + Δ²/4σ²) = 1.118. The sample size is proportional to the variance, so a blinded design at this effect asks for 25% more units than it needs, and it does so systematically rather than by chance.

Choosing n after looking

Re-estimating the sample size from an interim is the one adaptation with a defence, and the defence is exactly what it costs: an analyst kept blind to the arms measures a spread that contains the effect, so the design overshoots by 1 + Δ²/4σ². Re-estimating the effect instead breaks the error rate.

adaptive · Stopping
Every split of 100 units, σ = 1 against 3. Each point is one integer split, with its variance computed exactly rather than simulated. The minimum is at 25:75, which is the ratio of the spreads 25:75, and equal allocation costs 25% more variance — the same as throwing away 20 of the 100 units. The shaded band is every split within 5% of the best, and it runs from 17% to 35%: sharp to state, flat to sit on.

Not half and half

The same units, the same measurements, the same analysis — and a different variance, decided before anything is measured. When the two arms have different spreads the best split is σ₁ : σ₂, equal allocation costs 2(σ₁²+σ₂²)/(σ₁+σ₂)², and at three to one that is a quarter of the experiment.

allocation · Allocation
Twenty series with a lag-one correlation of 0.8. Every series has a true mean of zero and 60 observations. The marks on the right are the twenty sample means. The variance of that mean is 8.3 times what 60 independent observations would give, so the series is worth about 7 of them.

The observations that repeat each other

Almost every standard error divides by √n, which claims the observations carry independent information. At a lag-one correlation of 0.8 a fifty-point series is worth about six independent observations, and its 95% interval covers 47%.

timeseries · Dependence
Six groups of 10, each fitting its own slope, then borrowing. Each faint line is one group's own least-squares slope through its own centre; each solid line is that slope after pooling towards the population slope of 0.79. Every group has the same 10 observations. The group whose x values span 0.4 has a slope standard error of 1.86 and moves 91% of the way in; the group spanning 2.0 has a standard error of 0.37 and moves 28%.

The slope that borrows

Pooling a mean makes it look as though how much a group borrows depends on how much data it has. Pool a slope instead and the illusion breaks — ten groups with ten observations each can borrow anything from 28% to 91%, decided entirely by where those ten observations were placed.

multilevel · Levels
Coverage of four nominal 95% intervals, n = 20. Computed exactly by summing over all 21 possible counts, not simulated. The Wald interval drops to 18.2% and is jagged everywhere; Clopper–Pearson never falls below 95% and pays for it in width.

What the 95% refers to

An interval that claims 95% is making a checkable statement about a procedure, not about the interval in front of you. Build every possible sample and count, and the interval taught first turns out to cover 87.6% of the time.

intervals · Coverage
weakly informative — Beta(2, 2), updated by 5 of 20. The prior is worth 4 observations. With 20 observations the posterior mean is 0.292, against a data proportion of 0.250 and a prior mean of 0.500.

What a prior is worth

A prior is not a philosophical position, it is a component with a stated size. For a proportion it is worth exactly a + b observations, which turns "how much does the prior matter" from an argument into a subtraction.

bayes · Prior
8 exponential draws, standardised, against the normal. The source is one-sided and skewed. At n = 8 the standardised sum has skew 0.695, and the theory says 2/sqrt(n) = 0.707 — so the convergence is visible AND its rate is predicted.

Sums of almost anything

The theorem says sums converge on one shape whatever they are sums of, which is remarkable and true. Watching it happen from a one-sided skewed source, with the rate of convergence predicted in advance, is more convincing than watching the shape appear.

normal · Clt
What x-bar plus or minus 2 sample standard deviations holds, at n = 10. The content of the band is a random variable. Across 20,000 normal samples of 10 it averages 91.1%, its fifth percentile is 74.7%, and it falls short of 95% on 59.9% of samples. The band that is drawn to show where 95% of the data lies.

Two standard deviations of what

The 95.45% inside two standard deviations is a fact about a curve whose centre and width are given. Drawn from ten observations, the same band holds 91.1% on average and less than 95% on 59.9% of samples — and the average is the reading that hides it.

estimated · Bands
How often a randomised trial reports the reversal, advantage 6 points. Simple randomisation against randomisation stratified by group, 4,000 trials at each size. The simple design reverses on 3.40% of trials at its worst size and 0.50% at 1280 units; the stratified design reverses on none of them, at any size.

The reversal a coin cannot prevent

Randomisation removes Simpson's reversal in expectation, which is not the same as removing it. A correctly randomised trial of eighty units, on a population where the treatment helps in both groups, reports it losing overall on 3.40% of trials — and stratifying the randomisation takes that to zero at every size.

reversal · Simpson
Which side a 95% t interval misses on, exponential source. Both tails should be 2.5%. At 8 observations the interval falls short of the mean on 9.75% of samples and overshoots on 0.31%. At 500 they are 3.31% and 2.05%, and the total is 5.36% — which a coverage table reports as very nearly right.

Where the two tails disagree

A 95% t interval on an exponential source at 120 observations covers 94.81%, which reads as very nearly right. It misses below the mean on 4.08% of samples and above on 1.11% — one tail 63% too heavy and the other 56% too light, and the total is the statistic that hides it.

tails · Student
The power trials actually have when sized for 80% from a pilot of 10. Four thousand pilots of 10 observations, each sizing a trial for 80% power at half a standard deviation from its own standard deviation. 55.9% of the trials have less than 80% power and 11.1% less than 50%; the median trial has 76.8%.

The spread a pilot supplies

A trial sized for 80% power from a pilot's standard deviation is sized from an estimate that is too small more often than not. With a pilot of ten, 55.9% of the trials it sizes have less than 80% power and 11.1% less than 50%, although the planned sample is right on average. Sizing from the pilot's 80% upper confidence limit instead leaves 19.8% short, at 1.65 times the sample; from its 90% limit, 10.0% short at 2.12 times.

planned · Power
Coverage against sample size, true proportion 0.15. Coverage does not improve monotonically. n = 19 covers 93.8% while the larger n = 20 covers 81.9%. The sample space is discrete, so the endpoints jump as n changes.

More data is not monotonically better

Coverage of an interval for a proportion does not improve smoothly as the sample grows. It oscillates, and there are larger samples that cover materially worse than smaller ones — a sample of twenty covers twelve points worse than a sample of nineteen.

intervals · Oscillation
The one candidate an effective sample size is right about. n/n_eff with the finite-sample inflation Σ(1 − |k|/n)ρ^|k| is not an approximation to tr(HΩ) for a fit with only an intercept — it is that trace, to machine precision, because the hat matrix of a constant column is 1/n everywhere and its trace against Ω is the mean of Ω. The quoted limit form n(1 − ρ)/(1 + ρ) is not even right about that one. And the average correction the table's fifteen candidates actually need is 3.318 per parameter, well below the scalar, so applying it to all of them over-charges every one.

One number for a table of candidates

An effective sample size is a real quantity, it is exactly right about one thing, and that thing is a mean. Substituted into Akaike's criterion it changes nothing at all, because the penalty it is meant to fix has no sample size in it.

effective · Dependence
A budget of 4,000, at 1 and 20 a unit. Every affordable pair, enumerated. The best is 280 cheap units and 186 expensive ones — a ratio of 1.51, against the σᵢ/√cᵢ rule's 1.49. The unit rule, which says buy in the ratio of the spreads, lands at 66:197 and costs 17% more variance for the same money. Both rules are right about their own constraint; only one of them was asked.

The cost of a unit

Change the constraint from units to money and the allocation rule changes with it — from σᵢ to σᵢ/√cᵢ, which can point the other way. An arm that is noisy and expensive gets fewer units than the same arm would if the money were not the thing running out.

allocation · Allocation
Where the normal approximation converges, and where it does not. Relative error against the exact binomial. At n = 1280 the error at the median is 0.96% and three sigma out it is 25.7% — a factor of 27. The tail is where the approximation is used.

The tail converges last

The central limit theorem is usually shown as a shape arriving. What the demonstration leaves out is the rate — and the rate is wildly different in the middle and in the tail, which is where every approximation in the subject is actually read.

normal · Rate
No block size is best at both things the procedure claims. Two claims and one dial. The honest interval's half-width falls as the blocks get smaller, because the interval's degrees of freedom are the number of blocks: 0.2602 at blocks of two against 0.2933 at blocks of sixteen. The fixed-width claim — that the mean is within 0.25 of the truth — gets more reliable as they get larger, because the sample size is less variable: 92.40% against 95.00%. Both are computed from the same runs, and the second is reproduced to within a tenth of a point by E[2Φ(d√N/σ) − 1], which needs the sample-size distribution and nothing else. The schedules sit at the bottom left: as narrow as the smallest fixed block and as few observations, with the rule's spread estimate on half as many degrees of freedom again.

Two degrees of freedom, one total

The block size is a dial, and the two things a fixed-width procedure claims move in opposite directions along it. Divide the width by the square root of the sample size and one of them turns out to depend on the number of blocks and on nothing else.

pace · Width
What a 95% credible interval covers, n = 20. Computed by summing over all 21 possible counts rather than by simulating them. Jeffreys' prior covers close to 95% across the range; a confident prior centred in the wrong place covers almost nothing where the truth is far from it.

What a credible interval covers

A credible interval makes the statement everyone wants and does not claim to have a coverage. It has one anyway, it can be summed over the sample space exactly, and on a reasonable prior it beats the interval taught first.

bayes · Credible
Power at an effect of 0.5 standard deviations. The curve is the non-central t on 2n − 2 degrees of freedom with δ = d√(n/2); the dots are 4,000 experiments run at each size. Reaching 80% power needs 64 per arm.

What a p-value does not say

The same p of 0.04 corresponds to a large effect in ten observations and a negligible one in two thousand. A p-value alone cannot be interpreted, and the number that makes it interpretable is almost never printed beside it.

testing · Power
One distribution, two effects. Three causal structures fitted to one covariance matrix over a treatment, a covariate and an outcome. Each reproduces it exactly — the largest entry-wise disagreement across all three is 4.4e-16 — so no sample of any size distinguishes them. The regression of the outcome on the treatment and the covariate returns 0.500 in all three, to within 4.4e-16, because that coefficient is a function of the covariance and of nothing else. The effect the three worlds hold is 0.500, 0.848 and 0.848: adjusting is exactly right in the first and off by −0.348 in the other two. The arithmetic cannot see the difference and the difference is the whole question.

The two worlds that look the same

Three causal structures were fitted to one covariance matrix and agree with it to 4.4·10⁻¹⁶. The regression returns 0.5000 under all three; the effect they hold is 0.5000, 0.8481 and 0.8481. What separates structures is a missing edge, and the signature of one is a correlation of exactly zero.

collider · Conditioning
Three bands called 95%, at n = 20. Half-widths in sample standard deviations: 0.468 for the mean, 2.14 for one future observation, 2.75 to hold 95% of the population. The first two differ by exactly the square root of n + 1, which is 4.58 here.

A tenth as wide, and both of them right

The interval for a mean and the interval for one future observation are both labelled 95%, and at a hundred observations one is 10.05 times the other — exactly the square root of n + 1. Read the narrow one as the wide one and it covers a new value 15.7% of the time.

estimated · Bands
Two studies with an R-squared of 0.85. The left study's points sit 0.50 from the line and the right study's 2.00 — a factor of 4.0. Both report an R-squared of 0.85 and, at the same sample size, the same standard error for the slope. The design was chosen to make it so, and it can always be chosen.

The t statistic wearing different clothes

For a simple regression, t² = (n − 2)R²/(1 − R²), exactly, on every dataset — checked to sixteen significant figures over five hundred fits. So a paper reporting R² and a p-value has reported one number twice, and two studies with the same R² have points four times further from the line.

spread · Summary
The pooled two-sample test's size, with a true null everywhere. Forty units split between two groups, with the second group's variance a stated multiple of the first's, and the two population means equal. A 5% test should reject 5% of the time. The pooled test runs from 0.55% to 18.91% across this region; Welch's runs from 4.63% to 5.51%.

A degrees of freedom that is not a count

The pooled two-sample test's size runs from 0.55% to 18.91% across forty units split five ways against five variance ratios, with a true null in every cell. Welch's runs from 4.63% to 5.51% — bought with a degrees of freedom that is a function of the data, not an integer, and not a count of anything.

tails · Student
The chance a trial succeeds against its size, when the expected effect of 0.5 is uncertain by four amounts. With the effect known, 80% is reached at 63 per arm. With the effect uncertain by 0.25 standard deviations it takes 113; by 0.5, 1268; by 0.75, no sample size at all, because the chance can never exceed the 74.8% prior probability that the effect is positive.

The chance a trial succeeds

A trial of sixty-four per arm has 80% power at an effect of half a standard deviation. If the effect is only believed to be about half a standard deviation, give or take a quarter, the chance the trial reaches significance is 69.2%; give or take a half, 61.4%. Reaching 80% then takes 113 per arm, or 1,268 — and when the belief is uncertain by three quarters of a standard deviation no number of patients reaches 80%, because the chance can never exceed the 74.8% probability that the effect is positive at all.

planned · Power
The best of 8 arms, tested as though it were the only one. 8,000 trials with no effect in any arm. Stage one runs 8 arms at 60 each, the best is carried forward, and stage two adds 30 more to it and to the control. The histogram is where the final statistic lands and the curve is the standard normal it is being read against — shifted right, because the arm was chosen for being ahead. 10.5% of these trials clear 1.96 against a claimed 5%, and the value that actually holds the rate for this design is 2.313.

Dropping the losers

Carrying the best of eight arms forward and testing it at 1.96 rejects a true null 10.3% of the time — the hypothesis was chosen by looking at the data, so the statistic is a maximum wearing a single comparison's clothes. The value that holds the rate is 2.313, and it has to be solved for.

adaptive · Stopping
R² against the number of useless predictors, n = 30. The response is pure noise and so is every predictor, so the true relationship is nothing at all. R² rises from 0.000 to 0.648 anyway, following k/(n − 1) — which is what a criterion that rewards higher R² is actually rewarding.

R² is not a measure of fit

Adding a predictor with no relationship to anything cannot reduce R², and in expectation raises it by 1/(n − 1). Twenty useless predictors on thirty points give an R² of 0.69 from pure noise.

regression · Summary
One rule keeps its promise and the other keeps its budget. Both stopping rules at five requirements, 1,500 experiments each, with a first stage of 5. The upper curve is the two-stage rule: 97.1%, 96.2%, 95.9%, 96.1%, 96.0% — at or above 95% at every point, which is a theorem rather than a tendency, because its interval is built from a spread estimated before the stopping point was chosen. It pays 2.06×, 2.01×, 2.00×, 1.99×, 1.99× the observations that knowing σ would need. The lower curve is the rule that re-estimates after every observation: 94.5%, 89.3%, 91.0%, 91.5%, 94.3%, on 0.98×, 0.87×, 0.88×, 0.93×, 0.96×. The second rule is the one anybody would run and the first is the one whose claim is true.

Stopping when it is precise enough

An experiment that runs until its estimate is precise enough is the natural design and the one with a theorem against it. Its two-stage cousin keeps its promise exactly, for every unknown spread, and pays twice the observations for it.

guarantee · Stopping
Power to find a real effect of 3 standard errors, 10 of 20 real. no correction finds 85.1%, Bonferroni finds 49.1%, Holm finds 52.5%, Benjamini–Hochberg finds 74.9%. The uncorrected procedure finds the most and controls nothing.

The price of control

Every correction is paid for in power, and the exchange rate can be measured. Holm buys familywise control for 33 percentage points of power; Benjamini–Hochberg buys a weaker guarantee for 10. Neither is free and neither is a matter of taste.

multiplicity · Multiplicity
Exact coverage, at every block size. Coverage of the interval each rule reports, at a nominal 95%, over 2,500 runs each with a standard error of 0.44 points. The blinded rule stops on the within-block contrasts and reports an interval built from the block means, and those two are independent whatever the rule does — so the interval is an ordinary t interval on b − 1 degrees of freedom and its coverage is exact. It is exact at every block size drawn. The interval a practitioner writes at the purely sequential rule's stopping time covers 91.72%, and Stein's two-stage rule is exact for the same reason as the blinded rule and spends 2.10 times the observations to be so. The bars are truncated at 86% so the differences can be seen.

The rule that cannot see the mean

A sequential rule stops when its own estimate of the spread is small, which is more often on the samples whose spread came out low — so the interval afterwards is short. There is a way to keep updating the estimate and stop being able to see the mean at all.

blind · Stopping
12,000 studies of a real effect of 0.3, n = 16. Power is 21%. The studies that reached significance report a mean effect of 0.621 — 2.07 times the truth. Every one of them is honest; the selection did the inflating.

The winner's curse

Filter honest studies down to the ones that reached significance and the effects they report are systematically too large. At low power the inflation is a factor of two, nobody has done anything wrong, and the selection did all of it.

testing · Curse
Twenty 95% intervals for a proportion that really is 0.35. 2 of the twenty miss the true value. The 95% is a property of the procedure across repetitions — no single interval has a 95% chance of anything, because it either contains 0.35 or it does not.

Twenty intervals and one expected miss

The 95% belongs to the procedure, not to the interval in front of you. Twenty intervals from twenty samples make that visible in a way no definition does, and the one that misses is not a mistake.

intervals · Repetition
The overshoot is the last block size and nothing else. A run stops at a multiple of its own block sizes and cannot land between them, so it ends past its own target by about half a block. Fixed sizes overshoot by 1.5, 2.2, 3.1, 4.7, 8.5 observations as the size goes 2, 3, 5, 8, 16. Every schedule here ends in blocks of two and every one of them lands where blocks of two land — 1.62, 1.32, 1.37 against 1.48 — while having spent most of the run inside blocks four and eight times larger. That is the one thing on this page a schedule genuinely takes from both ends.

What a schedule actually buys

Big blocks early and small blocks late is the right instinct and it does not take both ends of the trade, because there are not two ends to take. What it does take is the overshoot — about four per cent of the observations — and a steadier stopping point.

pace · Nuisance
What differencing fixes, and what it costs, 100 steps. The first pair is the false-positive rate for two independent random walks: 77% on the levels, 4.9% on the differences. The second pair is how much of a real relationship survives: R² falls from 0.91 to 0.33. The same operation does both.

What differencing costs

Differencing takes the false-positive rate between two unrelated walks from 76.7% to 4.9%, and takes a genuine relationship's R² from 0.91 to 0.33. Applied to a series that did not need it, it doubles the variance and installs a correlation of −0.5 that the data never had.

timeseries · Spurious
Twenty samples of 40, every one of them genuinely normal. Each panel is a quantile-quantile plot of 40 draws from a normal distribution. The worst point in the worst panel sits 0.87 standard deviations off the line. Anything a reader would reject here would be a false alarm.

What normal actually looks like

A single quantile plot of forty normal points wanders enough to look suspicious. Twenty of them, all genuinely normal, show what the noise looks like — and any single panel a reader would have rejected is in there.

normal · Qq
The crossing barely moves. Both methods' costs in one unit — assignments evaluated per usable draw — as the tolerance tightens. A hunt costs 1/p and rises without limit: from 2.22 at a tolerance of 1.2 to 357.14 at 0.18. A walk costs its autocorrelation time and barely moves. The two cross at a tolerance of 0.190 at one swap and 0.195 at eight — the whole family of proposal sizes crosses inside a band of about two hundredths, because where the crossing is, the large proposal has already lost its advantage. A multi-swap proposal is worth a factor of 5.65 in the regime where the walk should not be used at all.

Where the gain is, and where the decision is

A bigger proposal is worth a factor of six at a loose tolerance and nothing at a tight one. The tolerances where it helps are the ones where a hunt costs two evaluations a draw, and the crossing barely moves.

blocks · Assignment
Robust, at the sample sizes it is reached for. Counted coverage of five 95% intervals for a slope, at six sample sizes, under an error variance leaning towards the edges of the design (γ = 0.8), over 20000 draws at the small end. The model-based interval sits at about 87.06% everywhere and does not improve with the sample, because it is a claim about a variance it is not estimating. The robust ones do improve: HC0 covers 88.73% at 20 rows, 92.70% at 50 and 94.93% at 1,000. Its promise is asymptotic and its use is not, and the gap between those two facts is this picture. The leave-one-out correction read against a t on n − 2 is the only line that is near its promise at the small end: 94.55% at 20 rows.

Robust is not free

A robust standard error's promise is asymptotic and its use is not. Its 95% interval covers 88.73% at twenty rows, and under mild heteroskedasticity it is the worse of the two intervals until a hundred.

sandwich · Misspecification
Twenty cells of an interval that is exactly 95%, 1,000 replications each. The t interval covers exactly 95% in every cell. Estimated at 1,000 replications its cells read 93.9% to 96.5%, and 2 of the twenty are flagged by their own ±1.96 standard errors.

A coverage table with its own error

Twenty cells estimating the coverage of an interval that is exactly 95%, at a thousand replications each, read from 93.9% to 96.5% — and a table like that flags at least one of its correct cells on 69.9% of honest runs. Ten times the replications does not repair it: at ten thousand the same table still flags one 63.3% of the time.

method · Seeds
Forty O'Brien–Fleming trials at a true effect of 0.16, with the boundary written as an effect. The dashed line is the smallest effect a trial can report and still stop at each look: 0.510 at 80 observations, 0.255 at 160 observations, 0.170 at 240 observations, 0.128 at 320 observations, 0.102 at 400 observations. The true effect is 0.16, so at 3 of the five looks a trial cannot stop without reporting more than it. 29 of these forty trials stop before the last look, each marked where it stopped.

The effect a stopped trial reports

An O'Brien–Fleming trial at 88.45% power holds its error rate exactly and reports an effect 9.6% too large on average. The 11.39% of trials that stop at the second look report 1.83 times the truth, the ones that cross at the last look report 0.80 times it, and pooling every trial by its size gives the truth back to the last digit.

sequential · Stopping
Estimating a prevalence of 0.10% from 1,000 tests. The positive rate reads 5.09%, which is 50.9 times the truth. The Rogan-Gladen correction averages 0.100% — unbiased — with a standard deviation of 0.820 points against the positive rate's 0.695, and it comes out negative on 48.6% of samples.

The prevalence the test has to estimate

Every predictive value takes a prevalence as given, and the prevalence is usually estimated from the same test's positive rate. At a true prevalence of one in a thousand that rate reads 5.09% — fifty times the truth — and the correction that inverts it is unbiased, 18% more variable, and negative on 48.6% of samples of a thousand.

screening · Baserate
How many observations the smallest and largest of them need. The interval between the extremes of n draws holds at least 95% of the population with probability 1 - n p^(n-1) + (n-1) p^n, whatever the population is. Reaching 95% confidence takes 93 observations.

Ninety-three observations, and nothing assumed

The interval between the smallest and largest of a sample holds a share of the population whose distribution does not depend on the population — Beta(n − 1, 2), for anything continuous. Buying the 95/95 that normality buys at ten observations costs 93 of them, and that number is the exchange rate between an assumption and data.

estimated · Bands
Two studies of the same effect, z statistics with mean 1.96: where one is significant and the other is not. Five hundred pairs. With the true effect identical in both, exactly one of the two is significant in 50.0% of pairs; among those, the difference between the two is significant in 9.7%. The dashed lines mark 1.96 on each axis; the diagonal lines mark a significant difference.

Significant in one, not in the other

Two studies of exactly the same effect, each with 50% power, disagree about significance half the time — and when they do, the test of the difference between them is significant in 9.75% of cases. A p of 0.01 beside a p of 0.20 is a difference with p = 0.36. Among four subgroups sharing one effect, at least one significant and one not happens 87.5% of the time, and the test that would tell a real difference apart needs four times the sample the effect itself needed.

alongside · Repetition
How many of a league table's top ten are small groups, ranked four ways. A hundred groups with sizes from 4 to 400, of which 36% have twenty units or fewer. Small groups make up 36.3% of the true top ten, 62.1% of the top ten by raw means, 13.4% by posterior means and 22.9% by the posterior chance of being in the top ten. The three rankings recover 4.43, 5.38 and 5.47 of the true top ten.

A league table of a hundred

A hundred groups with sizes from 4 to 400, and a top ten to publish. Ranked by their own means, small groups fill 62.1% of the top ten against their 36.3% share of the true top ten. Ranked by posterior means they fill 13.4%. The ranking built from each group's chance of being in the top ten recovers 5.47 of the true ten, the best of three and barely half; and the group ranked first could hold any rank from 1 to 31.

borrowed · Shrinkage
How much of a normal outcome's information survives cutting it into two, by where the cut is. For a small shift, a cut at the mean keeps 63.7% of the information, so the trial needs 1.57 times the sample. A cut at the top tenth keeps 34.2% and needs 2.92 times; a cut two standard deviations out keeps 13.1%.

An outcome cut in two

Replacing a measured outcome with whether it crossed a threshold keeps 63.7% of the information when the cut is at the mean, 34.2% at the top tenth and 13.1% two standard deviations out. A trial that needs 63 patients per arm on the measured outcome needs 102 cut at the mean and 185 cut at one and a half standard deviations. The responder rates that result read as a share of patients who respond — 50.0% against 69.1% — when every patient moved by the same amount; and a cut chosen after looking turns a 5% test into a 17.7% one.

planned · Power
What a schedule is allowed to read, and what happens when it reads more. The construction allows the block sizes to be anything at all as long as they are functions of the within-block contrasts, which are independent of every block mean. A schedule that shrinks the block whenever the between-block spread is running above what the contrasts say is a direct attempt to hold down the quantity the interval will be built from, and it succeeds: the estimate lands at 0.8373σ² against the honest 0.9831, and the coverage goes with it. Reading the running mean instead pushes the other way and over-covers — which is not a repair, it is the same violation with the sign reversed, and the level is no longer a property of the procedure at all.

A schedule that reads the mean

The block sizes may be anything at all provided they are functions of the contrasts. Two natural schedules break that, in opposite directions — and the most natural mistake of the three is not a schedule at all but a stopping rule, at 86.87% coverage and fewer observations.

pace · Stopping
What a pilot buys, σ = 1 against 3. Each point is 6,000 two-stage experiments of 100 units: a pilot of m per arm, then the rest split by the pilot's own estimate of the two spreads. Above the line the pilot has made the experiment worse than not bothering. The best pilot here is 8 per arm at 0.809, against 0.800 for a designer who knew the spreads — so the rule recovers 96% of what knowing them is worth. A larger pilot estimates the ratio better and has less left to apply it to, which is why the curve turns.

Allocating on a guess

Every allocation rule in this field is a function of quantities the experiment is being run to find out. Fed a pilot's estimate of them, the rule that minimises the variance makes the experiment worse than not bothering — until the arms differ by about a factor of two, which is further than anyone would guess.

allocation · Allocation
Power at an effect of 0.5 standard deviations. The curve is the non-central t on 2n − 2 degrees of freedom with δ = d√(n/2); the dots are 4,000 experiments run at each size. Reaching 80% power needs 64 per arm.

How many subjects

Sixty-four per arm for 80% power at half a standard deviation — a power figure that could only be simulated, with nothing to disagree with, until the non-central t was written. Two routes now, agreeing to within the simulation's own error.

design · Power
An AR(1) at φ = 0.5, 200 observations. The bars are the measured correlations; the curve is φᵏ, which is what an AR(1) must have. The band is ±1.96/√n, where an independent series would stay. The first bar is 0.53 against a band of ±0.14.

The check before the standard error

One number decides whether every interval in an analysis is trustworthy, and the check for it flags a lag-one correlation of 0.5 nine times in ten — and one of 0.2 only one time in five, where the interval already covers 88.6% instead of 95%.

timeseries · Dependence
Adding a run can make the design worse. The D-efficiency of the best N-run design at each size, against the optimal measure. It is not a rising curve. 13 runs reaches 99.77% and 14 falls to 99.44%, because the optimal weights are real numbers and N runs is an integer approximation to them, so how good a design can be depends on how well N divides. A Wald interval behaves the same way: a larger sample sometimes makes its coverage worse, for exactly this reason.

The design that has to be integers

The optimal design is a set of real weights and an experiment is a set of runs, so the theory's answer is never available. Thirteen runs reach 99.77% of it and fourteen reach 99.44% — adding a run makes the design worse per run, and the search that finds it does not always find the same one.

optimality · Criterion
Expected width against coverage, n = 30, p = 0.15. The Wald interval is the shortest and covers 94.2%. Clopper–Pearson covers 98.3% and is 13% wider. Shortness is not a virtue on its own — an interval of zero width is the shortest of all.

The shortest interval is the one that misses

Four intervals for the same data, with their widths and their coverage measured together. The narrowest is the one that fails its stated level, which is exactly why it looks the most appealing.

intervals · Width
Twenty residual plots from data where the model is exactly right, n = 24. Every panel is a correctly specified linear model with normal errors. The apparent curvature, funnelling and outliers are all produced by noise, and the largest single residual across the twenty is 2.13 standard deviations of the error. This is the reference nobody has when judging a real residual plot.

Twenty residual plots

Judging whether a residual plot looks wrong requires knowing what a correct one looks like, and almost nobody has seen twenty of those. Here they are, from a model that is exactly right, at the sample size that matters.

regression · Qq
What 20 clusters of 20 correlated observations do to a 95% interval. Each study has 400 observations arranged as 20 clusters of 20. The lower points are the counted coverage of the usual interval, which treats them as 400 independent observations; the curve through them is 2Φ(1.96/√deff) − 1 with deff = 1 + 19ρ, computed before any data was drawn. At ρ = 0.81 the interval covers 36% rather than 95%. The upper points treat the cluster as the unit and need no variance components at all.

Two levels at once

A third level of grouping adds no new arithmetic and produces one number — the design effect — that decides how many independent observations a clustered study is worth. It is the same quantity the time-series field computes for autocorrelated data, arrived at from a completely different picture.

multilevel · Levels
Free until the sums stop seeing what the differences see. Coverage with and without the block sums pooled into the interval's variance estimate. With one effect and one level they are free. With an effect that varies between blocks they are still free, because a block's sum picks that variation up exactly as its difference does. With a level that varies they make the interval 37% wider and conservative. And where the effect falls as the level rises — a ceiling, and not an exotic thing to suppose — the sums carry none of the between-block variation while the differences carry all of it, the pooled estimate is short, and the interval that uses it covers 88.75% on a width 20% narrower than the honest one.

What a two-arm rule may not pool

A spread computed "within the block" without the arm label carries a share of the effect, so the trial runs 173 observations at a null and 282 at an effect of 1.5. The stopping rule is reading the thing it exists to measure, and the phrase that produced it is one word long.

contrast · Allocation
What the exactness costs, and the dial it is bought with. The median half-width of the interval each rule reports, at a requirement of 0.4 and a first look after 5 observations. The flat line is the interval a practitioner writes at the purely sequential rule's stopping time, which covers 91.72% rather than 95%. The curve is the blinded rule, which covers its nominal level at every block size: it reads b − 1 degrees of freedom where the other reads n − 1, and pays for the exactness in width. The best block size is 3, at 0.4712. Larger blocks give the stopping rule a better estimate and the interval a worse one, and the two costs go opposite ways, which is what puts the minimum in the middle.

What the blindfold costs

The exactly-covering rule pays for it in the width of the interval, and the block size is a dial between two costs that run in opposite directions. And on an interval whose width was fixed in advance, the same repair buys nothing at all.

blind · Nuisance
Which samples Wilson and Clopper–Pearson each cover, n = 50, p = 0.2. Each bar is the probability of one count, shaded by which interval built on that count contains 0.2. Both cover 95.1% of samples, only Wilson 0.0%, only Clopper–Pearson 1.6%, neither 3.3%. The correlation between their hits is 0.810, so on shared draws the variance of their difference is 4.891 times smaller than on independent ones.

The same draws for both methods

Two intervals computed on the same simulated datasets give a difference in coverage whose variance can be 4.891 times smaller than on separate datasets — or, for a pair that covers different samples, 1.164 times larger. Which one a comparison gets is an exact sum over the counts each interval covers, and a standard error that ignores the sharing covers 100.00% for one pair and 93.07% for the other.

method · Seeds
Four 95% intervals for the mean of 30 exponential observations, split by the side they miss on. Both bars should read 2.5%. The t interval misses below the mean on 6.38% of samples. Widened until its total is exactly 5%, it misses below on 4.69% and above on 0.30%. Hall's transformation misses below on 3.31% and above on 1.96%.

The side a bound is read from

On thirty exponential observations the upper limit of a 95% t interval is exceeded by the true mean 6.38% of the time, against the 2.5% a safety margin set from it assumes. Widen the interval until its total coverage is exactly 95% and the upper limit is still exceeded 4.69% of the time. A symmetric repair fixes the number that is reported and not the one that is used; Hall's transformation, which bends the interval, takes the same rate to 3.31%.

tails · Student
The factor that holds all of the next m observations with 95% probability, from a sample of 10. From 10 observations, the band for one future value has factor 2.371, for ten 3.716, for a hundred 4.942 and for a thousand 6.008. The 95%-content tolerance factor is 3.382 and is passed by m = 10; the Bonferroni stretch of the prediction factor reaches 7.567 at a thousand.

All of the next ten

A warranty, a batch release or a monitoring rule promises something about every one of the next ten observations, not about one. From a sample of ten, the band that holds all ten with 95% probability reaches 3.716 sample standard deviations either side of the mean — already wider than the 3.382 of a tolerance interval for 95% of the population — and it keeps widening: 4.942 for a hundred, 6.008 for a thousand, with no ceiling. A 95% prediction interval, read as the answer, holds all ten 67.9% of the time: more than 0.95 to the tenth power, because the ten succeed and fail together.

estimated · Bands
Forty trials at a true effect of 0.16, under the rule "power at the trend < 10%". The upper line is the benefit boundary (4.56, 3.23, 2.63, 2.28, 2.04); the lower line is where the rule stops a trial for futility (0.40 at 80, 0.66 at 160, 0.95 at 240, 1.31 at 320). Of forty trials with a real effect, 29 cross for benefit and 11 are stopped for futility.

A boundary for giving up

Adding "stop if z is below zero" to an O'Brien–Fleming trial costs 5.20 points of power at the effect it was designed for and halves the observations a trial with no effect uses. Stopping when conditional power at the observed trend falls under 10% costs 13.23 points and stops 21.28% of trials with a real effect. Making that rule binding lowers the benefit boundary from 2.040 to 1.901, and a binding rule that is then ignored rejects a true null 3.523% of the time instead of 2.5%.

sequential · Stopping
What each estimator costs, against how many groups there are. Each point is 4,000 datasets, with the population spread estimated from the data rather than supplied. Partial pooling first beats BOTH of the estimators it sits between at 5 groups; below that, complete pooling — which estimates nothing at all — is the better answer. Its own cost falls from 1.751 at 2 groups to 0.795 at 40.

The fewest groups that can borrow

At three groups the estimator that shrinks towards its own data's mean returns the group means untouched, on every dataset, because its constant is J − 3. At two it expands instead of shrinking. And the number of groups at which partial pooling starts to be worth doing is five, or two, or never — it depends on how far apart the groups are.

hierarchical · Pooling
What the guess is worth, when it is worth anything. The variance cost of an even split relative to the variance-minimising one for a risk difference, against the first arm's proportion, with the second at 0.3. The cost is a pure number: it does not depend on the trial's size. It is exactly zero at 0.3 and at 0.70, where the two arms have the same p(1 − p); it is 0.19% at a half and 4.36% at a tenth. Across the whole range from a tenth to nine tenths it never exceeds 4.36%, which is what the variance-minimising rule is worth here — and what it is worth is the reason it is safe to use with a guess.

The arm whose variance is its answer

With a binary outcome the allocation rule is a function of the proportions the trial exists to estimate. It costs at most 4.36% of variance to ignore it anywhere between a tenth and nine tenths, because √(p(1−p)) stays within a factor of two of its peak across 98% of the unit interval.

allocation · Allocation
Student's t on 5 degrees of freedom, against the normal. The two-sided 95% critical value is 2.571 for t(5) and 1.960 for the normal — 31% wider. Using the normal at this sample size makes every interval too short by that much.

The correction for not knowing the spread

The t distribution exists because the standard deviation is estimated rather than known. At eight observations, using the normal instead makes every interval 12% too short — and the coverage that follows can be measured rather than argued about.

intervals · Student
What each group gains from being pooled, τ = 1. Eight groups whose sizes span a factor of 13.3. The smallest gains 2.076 of squared error, which is 36.8% of the total reduction; the largest gains 0.030, which is 0.5%. 2 of the eight account for half of everything pooling buys.

Where the borrowing goes

Pooling cuts the total squared error across eight groups by 56%. Two of the eight take 61% of that reduction, the four best-measured groups share 11% between them, and the largest group gets 1.5% of what the smallest does. The headline is a fact about the groups nobody was asking about.

hierarchical · Pooling
What a variance estimated from K units is worth. The between-unit mean square is a scaled chi-square on K − 1 degrees of freedom, so the estimator's whole distribution is decided by the number of units. At two units its interquartile range spans a factor of 13.03 and its ten-to-ninety range a factor of 171.3, and it comes out exactly zero on 26.7% of studies. The closed form and 3,000 simulated studies agree to 0.051 at every quantile.

A level with two units

A variance estimated from two units is a scaled chi-square on one degree of freedom. Its interquartile range spans a factor of thirteen, its ten-to-ninety range a factor of a hundred and seventy-one, and it comes out exactly zero on 26.7% of studies — so the design effect it decides runs from 1.00 to 7.01 against a truth of 4.69.

multilevel · Levels
A prior worth 35 observations, moved across the range — truth 0.1, n = 20. The same prior weight centred at each of 33 places. Its interval covers 100.0% where the centre is near the truth and 0.0% at its worst, while the mean width where it covers least is 0.221 against a flat prior's 0.263 on the same data.

When the prior is confident and wrong

A prior worth thirty-five observations, centred in the wrong place, produces a 95% interval that covers nothing at all — and reports a width 5% narrower than an honest one. It takes seventeen thousand observations to repair, not thirty-five, and the worst study to run is the one whose sample size equals the prior's weight, exactly.

bayes · Credible

Named alongside it

The objects these essays reach for when they reach for this one.

CoverageStatistical powerMonte CarloDegrees of freedomStandard deviationConfidence intervalStopping ruleEstimated varianceExperimental designFixed-width intervalBlindingBlocking

All concepts