The collection

Every essay — page 2

Essays 25 to 48 of 436, in the same order.

Tests, and the second number

A p-value alone cannot be read: the same 0.04 means different things at different sample sizes, and nothing at all without knowing how many analyses were available. Under a real effect it is a draw from a distribution several orders of magnitude wide, a replication of a result at 0.05 succeeds exactly half the time under any model centred on that result, and the standard ways of combining several p-values disagree about which evidence counts. Every figure here carries the number that makes it interpretable.

The false-positive rate against the number of analyses, on pure noise. The data has no effect in it. Each analysis is correct and each p-value is honest. With 20 correlated outcomes available, something reaches p below 0.05 58% of the time.

Twenty analyses of nothing

Twenty honest, correct analyses of data with no effect in it find something significant 57% of the time. Nobody p-hacked, every individual p-value is right, and the reported one is the smallest of twenty.

7 figures · Forking, part 3
The p-value of a study with 80% power, twenty thousand times. Twenty thousand two-sided z-tests, each on 25 observations whose true mean is 0.5603 standard deviations from the null, a noncentrality of 2.802. The bars are the counted share of p-values in bins a quarter of a power of ten wide, with the leftmost bin holding everything smaller; the line is the closed form. The middle eighty per cent of the p-values runs from 4.4×10⁻⁵ to 0.13, 3.46 orders of magnitude, the median is 0.0051, and 80.0% fall below 0.05, which is what the power means.

The p-value a replication gets

Under a true null a p-value is flat. Under a real effect its distribution is closed form and wide — a study with 80% power returns anything from 4.4×10⁻⁵ to 0.13 in eight runs of ten — and the chance that an exact replication of a p = 0.05 result is significant again is exactly one half, under both of the models people use without naming them.

8 figures · Uniformity, part 2
Three ways to reject with two studies, drawn where the two z statistics live. Two one-sided studies, each summarised by its z statistic. Fisher's combination rejects outside a curve that runs parallel to both axes, so one study past z = 2.378 decides it alone; Stouffer's rejects above the straight line z₁ + z₂ = 2.326; Tippett's rejects when either z passes 1.955. Each region holds exactly 5% of the standard bivariate normal — Fisher's in closed form, e^(−c/2)(1 + c/2) at c = 9.488 — and of 100,000 counted null pairs they catch 4.95%, 4.88% and 5.04%. Two alternatives carry the same Stouffer evidence: one study at 2.326 and the other at nothing, where the powers are 62.7%, 50.0% and 65.4%; and both at 1.163, where they are 47.7%, 50.0% and 38.3%.

Two ways to combine p-values

Fisher's and Stouffer's combinations are both exactly right when every null is true, for the single reason that each p-value is flat. Under a real effect they disagree about which evidence counts: with Stouffer held at 50% power across ten studies, Fisher is the more powerful while the signal sits in six or fewer of them and the less powerful from seven.

7 figures · Uniformity, part 3
How often each combination, and each union of them, rejects ten studies of nothing. Each combination alone rejects exactly 5% of null sets. Counted on 1,000,000 sets of ten null studies: Fisher or Stouffer 6.63%, Fisher or Tippett 8.05%, Stouffer or Tippett 8.96%, any of the three 9.66% — enclosed on a two-dimensional lattice between 9.18% and 10.09% — and all three together 1.05%. The three sizes add to 15%.

The smallest of three combinations

Reporting whichever of Fisher's, Stouffer's and Tippett's combinations is smallest is a test of its own, and on ten studies of nothing it rejects 9.66% of the time — not 5%, and nowhere near the 15% the three sizes add to, because the statistics are correlated at up to 0.903. Read at 2.448% each it is exact, and then it trails the best single combination by at most 7.45 points and leads the worst by at least 10.30.

6 figures · Uniformity, part 4

Reversals that are not errors

Simpson's reversal, the base rate, regression to the mean. Each is normally taught with one famous table. A table is a point; these are swept, so how much of the space behaves that way and how large it can get both have answers — including why two correct analyses of one baseline disagree, how far a group enrolled on a noisy reading falls with nothing done to it, and why the correlation predicts the regression of the extremes only when the population is normal.

Which allocations reverse the overall comparison. The per-group success rates are held fixed; only the split of each group between treatment and control changes. 32% of the allocations reverse, and the worst reverses by 13.1 percentage points.

Simpson's reversal is a region, not a table

The treatment wins in both groups and loses overall. That is normally shown with one famous table, which cannot answer the two questions a reader has — how often, and how large. Swept, it turns out to occupy 31% of the allocation space.

7 figures · Simpson, part 1
What a positive test means, sensitivity 90%, specificity 95%. At a prevalence of one in a thousand, 98 of every hundred positives are false. At one in 10, 33 are. The test has not changed.

What a positive test is worth

A test that is 90% sensitive and 95% specific sounds accurate. For a condition affecting one person in a thousand, 98% of its positive results are wrong, and a worse test on a commoner condition beats a better test on a rare one.

7 figures · Baserate, part 1
Two measurements of the same thing, correlated 0.60. Pick the worst 15% on the first measurement and their average rises by 0.78 on the second. Pick the best and theirs falls by 0.48. No treatment was given to anybody.

Regression to the mean

Select the worst performers, measure them again, and they improve. Select the best and they decline. No intervention is required for either, the size of the apparent effect is predictable from the correlation alone, and it is the reason so many things appear to work.

4 figures · Rtm, part 2
Two groups, a baseline and a follow-up, and nothing happening in between — baseline reliability 0.6. 600 units in two pre-existing groups whose true means are 1.00 apart, read once at baseline and once at follow-up, with no change for anybody. The two groups' mean changes are −0.075 and −0.032, so the change-score analysis reports a group difference of 0.043. The regression of follow-up on baseline and group reports 0.409, against a closed form of (1 − λ) × 1.00 = 0.400: at any one baseline reading the two groups' lines sit that far apart, because each group's units regress towards their own group's mean. The pooled slope in this sample is 0.614, the baseline's reliability.

Two analyses of one baseline

Two groups read at baseline and again at follow-up, with no change for anybody. Subtracting the baseline reports a group difference of −0.0014 and adjusting for it reports 0.4008 — and each analysis is exactly right about one reason the groups started apart and wrong by 0.40 about the other.

6 figures · Rtm, part 3
A cohort screened once, the top tenth enrolled, and followed up with nothing given — correlation 0.6. 2000 people read once at screening and once at follow-up, with a test–retest correlation of 0.6 and no treatment. The 215 above a cut at the top ten per cent of one reading (1.282 standard deviations) are enrolled. Their mean screening reading is 1.744 and their mean follow-up reading 0.982, a fall of 0.762 ± 0.055 with nothing done to anyone. The closed form for the fall is (1 − ρ) times the truncated-normal mean, (1 − 0.6) × 1.755 = 0.702.

The measurement that got them enrolled

Enrol the top tenth of one screening reading and give them nothing, and they fall by 0.702 standard deviations at follow-up. Measured from a fresh reading taken after enrolment they fall by nothing. Averaging ten screening readings still leaves 0.101, and it takes twenty-one to get under 0.05.

6 figures · Rtm, part 4
The top of a heavy-tailed population keeps its lead; the top of a light-tailed one gives it back. Select the top share on the first reading and read the group again: the share of its mean lead the second reading keeps, by integration over the true score (lines) and counted on 400,000 draws a parent in 20 batches (points, with two standard errors). The normal keeps exactly 0.6 at every selection. At the top half the Laplace keeps 0.541, the t 0.535 and the uniform 0.648 — the heavy tails keep LESS than the correlation. By the top one per cent the order has reversed: 0.761, 0.784 and 0.443. At one in ten thousand the t keeps 0.977 and the uniform 0.346.

A lead that a heavy tail keeps

Four populations whose readings all correlate at exactly 0.6, and whose least-squares slopes all read 0.6. Select the top one per cent on one reading and measure them again: they keep 60% of their lead if the true scores are normal, 76.1% if they are Laplace, 78.4% if they are a t on four degrees of freedom — and 44.3% if they are uniform. The correlation predicts the regression of the extremes for one shape of population only.

6 figures · Rtm, part 5
Twenty studies of a thousand readings estimate the regression of the extremes: the t, four degrees parent. Each thin line is one study of 1000 readings from the t, four degrees parent: Tweedie's formula with the log-density's slope estimated by a degree-5 log-spline, drawn up to that study's largest reading. The thick line is the share kept by integration over the true score, and the dashed line the correlation, 0.6. The top ten readings of the median study begin at 2.45. Over 400 studies the corrected share kept by the top one per cent averages 0.8190, with a spread of 0.0806, against 0.7842 by integration; a Gaussian kernel averages 0.7914 with a spread of 0.0853.

The slope of a density nobody can see

Tweedie's formula corrects a reading by the slope of the readings' own log-density, and a study has its readings. Estimated from a thousand of them, the correction for the top one per cent beats the correlation's linear rule on 84.0% to 98.0% of studies from heavy-tailed populations and on 75.0% to 81.5% from a bounded one — and costs an error of 0.09 to 0.12 where the population is normal and the rule was already exact. At 250 readings the log-spline loses to the rule it replaces, and at 16,000 the same log-spline gets worse on a power tail.

6 figures · Rtm, part 6

The prior, doing visible work

A credible interval says the thing everyone wants a confidence interval to say, and it needs a prior to do it. So the prior is treated as a component with a measurable effect: it is worth a stated number of observations, and the interval it produces has a coverage that can be summed over the sample space like any other.

weakly informative — Beta(2, 2), updated by 5 of 20. The prior is worth 4 observations. With 20 observations the posterior mean is 0.292, against a data proportion of 0.250 and a prior mean of 0.500.

What a prior is worth

A prior is not a philosophical position, it is a component with a stated size. For a proportion it is worth exactly a + b observations, which turns "how much does the prior matter" from an argument into a subtraction.

7 figures · Prior, part 1
What a 95% credible interval covers, n = 20. Computed by summing over all 21 possible counts rather than by simulating them. Jeffreys' prior covers close to 95% across the range; a confident prior centred in the wrong place covers almost nothing where the truth is far from it.

What a credible interval covers

A credible interval makes the statement everyone wants and does not claim to have a coverage. It has one anyway, it can be summed over the sample space exactly, and on a reasonable prior it beats the interval taught first.

7 figures · Credible, part 2
A normal mean with a flat prior — one interval, two readings. Both intervals are [1.878, 3.446]. The frequentist reading is that the procedure captures the truth 95% of the time; the Bayesian reading is that the parameter is in this interval with probability 0.95. The endpoints are identical to machine precision.

Where the two schools agree

With a flat prior on a normal mean, the credible interval and the confidence interval are the same interval, endpoint for endpoint. Knowing exactly when that stops being true is more useful than either camp's general argument.

6 figures · Credible, part 3
What a positive test means, sensitivity 90%, specificity 95%. At a prevalence of one in a thousand, 98 of every hundred positives are false. At one in 10, 33 are. The test has not changed.

The base rate was always Bayes

The screening arithmetic everybody finds counter-intuitive is a posterior update with a prior of one in a thousand. Naming it that way turns a famous puzzle into an instance of a rule, and makes the sequential version obvious.

7 figures · Baserate, part 2
Two 95% intervals for 2 of 20, under Jeffreys — Beta(½, ½). The equal-tailed interval runs from 0.0214 to 0.2839 and is 0.2625 wide; the shortest interval runs from 0.0093 to 0.2540 and is 0.2447 wide. Both hold 95% of the posterior, and the shorter one buys its 6.8% by moving its lower endpoint towards the denser side.

The shortest interval, and the one that does not move

Two 95% intervals come out of every posterior and they are not the same set. The shorter one is shorter by 4.86% on average and 22.41% at its best, it covers 86.72% where the other covers 95.68%, and it is not even the shortest once the parameter is written a different way.

7 figures · Credible, part 5
Four 95% intervals for the odds after 2 of 20. credible, transformed: 0.0218 to 0.3964. Wald, transformed: -0.0305 to 0.3012. delta method on the odds: -0.0512 to 0.2734. delta method on the log-odds: 0.0258 to 0.4789. The first two are the same intervals for the proportion with their endpoints put through the odds; the last two are fresh approximations made on the new scale.

An interval for something else

An interval for the odds is free — put the endpoints through the odds and the coverage does not move, exactly, for any interval at all. The method everyone uses instead computes a new standard error on the new scale, and at twenty trials that costs four points of coverage, produces negative odds, and has no value at all when nothing was observed.

7 figures · Credible, part 6
A prior worth 35 observations, moved across the range — truth 0.1, n = 20. The same prior weight centred at each of 33 places. Its interval covers 100.0% where the centre is near the truth and 0.0% at its worst, while the mean width where it covers least is 0.221 against a flat prior's 0.263 on the same data.

When the prior is confident and wrong

A prior worth thirty-five observations, centred in the wrong place, produces a 95% interval that covers nothing at all — and reports a width 5% narrower than an honest one. It takes seventeen thousand observations to repair, not thirty-five, and the worst study to run is the one whose sample size equals the prior's weight, exactly.

8 figures · Credible, part 7

Regression, and what the summary hides

A slope, a standard error and an R² can all be computed from data the model is grossly wrong about, and none of them says so. Leverage is a number known before the outcome is looked at, influence is a number, and whether a residual plot looks bad is a question with a calibrated answer. Two wrong rows placed together hide from every single-row diagnostic, and a loss that bounds a large residual does nothing about a far one.

Twenty points and one more, at leverage 0.74. Without the distant point the slope is 0.495; with it the slope is -0.389. Its leverage is 0.737 and its Cook's distance is 24.1, against a conventional threshold of 1.

The line that one point drew

A single observation among twenty-one reverses the sign of a fitted relationship. Its leverage is known from its x value before the outcome is looked at, so this is a property of the design rather than a surprise in the data.

6 figures · Leverage, part 1
Four datasets, slope 0.50, R² 0.67. Every one of these fits reports the same slope to two decimals and the same R². Only the first is a linear relationship with noise: the second is a curve, the third is a line with one outlier, and the fourth has its slope set by a single point.

Four datasets, one summary

Four datasets agree on slope, intercept and R² to two decimals. One is a linear relationship, one is a curve, one is a line with an outlier, and one has its slope set by a single point. The summary cannot tell them apart and neither can any other summary.

6 figures · Summary, part 1
R² against the number of useless predictors, n = 30. The response is pure noise and so is every predictor, so the true relationship is nothing at all. R² rises from 0.000 to 0.648 anyway, following k/(n − 1) — which is what a criterion that rewards higher R² is actually rewarding.

R² is not a measure of fit

Adding a predictor with no relationship to anything cannot reduce R², and in expectation raises it by 1/(n − 1). Twenty useless predictors on thirty points give an R² of 0.69 from pure noise.

6 figures · Summary, part 2
Twenty residual plots from data where the model is exactly right, n = 24. Every panel is a correctly specified linear model with normal errors. The apparent curvature, funnelling and outliers are all produced by noise, and the largest single residual across the twenty is 2.13 standard deviations of the error. This is the reference nobody has when judging a real residual plot.

Twenty residual plots

Judging whether a residual plot looks wrong requires knowing what a correct one looks like, and almost nobody has seen twenty of those. Here they are, from a model that is exactly right, at the sample size that matters.

6 figures · Qq, part 3
Two far rows, and the line with one of them deleted. Twenty clean points and two rows near x = 9. The slope is −0.511 with every row, −0.376 with one far row deleted, and 0.495 with both deleted. Deleting one of them barely moves the line, because the other is still there.

Two points that hide each other

One far observation among twenty-one has a Cook's distance of 24.1. Put a second beside it and the two read 0.966 and 0.772, neither crossing 1, while together they reverse the slope and deleting both moves the fit by 53.3.

7 figures · Leverage, part 2
One row at x = 9 on a wrong line, fitted three ways. Least squares gives slope −0.389, Huber 0.171 with the extra row at weight 0.115, and least trimmed squares 0.420, fitted to the 12 rows it keeps. The twenty clean rows alone give 0.495. Open circles are the rows the trimmed fit leaves out.

A robust loss and a far x

One far row drags least squares to a slope of −0.389. Huber's loss, the standard robust line, reaches only 0.171, and carried further out the same row gets its full weight back. Least trimmed squares reads 0.420 at every distance, and at the normal model keeps 7.13% of least squares' efficiency to do it.

9 figures · Leverage, part 3

FieldsThreadsSeriesConceptsFigure librarySearch