Every essay
Grouped by field. Each heading is one of 93 fields, which is the coarsest of the five ways through this collection — the others are series, threads, figures and concepts.
The distribution itself 7The tail past the last observation 7Intervals, counted 7Tests, and the second number 7Reversals that are not errors 7The prior, doing visible work 7Regression, and what the summary hides 7A standard error for a model that is wrong 7A variable that moves one thing only 7What conditioning on a variable does 7Corrections, and what each controls 7When the data stops early 7The value that is not there 7Stopping rules 6Groups that borrow 7Decided before the data 7Weighting one sample into another 7When the observations repeat each other 7The spread, and its own uncertainty 4Hierarchy past one number 7Series that move together 4Three series, and a count 7The surface between the corners 7A design chosen rather than looked up 7Splitting the units 6Designs that change while they run 4The reference distribution the design supplies 6Coverage without a distribution 7The observation that has not happened 4The criterion, and what it assumes 4Balancing on what was recorded first 4Comparing two forecasters 7A forecast that is a probability 7A design that assumes less 4More arms than two 4The best of a set, and what the search costs 4What the design is asked to guarantee 4Balancing what has no levels 4Searching among fitted models 4What the procedure may not read 4The shape the covariate enters by 4A search with no fixed point 4Choosing what the rule reads 4The block size as a schedule 4Scoring a search without spending data 4When the set is too large to walk 4A promise about two arms 4Counting what is independent 5When the two are not independent 4The weights the corner needs 3Estimating the dependence, not naming it 4A cut point, at a correlation 3What a block may vary 5The shape a dependence has 4A block, weighted inside itself 2What a dictionary buys and what it costs 4When a fixed width is reached 4Fitted together, or fitted after 4Paying for a search 4What a chain cannot report 4A guarantee that needed a symmetry 4Where a taper's case begins 4A covariance with no parameter 4How long the list is 3Two searches over one sample 3The diagnostic after the trial 4The other half of the dependence 3A block length chosen from the data 3A charge for a covariance's own dimension 3The rate and the size of a disagreement 3Two searches over different features 3A probe chosen rather than picked 3Both halves of the dependence at once 4The block length read on a quantile 4A charge that is not a straight line 6What decides whether a tuning list decides 4Overlap and complementarity, separated 3A probe from what the rule blocks 5The same table at seven correlations 4The interval, studentised 5The interval that holds observations, not a mean 4When the stratified answer and the pooled one disagree 3The analyses that were available and not run 3What a summary of a scatter is a property of 3Shape, and what it does to a two-sample test 4Two tests, a threshold, and the rate they are read against 3What a diagnostic plot is showing 3Past the first term of the normal approximation 3A proportion's interval near the boundary, and the coin 3An interval read beside something else 3What partial pooling does to one group, to the set, and to a ranking 3What a sample-size calculation was given 3What makes it checkable 7
The distribution itself
Sums converge on it, which is the most-quoted theorem in the subject. What the quoting leaves out is the rate: the middle converges quickly and the tail does not, and the tail is where the approximation is actually read. The same theorem carried through a function stops being a normal limit where the function is flat, and carried through a ratio it forbids every interval that is always finite.
Sums of almost anything
The theorem says sums converge on one shape whatever they are sums of, which is remarkable and true. Watching it happen from a one-sided skewed source, with the rate of convergence predicted in advance, is more convincing than watching the shape appear.
The tail converges last
The central limit theorem is usually shown as a shape arriving. What the demonstration leaves out is the rate — and the rate is wildly different in the middle and in the tail, which is where every approximation in the subject is actually read.
What normal actually looks like
A single quantile plot of forty normal points wanders enough to look suspicious. Twenty of them, all genuinely normal, show what the noise looks like — and any single panel a reader would have rejected is in there.
The shape, and where its mass is
68, 95, 99.7 is recited more often than any other set of numbers in the subject. They are integrals of a specific curve, they are worth computing rather than remembering, and the third one is the one people misuse.
Where the derivative is zero
The delta method reads a standard error off a tangent line, and at a flat point the tangent says the spread is zero. The interval built on it for a squared mean covers 99.991% there and 85.978% one and a half standard errors away, with nearly every miss on the same side — and the law it should have used is a χ², not a normal.
A ratio whose interval has to be the whole line
The delta interval for a ratio of two means covers 95.61% when the denominator is eight standard errors from zero and 1.10% at a ten-thousandth of one, and ten times as wide it still covers only 3.48%. Linearising is not the fault. Gleser and Hwang proved that every interval that is always finite fails the same way, so an interval that keeps its promise has to be the whole line some of the time.
A flat point with more than one direction
At a stationary point of a function of several means the second-order law is ½ Z′HZ, so the bias is half the Hessian's trace — 2.008 for a bowl, 5.028 for a valley, and −0.006 for a saddle, where the eigenvalues cancel. The saddle's coverage is the worst of the three.
The tail past the last observation
The distribution a sum converges on has one limit; the distribution a maximum converges on has three, and every extreme-value analysis is a claim about which. The claim is also an extrapolation: a hundred-block level read from fifty blocks sits above the largest reading in the record half the time. So each measurement here reports not what the estimate is but how much of it came from data.
Three shapes, one limit
A normalised sum has one limit and a normalised maximum has three, indexed by a single number. Twenty blocks put the sign of that number right 97.3% of the time — and naming the family from a light-tailed record gets worse as the record grows, from 83.0% at twenty blocks to 4.8% at five hundred.
The maximum converges slowly
The rate at which a normalised maximum reaches its limit law is computable rather than simulable, because the exact law of a maximum is always available. For a normal parent the distance falls like one over the logarithm of the block and is still 0.0091 at a million readings; for an exponential parent, with the same limit, it is 2.707×10⁻⁷.
The threshold is a dial
A peaks-over-threshold analysis has one knob, and raising it buys accuracy with exceedances. For a normal parent the error is smallest at the 0.925 quantile and 80.6% of it is still bias there — and both diagnostics practitioners use to set the knob lose to a fixed 0.90 rule, one by a factor of 1.590 and one by 11.881.
A level with no data in it
The largest of fifty block maxima is a 51-block event by its own plotting position, so a hundred-block level is read 1.96 times past the longest event the record contains — and it lands above the largest reading on 52.4% of records. The estimate stays nearly unbiased out there; what grows is its error, sixfold from ten blocks to a thousand.
Two intervals for one return level
Two 95% intervals read off the same fits of the same records, against a level known in closed form. The symmetric one covers 80.3% at twenty-five blocks and reaches only 89.0% at two hundred — and 99.24% of its misses are the interval sitting entirely below the truth, which is not the endpoint anybody expects to fail.
The clustering the tail has
Every threshold method counts exceedances as though they were independent pieces of information, and in a dependent series they arrive in clusters. Ignoring that overstates a return level by the reciprocal of the extremal index — ×3.527 counted where the mean cluster holds four — and leaves a reported standard error 2.151 times too small.
The run length a declustering chooses
The runs estimator of an extremal index carries a constant nobody derives. Where a cluster is a run of neighbouring exceedances the constant barely matters; where a cluster's members fall six steps apart, the estimate is 0.9069 at a run length of six and 0.3649 at seven against an index of 0.40, and a run length of four removes under a tenth of the overstatement declustering exists to remove. A rule that reads the run length off the data has the smallest worst error of the three.
Intervals, counted
An interval that claims 95% is making a checkable statement about a procedure. Build every possible sample and count. The interval taught first fails, the failure is worst where proportions are most often reported, and more data does not monotonically help.
What the 95% refers to
An interval that claims 95% is making a checkable statement about a procedure, not about the interval in front of you. Build every possible sample and count, and the interval taught first turns out to cover 87.6% of the time.
More data is not monotonically better
Coverage of an interval for a proportion does not improve smoothly as the sample grows. It oscillates, and there are larger samples that cover materially worse than smaller ones — a sample of twenty covers twelve points worse than a sample of nineteen.
Twenty intervals and one expected miss
The 95% belongs to the procedure, not to the interval in front of you. Twenty intervals from twenty samples make that visible in a way no definition does, and the one that misses is not a mistake.
The shortest interval is the one that misses
Four intervals for the same data, with their widths and their coverage measured together. The narrowest is the one that fails its stated level, which is exactly why it looks the most appealing.
Where the bootstrap lies
Resampling is the most generally useful trick in the subject and it has a failure mode that is easy to state: it cannot see past the data. For a statistic that lives at the edge of the sample, coverage collapses from 95% to almost nothing.
The correction for not knowing the spread
The t distribution exists because the standard deviation is estimated rather than known. At eight observations, using the normal instead makes every interval 12% too short — and the coverage that follows can be measured rather than argued about.
An interval that covers and says nothing
A procedure returning the whole line 95% of the time and the empty set otherwise has coverage exactly 95% at every parameter value. Two real intervals at forty observations have expected widths of 0.2418 and 0.2417 and worst-case coverages of 55.31% and 92.21%.
Tests, and the second number
A p-value alone cannot be read: the same 0.04 means different things at different sample sizes, and nothing at all without knowing how many analyses were available. Under a real effect it is a draw from a distribution several orders of magnitude wide, a replication of a result at 0.05 succeeds exactly half the time under any model centred on that result, and the standard ways of combining several p-values disagree about which evidence counts. Every figure here carries the number that makes it interpretable.
A p-value that is not flat is not a p-value
Under a true null, p-values are uniform. That is stronger than saying the test rejects 5% of the time, it constrains the whole distribution rather than one point of it, and it catches implementation errors that a rejection rate sails past.
What a p-value does not say
The same p of 0.04 corresponds to a large effect in ten observations and a negligible one in two thousand. A p-value alone cannot be interpreted, and the number that makes it interpretable is almost never printed beside it.
The winner's curse
Filter honest studies down to the ones that reached significance and the effects they report are systematically too large. At low power the inflation is a factor of two, nobody has done anything wrong, and the selection did all of it.