Concept

Correlation — where it appears

Covariance divided by the two standard deviations, so it lies between −1 and 1 and does not depend on the units. Two quantities can be strongly correlated and still leave a variance estimate built on one independent of a mean built on the other, because a quadratic and a linear form have no third moment to couple them.

Named by 33 essays across 20 fields — each of them below, with the objects they name alongside it.

The expansion that never terminates. The Hermite coefficients of a median split, in magnitude, against the reference j to the power −3/4, anchored at the first one. Every even order is exactly zero because sign is an odd function, and every odd order is not, so no truncation is exact — where a polynomial of degree d is exact at any order past d. Summed, the tail past J falls like 1/√J: sixty orders still leave 6.6% of the variance outside. That statement is what made a cut dictionary's geometry unavailable in closed form, and it is a statement about the function against itself. What it is not is the accuracy of an inner product between two correlated variables, where every term past J carries a factor of ρ^m as well.

A cut is not a polynomial, and it does not have to be

A threshold's expansion never terminates, which is why a balancing dictionary's geometry was closed for powers and taken to draws for cut points. Conditioning on the second variable closes it for both.

splits · Blocking
Two covariates make the dictionary an outer product. Four functions of each covariate, and everything a balancing rule may be handed. The margins are the 8 main effects and the block between them is the 16 interactions, which are 66.7% of the dictionary. Every inner product in it is closed form — ⟨f₁g₁, f₂g₂⟩ = ⟨f₁,f₂⟩⟨g₁,g₂⟩ when the covariates are independent — so nothing about the geometry gets harder. What gets harder is the counting: choosing k of 24 is C(24, k), which is 10,626 at four and 735,471 at eight.

A dictionary that is a product

Two covariates make what a balancing rule may read an outer product — eight main effects and sixteen interactions — and every inner product in it is still closed form. What a rule holding all eight main effects removes of a pure interaction is not small. It is zero.

product · Blocking
The rule is parity, and it runs both ways. At a correlation of 0.5, four combinations of a dictionary and an outcome shape. The joint sign flip (X, Y) → (−X, −Y) leaves the bivariate normal alone at every correlation, so a function that changes sign under it is orthogonal to one that does not. A product of two odd functions is even; a product of an odd and an even one is odd. So an odd dictionary removes exactly none of the first and something of the second, and an even dictionary does the reverse — which it does, to machine precision, in both of the two rows that should be zero. This is one rule where there had been two: that a median split's square is constant, and that a polynomial dictionary contains the products a correlation generates.

A dictionary that is neither

A rule handed two median splits removes none of their interaction; a rule handed two covariates removes none of their product. Those were two results with two explanations, and they are one result with one — and finding it corrected the number underneath both.

dict · Criterion
The fourth-order expectation, by two routes. Every inner product in an eight-term slice of the dictionary at ρ = 0.5, computed from the linearisation and Mehler's formula and counted from two hundred thousand draws of a correlated pair. The entries that matter are the ones off the main effects: ⟨f(X)u(Y), g(X)v(Y)⟩ is a fourth-order expectation, which the independent-covariate field could not write down. The worst departure is 1.99 standard errors over 36 pairs, measured in each pair's own error because the entries differ in size by two orders of magnitude.

The fourth moment that was missing

Mehler's formula makes the main effects exact at any correlation and stops there, because the interactions need an expectation of four Hermite functions rather than two. A linearisation turns the four into two, and the whole geometry becomes closed again.

joint · Blocking
Twenty series with a lag-one correlation of 0.8. Every series has a true mean of zero and 60 observations. The marks on the right are the twenty sample means. The variance of that mean is 8.3 times what 60 independent observations would give, so the series is worth about 7 of them.

The observations that repeat each other

Almost every standard error divides by √n, which claims the observations carry independent information. At a lag-one correlation of 0.8 a fifty-point series is worth about six independent observations, and its 95% interval covers 47%.

timeseries · Dependence
What 20 analyses of one dataset are worth. The threshold giving a 5% family-wise error rate, read back as a number of independent analyses. At no correlation it is 20.05; at 0.6 it is 11.37; at 0.95 it is 2.58. Bonferroni divides by 20 throughout.

How many analyses there really were

Bonferroni divides by twenty because twenty analyses were run. Twenty analyses of one dataset are worth 11.37 independent ones at a correlation of 0.6 and 2.58 at 0.95, and the threshold that controls exactly the same error rate is measurable rather than assumed.

paths · Forking
One relationship at five designs, residual spread 1.00. Every panel has the same slope of 1, the same intercept of 0 and the same residual standard deviation of 1.00. Only the range of x differs. R-squared runs from 0.021 to 0.849, and the estimated residual spread is 0.9932 in all five.

R² is a property of the design

One line, one residual spread, five studies that differ only in how far apart they placed their x values. R² runs from 0.021 to 0.849 and the estimated residual spread is 0.993 in every one of them. Nothing about the relationship changed.

spread · Summary
A model of the active set, and the active set. Each of the 14 units of one design, at the share of its exchanges the tolerance box blocks — computed from the design's columns and the tolerance under a uniform position in the box, against counted over all 116 admissible assignments. The diagonal is where the two would agree. Over 192 designs they agree about the ordering of the units at a correlation of 0.8141 ± 0.0112, negative on 0.5% of them, and disagree about the level: 0.8442 counted against 0.8170 modelled, a gap of 0.0272 ± 0.0051. An admissible assignment does not sit uniformly in its box, and this is the size of that.

A model and a count

The share of a unit's exchanges a tolerance box refuses can be modelled from the design or counted over the admissible set. They order the units the same way at a correlation of 0.81 and disagree about the level by 0.027.

blocked · Randomisation
A threshold in the tail is a threshold nothing balances. The share of a coin's imbalance in an indicator 1{x > c} that survives a rule which balances the covariate itself. The smooth curve is 1 − ρ² with ρ = φ(c)/√(p(1−p)), a closed form with no trial in it; the points are counted over 500 trials of 200 units at each threshold. At the median the two agree that about a third survives — the removed share is exactly 2/π — and by two standard deviations 86.9% survives. The closed form is exact in the limit and optimistic by a few points at this many units, because the rule balances the sample's mean rather than the population's.

A threshold in the tail

How much of a threshold's imbalance a balanced covariate removes is a correlation, and the correlation is a closed form. At the median it is exactly 2/π — the same 2/π a median split throws away — and two standard deviations out it is an eighth.

shape · Blocking
The zero was a fact about independence. What a balancing rule handed every main effect of both covariates removes of a pure interaction, as the covariates are allowed to move together. At ρ = 0 it is exactly nothing — at machine precision, at any number of main effects — which is the independent-covariate result and is correct. It is not small anywhere else: the product of the two covariates loses 64.0% of itself by ρ = 0.5, because h₁h₁ = h₀ + √2·h₂ and Mehler pairs h₂ with h₂ at ρ². Four interactions are drawn and none of them keeps the zero.

A zero that was an assumption

A rule handed every main effect of both covariates removes exactly none of a pure interaction. That is true at machine precision, it is a fact about independence, and it dies as the square of the correlation.

joint · Criterion
Sheppard's arcsine, by two routes. Corr(sign X, sign Y) as the covariates' correlation runs from zero to one, drawn twice. One route is a sixty-four-node quadrature of the orthant probability over the correlation — the general construction, which works at any pair of cut points; the other is (2/π) arcsin ρ, which is elementary and works only at the median. They agree to 3.3e-16 at every one of 81 correlations, which is what licenses the quadrature everywhere else. The curve is above the diagonal at small ρ and below it at large: two signs agree with probability ½ + arcsin(ρ)/π, so a correlation of 0.5 gives exactly ⅓ and a correlation of 0.8 gives 0.5903.

The arcsine that closes it, and the error that was overstated

Two median splits of a correlated pair agree with probability ½ + arcsin(ρ)/π, exactly. And the truncation the field was avoiding falls geometrically in the correlation, not algebraically in the order.

splits · Routes
Two independent random walks, 100 steps. Nothing connects these two series: each is generated from its own independent draws. Regressing one on the other gives a slope with t = -10.9, R² = 0.55 and p = 0.0e+0 — a result that would be reported as a finding by any standard output.

Two walks and a finding

Regress one random walk on another, independently generated, and the slope is significant 76.7% of the time with a median R² of 0.17. Nothing connects the two series, nothing in the output says so, and more data makes it worse.

timeseries · Spurious
What the worst case is worth, one function at a time. The smallest share each dictionary removes, over seven outcome shapes, at a correlation of 0.5. A rule balancing the mean of each covariate has a worst case of exactly zero — against the square, and against both products. Adding a median split to it, which is the second thing every trial balances, leaves the worst case at exactly zero, because a median split is odd and so is a mean. Adding the square instead moves it to 6.8%, and the extra functions after that move it to 7.4%. The worst case is decided by which parities the dictionary contains rather than by how many functions are in it.

What the extra function buys

A rule balancing the mean of each covariate has a worst case of exactly zero. Adding the median split — the other thing every trial balances — leaves it at exactly zero, and one square moves it.

dict · Criterion
Two 95% intervals 2.772 standard errors of the difference apart, standard errors in the ratio 1. The intervals are separated, and the test of the difference gives p = 0.0056. Two 95% intervals with equal standard errors just touch at p = 0.0056.

Two intervals that overlap

Two 95% intervals that just touch are read as a difference at the edge of significance. With equal standard errors their difference has p = 0.0056, not 0.05; two intervals can overlap by 58.6% of an arm and still differ at exactly 5%; standard-error bars that just touch mark p = 0.157; and when the two estimates are correlated at 0.8, touching intervals conceal a difference of 6.2 standard errors. Read as a test, non-overlap needs 1.66 times the sample for the same power.

alongside · Repetition
Three copulas that break nothing, and a factor of two between them. The three radially symmetric copulas, at a matched Spearman correlation of 0.40, against the covariate's marginal. All three leave exactly nothing with a symmetric covariate — that is the guarantee, and it holds to twenty decimal places. What they do to a skewed covariate is not the same at all: at a skewness of 2.26 a Frank copula leaves 12.118% where a Gaussian leaves 21.539% and a t on four degrees of freedom leaves 23.640%. A factor of 2.0 between two copulas that are both symmetric, both matched on rank correlation, and both harmless on their own. So the copula matters to the marginal's leak without breaking any symmetry of its own, which is a milder version of the same finding and applies to every trial rather than to the asymmetric ones.

A copula that halves a marginal

Three copulas break nothing on their own and put a factor of two between the same skewed covariate's leaks — 12.118% under a Frank against 23.640% under a t, at the same rank correlation.

compound · Adjustment
3 arms against one control, 360 units in all. Every control size, enumerated. The best is 132 on the control and 76 on each arm — a ratio of 1.74, against √3 = 1.73. Splitting the units evenly over all 4 groups costs 7.2%, which is small; what the larger control also does is lower the correlation between the comparisons, from 0.50 to 0.37, and that changes which multiplicity correction is right.

One control, many arms

The control appears in every comparison, so it is worth √k treatment arms — and the same sharing makes the k tests correlated at n/(n+n₀), which is the quantity Bonferroni ignores. Both facts come out of one design decision, and it is the size of the control.

allocation · Multiplicity
Two measurements of the same thing, correlated 0.60. Pick the worst 15% on the first measurement and their average rises by 0.78 on the second. Pick the best and theirs falls by 0.48. No treatment was given to anybody.

Regression to the mean

Select the worst performers, measure them again, and they improve. Select the best and they decline. No intervention is required for either, the size of the apparent effect is predictable from the correlation alone, and it is the reason so many things appear to work.

paradox · Rtm
Two arms leave one degree of freedom per block unaccounted for. Each point is one run. The one-mean field's identity is (b − 1) + (N − b) = N − 1, and every schedule moves along that line rather than off it. Two arms give the rule N − 2b and the interval b − 1, which come to N − b − 1 — short of the N − 2 two arms leave by exactly one per block, since a block's arm counts absorb one degree of freedom each and only one of the two directions carries the difference. The hollow points add what the block sums are worth, b − 1 more, and land on the total. The missing degrees of freedom are not lost; they are in a place the interval has to be shown it may read.

The degrees of freedom in the sums

One arm partitions N − 1 exactly. Two arms give the rule N − 2b and the interval b − 1, which is short by one per block — and the missing ones are in the block sums, which are correlated with the differences at −0.79 and are usable anyway.

contrast · Blocking
The guarantee that survives a correlation, and the one that does not. What a balancing rule handed both main effects removes of the pure interaction between them, as the covariates become dependent. For median splits it is exactly zero at every correlation, because sign(x)² = 1: the interaction sign(X)sign(Y) is orthogonal to sign(X) and to sign(Y) whatever ρ is. For the product of the raw covariates it is 4ρ²/(1+ρ²)² — 64.00% by ρ = 0.5, rising to all of it at perfect correlation. A cut away from the median sits between them and is not small: 23.01% at a cut of one. The zero is not a fact about interactions. It is a fact about a dictionary whose functions square to a constant, which a polynomial one does not.

The zero that survives a cut

A rule holding both main effects removes half of a pure interaction between correlated powers and exactly none between correlated median splits. The guarantee that a correlation destroyed was never about interactions.

splits · Criterion
What differencing fixes, and what it costs, 100 steps. The first pair is the false-positive rate for two independent random walks: 77% on the levels, 4.9% on the differences. The second pair is how much of a real relationship survives: R² falls from 0.91 to 0.33. The same operation does both.

What differencing costs

Differencing takes the false-positive rate between two unrelated walks from 76.7% to 4.9%, and takes a genuine relationship's R² from 0.91 to 0.33. Applied to a series that did not need it, it doubles the variance and installs a correlation of −0.5 that the data never had.

timeseries · Spurious
How far apart four summaries put the quartet. Each bar is the largest value minus the smallest across the four datasets. Pearson spans 0.0018, which is the construction working. Distance correlation — the measure that is zero if and only if the variables are independent — spans 0.101, and Spearman spans 0.491.

The summary that was meant to work

Distance correlation is zero if and only if two variables are independent, which is exactly the guarantee a correlation coefficient lacks. Run on the four datasets that share a correlation, it spreads them by 0.10 — and Spearman, which guarantees nothing, spreads them by 0.49.

spread · Summary
An AR(1) at φ = 0.5, 200 observations. The bars are the measured correlations; the curve is φᵏ, which is what an AR(1) must have. The band is ±1.96/√n, where an independent series would stay. The first bar is 0.53 against a band of ±0.14.

The check before the standard error

One number decides whether every interval in an analysis is trustworthy, and the check for it flags a lag-one correlation of 0.5 nine times in ten — and one of 0.2 only one time in five, where the interval already covers 88.6% instead of 95%.

timeseries · Dependence
The false-positive rate against the number of analyses, on pure noise. The data has no effect in it. Each analysis is correct and each p-value is honest. With 20 correlated outcomes available, something reaches p below 0.05 58% of the time.

Twenty analyses of nothing

Twenty honest, correct analyses of data with no effect in it find something significant 57% of the time. Nobody p-hacked, every individual p-value is right, and the reported one is the smallest of twenty.

testing · Forking
What 20 clusters of 20 correlated observations do to a 95% interval. Each study has 400 observations arranged as 20 clusters of 20. The lower points are the counted coverage of the usual interval, which treats them as 400 independent observations; the curve through them is 2Φ(1.96/√deff) − 1 with deff = 1 + 19ρ, computed before any data was drawn. At ρ = 0.81 the interval covers 36% rather than 95%. The upper points treat the cluster as the unit and need no variance components at all.

Two levels at once

A third level of grouping adds no new arithmetic and produces one number — the design effect — that decides how many independent observations a clustered study is worth. It is the same quantity the time-series field computes for autocorrelated data, arrived at from a completely different picture.

multilevel · Levels
Which samples Wilson and Clopper–Pearson each cover, n = 50, p = 0.2. Each bar is the probability of one count, shaded by which interval built on that count contains 0.2. Both cover 95.1% of samples, only Wilson 0.0%, only Clopper–Pearson 1.6%, neither 3.3%. The correlation between their hits is 0.810, so on shared draws the variance of their difference is 4.891 times smaller than on independent ones.

The same draws for both methods

Two intervals computed on the same simulated datasets give a difference in coverage whose variance can be 4.891 times smaller than on separate datasets — or, for a pair that covers different samples, 1.164 times larger. Which one a comparison gets is an exact sum over the counts each interval covers, and a standard error that ignores the sharing covers 100.00% for one pair and 93.07% for the other.

method · Seeds
The false discovery rate of twenty correlated tests, against the correlation. BH, every null true: 5.08% at 0, 4.86% at 0.3, 3.70% at 0.6, 2.34% at 0.9. BH, 10 of 20 real: 2.55% at 0, 2.53% at 0.3, 2.26% at 0.6, 1.66% at 0.9. BY, every null true: 1.46% at 0, 1.31% at 0.3, 1.03% at 0.6, 0.69% at 0.9. BY, 10 of 20 real: 0.72% at 0, 0.75% at 0.3, 0.64% at 0.6, 0.50% at 0.9. 20,000 families at each correlation.

False discoveries that arrive together

Correlate twenty tests and Benjamini–Hochberg still holds its false discovery rate — 1.66% at a correlation of 0.9 with ten real effects, against 2.55% when the tests are independent. What changes is how the errors come. A family of true nulls reports anything 2.34% of the time instead of 5.08%, and when it does, it reports 16.56 false findings out of twenty.

multiplicity · Multiplicity
The line is the sample, and the sample is the finding. 900 draws of two independent standard normal causes, with the 453 of them past a threshold of 0.00 marked and the 447 that fall short left pale. In the population the two are independent by construction. Inside the selected sample the correlation is -0.4669 in closed form and -0.5050 counted on these 453 rows, and the least-squares line through them has a slope of -0.545. The mechanism is visible in the picture rather than argued: the threshold removes one corner of the cloud, and a cloud with a corner missing is a cloud whose two coordinates carry information about each other.

The sample is a condition

Two independent standard normals, selected on their sum exceeding its median, read a correlation of exactly −1/(π − 1) = −0.4669 inside the sample. Nothing is measured badly and nothing is missing — and both halves of that split read it, in the same direction, while the population containing both reads zero.

collider · Conditioning
Storey's estimate of the share of true nulls over twenty thousand families, independent and correlated at 0.6. The true share is 0.5. Independent tests: mean 0.610, spread 0.160, below half the truth in 0.92% of families. Correlated at 0.6: mean 0.609, spread 0.240, below half the truth in 9.33%.

Estimating how many nulls are true

Benjamini–Hochberg at 5% delivers 2.55% when half of twenty nulls are false, because it cannot tell how many are. Storey's estimate of that share, read off the p-values above one half, spends the rest and finds 81.93% of the real effects instead of 74.70% on independent tests. Correlated at 0.9, the same procedure reports a finding in 19.29% of families in which every null is true.

multiplicity · Multiplicity
What 240 observations are worth, by which question is asked. 8 rows and 10 columns with 3 observations in each cell — 240 in all, each belonging to one row and one column, neither nested in the other. the overall mean: variance 0.1703 against a naive 0.0060, a design effect of 28.4 and 8.5 effective observations; a difference between two rows: variance 2.0682 against a naive 0.0960, a design effect of 21.5 and 11.1 effective observations; a difference between two columns: variance 1.0845 against a naive 0.1200, a design effect of 9.0 and 26.6 effective observations.

Two groupings that cross

Pupils belong to a school and to a neighbourhood, and neither is nested in the other. There is then no design effect: the overall mean is worth 8.5 independent observations out of 240, a row difference 11.1 and a column difference 26.6, and which grouping matters depends on the question rather than on the study.

multilevel · Levels
The top of a heavy-tailed population keeps its lead; the top of a light-tailed one gives it back. Select the top share on the first reading and read the group again: the share of its mean lead the second reading keeps, by integration over the true score (lines) and counted on 400,000 draws a parent in 20 batches (points, with two standard errors). The normal keeps exactly 0.6 at every selection. At the top half the Laplace keeps 0.541, the t 0.535 and the uniform 0.648 — the heavy tails keep LESS than the correlation. By the top one per cent the order has reversed: 0.761, 0.784 and 0.443. At one in ten thousand the t keeps 0.977 and the uniform 0.346.

A lead that a heavy tail keeps

Four populations whose readings all correlate at exactly 0.6, and whose least-squares slopes all read 0.6. Select the top one per cent on one reading and measure them again: they keep 60% of their lead if the true scores are normal, 76.1% if they are Laplace, 78.4% if they are a t on four degrees of freedom — and 44.3% if they are uniform. The correlation predicts the regression of the extremes for one shape of population only.

paradox · Rtm
How often each combination, and each union of them, rejects ten studies of nothing. Each combination alone rejects exactly 5% of null sets. Counted on 1,000,000 sets of ten null studies: Fisher or Stouffer 6.63%, Fisher or Tippett 8.05%, Stouffer or Tippett 8.96%, any of the three 9.66% — enclosed on a two-dimensional lattice between 9.18% and 10.09% — and all three together 1.05%. The three sizes add to 15%.

The smallest of three combinations

Reporting whichever of Fisher's, Stouffer's and Tippett's combinations is smallest is a test of its own, and on ten studies of nothing it rejects 9.66% of the time — not 5%, and nowhere near the 15% the three sizes add to, because the statistics are correlated at up to 0.903. Read at 2.448% each it is exact, and then it trails the best single combination by at most 7.45 points and leads the worst by at least 10.30.

testing · Uniformity
What one lost run costs a 16-run factorial fitting 11 coefficients. Every run is worth the same: dropping any one multiplies every coefficient's variance by 1.2000, which is 1 + 1/(N − p) with N = 16 and p = 11, and gives every pair of coefficients a correlation of 0.1667 where the complete design had exactly zero.

The run that did not happen

Lose one run from any orthogonal design and every coefficient's variance is multiplied by exactly 1 + 1/(N − p), and every pair of coefficients acquires a correlation of exactly 1/(N − p + 1) where there was none. The price is set by the design's spare capacity and by nothing else, and a saturated design cannot survive it at all.

design · Factorial
Two companions on one simulation, two hundredfold apart. How many times as many draws each companion is worth, on the same 4,000 simulated samples of 40 observations. The coverage of the interval is estimated with the observed count as its companion, whose expectation is 12 exactly; they correlate at 0.2665 and the companion is worth 1.08 times the draws. The expected width is estimated with p̂(1 − p̂) as its companion, whose expectation is 0.20475 exactly; they correlate at 0.9977 because the width is a monotone function of it, and the companion is worth 214 times the draws — 856 thousand simulated samples' worth of precision from four thousand.

The check worth more than the check

The same exactly known companion that verifies a simulation can sharpen it. On one set of four thousand draws, one companion is worth 1.08 times the draws and another is worth 214 times them, and the factor is 1 − ρ² with nothing else in it.

method · Routes

Named alongside it

The objects these essays reach for when they reach for this one.

Closed formCovariate balanceDependenceExperimental designOrthogonalityMonte CarloInteractionProjectionBasis functionsHermite polynomialsBasisMultiple comparisons

All concepts