The collection

Every essay — page 4

Essays 73 to 96 of 436, in the same order.

Corrections, and what each controls

Bonferroni bounds the chance of any false positive; Benjamini–Hochberg bounds the share of the findings that are false. Both get called correcting for multiple comparisons and they are different promises, so every procedure here is made to report both rates and the power each one costs.

Power to find a real effect of 3 standard errors, 10 of 20 real. no correction finds 85.1%, Bonferroni finds 49.1%, Holm finds 52.5%, Benjamini–Hochberg finds 74.9%. The uncorrected procedure finds the most and controls nothing.

The price of control

Every correction is paid for in power, and the exchange rate can be measured. Holm buys familywise control for 33 percentage points of power; Benjamini–Hochberg buys a weaker guarantee for 10. Neither is free and neither is a matter of taste.

6 figures · Multiplicity, part 3
The false discovery rate of twenty correlated tests, against the correlation. BH, every null true: 5.08% at 0, 4.86% at 0.3, 3.70% at 0.6, 2.34% at 0.9. BH, 10 of 20 real: 2.55% at 0, 2.53% at 0.3, 2.26% at 0.6, 1.66% at 0.9. BY, every null true: 1.46% at 0, 1.31% at 0.3, 1.03% at 0.6, 0.69% at 0.9. BY, 10 of 20 real: 0.72% at 0, 0.75% at 0.3, 0.64% at 0.6, 0.50% at 0.9. 20,000 families at each correlation.

False discoveries that arrive together

Correlate twenty tests and Benjamini–Hochberg still holds its false discovery rate — 1.66% at a correlation of 0.9 with ten real effects, against 2.55% when the tests are independent. What changes is how the errors come. A family of true nulls reports anything 2.34% of the time instead of 5.08%, and when it does, it reports 16.56 false findings out of twenty.

8 figures · Multiplicity, part 9
Storey's estimate of the share of true nulls over twenty thousand families, independent and correlated at 0.6. The true share is 0.5. Independent tests: mean 0.610, spread 0.160, below half the truth in 0.92% of families. Correlated at 0.6: mean 0.609, spread 0.240, below half the truth in 9.33%.

Estimating how many nulls are true

Benjamini–Hochberg at 5% delivers 2.55% when half of twenty nulls are false, because it cannot tell how many are. Storey's estimate of that share, read off the p-values above one half, spends the rest and finds 81.93% of the real effects instead of 74.70% on independent tests. Correlated at 0.9, the same procedure reports a finding in 19.29% of families in which every null is true.

7 figures · Multiplicity, part 10
Twenty hypotheses tested in a declared order, the ten real effects listed first. Effects of three standard errors, ten real, familywise 5%. fixed sequence: 85.3% at position 1, 45.0% at 5, 20.4% at 10; overall power 46.10%; fallback: 49.1% at position 1, 56.4% at 5, 57.9% at 10; overall power 55.87%; Holm: 52.5% at position 1, 52.2% at 5, 52.5% at 10; overall power 52.53%.

An order that spends the error rate

Test twenty hypotheses in a declared order, each at the full 5% and each only if every one before it was rejected, and the first is found 85.3% of the time where Holm finds it 52.5%. The tenth is found 20.4% of the time, the product of the powers before it. Move one true null to the head of the list and every real effect behind it is found no more than 4.3% of the time.

6 figures · Multiplicity, part 11
Every finding Benjamini–Hochberg made in thirty families, with its interval, at real effects of 2. 85 findings, sorted by their estimate. 12 of their ordinary 95% intervals miss the true effect, every one of them on the far side; 6 of the wider false-coverage-rate intervals miss.

Intervals for the findings

Benjamini–Hochberg's findings usually go out each with its ordinary 95% interval. With ten real effects of two standard errors among twenty tests, 11.59% of those intervals miss their effect, every miss on the far side, and the interval around the most prominent finding covers 72.36% of the time — 2.38% when the effects are one standard error. Intervals widened for the number of findings hold the share that miss under 5%.

5 figures · Multiplicity, part 12

When the data stops early

A subject still event-free when a study ends is not missing and not observed — it is known to exceed something, which is a third state most tools have no slot for. The estimator that gives it that slot recovers the curve, and everything read off the curve afterwards has a condition attached: the interval printed around its end stops covering where the curve is read, a dropout that carries information leaves a record identical to one that does not, one minus the curve is not a risk once something else can end observation first, and a hazard ratio is an average whose weights the length of follow-up chose.

The same study read three ways. At time 2 the truth is 0.497. Kaplan–Meier gives 0.532; dropping the censored subjects gives 0.180; treating the censoring time as the event time gives 0.392. Both naive readings understate survival, because the subjects they mishandle are the ones doing well.

The data that stops early

A subject still event-free when a study ends is not missing and not observed. It is known to exceed something, which is a third state most tools have no slot for — and the two obvious ways of forcing it into one are wrong by 31 and 13 percentage points.

6 figures · Censoring, part 1
Kaplan–Meier from 120 subjects, 59 of them censored. The step curve is the estimate, the smooth curve is the truth it is trying to recover. 59 of 120 subjects were still event-free when observation stopped; they are not dropped, and they are not counted as events — they leave the risk set at the time they were last seen.

The curve that survives censoring

Kaplan–Meier recovers the true survival curve to within a fraction of a point at every censoring level from 37% to 71%, where dropping the censored subjects is off by 28 and then by 43. The estimator is a running product and the reason it works is in its denominator.

7 figures · Censoring, part 2
One cohort of 40, and two intervals around the end of its curve. A single simulated study of 40 subjects with exponential survival at rate 0.35, dropout at rate 0.15 and follow-up to 6 — the first seed from 8811 upward whose plain band reaches below −0.05, chosen to show the failure rather than its frequency. The step curve is Kaplan–Meier and the smooth curve the truth. The plain band, the estimate plus and minus 1.96 Greenwood standard errors, first dips below zero at t = 3.78 and reaches −0.052; early on it also rises to 1.023, above one. At t = 5 the estimate is 0.069 with 1 subject still under observation, the plain interval runs from −0.052 to 0.191 and the log-log interval from 0.006 to 0.251, against a truth of 0.174. The log-log band is built on a scale that cannot leave [0, 1], and it bends away from the edge rather than through it.

The interval at the end of the curve

The interval most software prints around a survival curve covers 89.7% at five years, where 3.3 of forty subjects are still being watched and where the curve is actually read. The same variance carried on a log–log scale covers 94.8% there — and the failure was never the width.

6 figures · Censoring, part 3
Two worlds, one Kaplan–Meier curve, two truths. World A gives each subject a frailty with mean one and variance 1, and multiplies both its event hazard (0.35) and its dropout hazard (0.5) by it, so the subjects likeliest to leave are the ones likeliest to fail. World B has independent event and dropout times whose hazards are world A's crude hazards. Kaplan–Meier over 1000 studies of 400 gives the same curve from both — 0.7750 and 0.7763 at t = 1; 0.6627 and 0.6641 at t = 2; 0.5924 and 0.5932 at t = 3; 0.5047 and 0.5042 at t = 5 — and that curve is world B's truth, 0.5052 at t = 5. World A's truth is 0.3636 there. The dashed lines are the two bounds that assume nothing, from every dropout failing on leaving (0.1905 at t = 5) to none ever failing (0.6667).

A dropout the data cannot see

Two worlds leave the same record to the last detail a study can write down — the same times, the same share ending in the event, the same share leaving first — and a log-rank test between them rejects at its own 5% level at every sample size from a hundred to sixteen hundred. Kaplan–Meier converges on 0.5052 at t = 5 from both. The truth is 0.5052 in one and 0.3636 in the other, and what is left to argue about is where between two bounds to stand.

6 figures · Censoring, part 4
The risk of one cause, estimated two ways. Two causes of an ending event with constant hazards 0.2 (the one of interest) and 0.3 (the competitor), random dropout at 0.1 and follow-up to 6. The lower line is the cumulative incidence, (0.2/0.5)(1 − e^(−0.5t)), the chance of actually having had this event by t; the dots on it are the Aalen–Johansen estimate over 2000 studies of 300, 0.3670 at t = 5 against 0.3672. The upper line is 1 − e^(−0.2t), and the dots on it are one minus Kaplan–Meier with the competing event treated as censoring: 0.6318 at t = 5 against 0.6321. The second is larger by a factor of 1.722 at t = 5, and it is not an error of estimation. It estimates, correctly, the risk in a population where the competing cause does not exist.

One minus Kaplan–Meier is not a risk

With two ways for observation to end, one minus Kaplan–Meier for one cause reads 0.6318 at t = 5 where the chance of actually having had that event is 0.3670. Added across the two causes, the complements reach 1.4088 — more than the whole cohort. Nothing is estimated badly: the complement estimates, correctly, the risk in a world where the other cause does not exist.

6 figures · Censoring, part 5
One treatment, a different hazard ratio at every follow-up. The hazard ratio a Cox model converges to, found as the root of its expected score by numerical integration, as the trial runs longer; dropout at 0.1 throughout. The proportional treatment reads 0.5 at every τ. The waning treatment reads 0.5000 at τ = 1, 0.6362 at τ = 3 and 0.7890 at τ = 8 — the same two arms, the same effect in the same first year, and a number that drifts towards one as later, effect-free events are added to the average. The dots are the mean of 400 Cox fits with 400 subjects an arm: 0.5014 at τ = 1, 0.7020 at τ = 2, 0.7635 at τ = 3, 0.8034 at τ = 5, 0.8208 at τ = 8. The crossing treatment reads 0.3429 at τ = 1, exactly 1 at τ = 3 by construction, and 1.0611 at τ = 8: beneficial, null or harmful according to when the trial stopped.

The hazard ratio the follow-up chose

A treatment that halves the hazard for one year and then does nothing has a Cox hazard ratio of 0.5000 if the trial stops at one year, 0.7617 at three and 0.8194 at eight. Nothing about the treatment differs between those numbers. When hazards are not proportional the hazard ratio is an average, and the length of follow-up and the dropout rate choose its weights.

7 figures · Censoring, part 6
The difference in restricted mean survival at every horizon, in three worlds. Treatment minus control, in closed form, with dropout irrelevant to the truth. The proportional treatment's difference grows to 0.4766 at τ = 3 and the waning treatment's to 0.2675. The crossing treatment's rises to 0.1776 at τ = 2, near where the two survival curves cross, and falls back to 0.1366 at τ = 3. The ticks along the bottom are the eleven horizons, from 0.5 to 3 in quarters, at which a trial below reads its differences.

A horizon chosen after looking

A difference in restricted mean survival read at whichever of eleven horizons looks most convincing rejects 11.24% of trials in which the treatment does nothing, against 4.70% at a horizon fixed in advance. The correlation of the differences across horizons is closed, and the Gaussian process it defines prices the choice at a critical value of 2.317 — which brings the counted size back to 4.99% and keeps 96.92% of the power that a horizon nobody could have known to fix would have had.

6 figures · Censoring, part 7

The value that is not there

A missing value is missing conditional on something, and which something decides everything that follows. Dropping the incomplete rows leaves a regression slope exactly right when the chance of being observed depends on the regressor, however strongly, and wrong by a quarter of itself when it depends on the outcome — and the data cannot tell those two cases apart. Filling the gaps in is not free either: one filled value is not an observation, and the arithmetic that makes several of them into one is two corrections rather than the one everybody quotes.

Three mechanisms leave the slope alone; one does not. The bias of the complete-case slope under each of four missingness rules, counted over 4000 studies of 200 rows at 35.0% missing, with the closed form printed beside each count. Missingness that depends on nothing, on the regressor, or on the second covariate leaves the slope exactly where it was — the closed forms are zero to machine precision and the counts are -0.0005, -0.0005 and -0.0011 against standard errors of about 0.0018. Missingness that depends on the outcome moves it by -0.1635, which is 27.3% of the slope being estimated. The same share of rows is lost in every case.

Three mechanisms and one dataset

Four rules for which outcomes go missing, each calibrated to lose the same 35% of the rows and each leaning on what it reads with the same coefficient. Three leave the fitted slope exactly where it was, and the one that reads the outcome moves it by 0.163531.

6 figures · Missingness, part 1
Unrepresentative in every respect but the one that matters. Three properties of the complete cases as the chance of being observed leans harder on the regressor, in closed form, at 35.0% of outcomes missing throughout. The mean of the regressor among the rows kept climbs from 0.0000 to 0.5528 against a population mean of zero, and the mean of the outcome from 0.0000 to 0.3980 above its own. The bias in the fitted slope is exactly zero at every one of the ten settings, because selection acting on the regressor alone leaves the conditional law of the outcome given the regressor untouched and least squares conditions on exactly that. The sample is wrong about almost everything and right about the one quantity being estimated.

Dropping the incomplete rows

Push the missingness until the rows that survive have a covariate mean of 0.543905 against a population zero and a variance of 0.5041 against one, and the fitted slope is still exactly right. Where the rule reads the outcome instead, the same sweep takes coverage to 2.42% at eight hundred rows.

6 figures · Missingness, part 2
A filled value is not an observation. What a 95% interval for the slope actually covers after each way of handling 35.0% missing outcomes, counted over 4000 studies of 200 rows. Dropping the incomplete rows covers 95.93%. Filling with the observed mean covers 13.85%, because the estimate itself has moved. Filling with a fitted value covers 80.85% against a closed prediction of 79.73%: the estimate is right and the reported standard error is short by a factor of 0.6567 against a predicted 0.6500, because the residual sum of squares is divided by the whole sample's degrees of freedom. Adding residual noise recovers the spread and covers 85.78% against a predicted 84.62%, since the interval still ignores the variance of having imputed at all.

One imputation is not an observation

Three ways of filling a missing outcome, under a mechanism that makes dropping the rows beyond reproach. Filling with the observed mean covers 13.85%, filling with a fitted value covers 80.85%, adding noise covers 85.78%, and the thing all three were meant to improve on covers 95.93%.

6 figures · Missingness, part 3
The extra 1/m, and the correction nobody quotes. What a pooled 95% interval covers against the number of imputations, counted over 2000 studies of 200 rows at 35.0% of outcomes missing. Rubin's rules — total variance W̄ + (1 + 1/m)B, read against a t distribution on (m − 1)(1 + W̄/((1 + 1/m)B))² degrees of freedom — cover 94.10% at two imputations and reach their promise by 5, at 95.25%. Dropping the (1 + 1/m) factor takes two imputations to 93.10%; using a normal quantile instead of the degrees-of-freedom correction takes it to 92.55%; dropping both takes it to 91.45%. The median degrees of freedom at two imputations is 12.95, which is why the second correction is the larger.

The variance between imputations

Pooling several filled datasets covers 94.10% at two imputations and reaches its promise at five, where a single fill covered 85.78%. The correction everybody quotes is the smaller of the two doing the work — 1.00 ± 0.22 points against 1.55 ± 0.28.

6 figures · Missingness, part 4
The damage does not stay in the term that was left out. Where each coefficient lands when the model that fills the missing outcomes and the model that analyses them disagree, over 1500 studies of 200 rows at 35.0% missing and 20 imputations. An imputer that omits a covariate the analysis fits attenuates that covariate's coefficient by exactly the missing fraction — -0.1405 counted against a closed -0.1400 — and pushes the coefficient it did impute on the other way by exactly the product of the omitted coefficient, the covariates' correlation and the missing fraction: 0.0402 counted against 0.0420. Both closed forms come out of the same two-by-two solve. Matching models leave both alone, and so does an imputer that knows more than the analysis.

An imputation model the analysis does not contain

A model that fills the gaps without a covariate the analysis fits attenuates that covariate's coefficient by exactly the missing share, 0.4 to 0.26, and moves the one it did carry by exactly γρf, 0.6 to 0.642. The reverse case is supposed to inflate the interval, and at four strengths of the extra knowledge it does not.

6 figures · Missingness, part 5
An estimate reported as a function of an assumption. What the slope really is, against a shift in the outcomes nobody saw — line from the closed form, dots counted over 2000 studies of 200 rows at 35.0% missing. Every point on this line produces exactly the same observed data, and the complete-case estimate is the flat line at 0.5996 regardless. The truth moves at -0.2845 per unit of shift, which is a function of the missingness model and the missing fraction and of nothing that can be estimated: across the swept range the true slope runs from 0.8845 to 0.3155, a span of 0.5691 against a value of 0.60 in the world where the shift is zero. Reporting the line is the honest form of the answer.

The mechanism the data cannot see

Two worlds produce identical covariates, identical patterns of what is recorded and identical recorded outcomes, to the last bit. Their true slopes are 0.6 and 0.315452, and the truth moves at 0.284548 per unit of an assumption nothing in the data can inform.

6 figures · Missingness, part 6
A peak where the recorded data have none. The profile log-likelihood of a selection model in cy, the coefficient that lets the chance of being recorded depend on the outcome itself, for one study of 800 rows whose missingness is at random, with residuals normal; every other parameter is maximised at each fixed value. The model assumes the outcome is normal given the covariates. The curve peaks at cy = 0.35, where the fitted slope is 0.839, and the values of cy within the 95% cut run from −0.13 to 0.75; the likelihood-ratio statistic against cy = 0 is 1.47. The study was drawn with cy = 0.00. With the outcome's law left free, every value of cy fits the recorded rows equally well and this curve would be flat: its curvature is the normal assumption.

The assumption that identifies the mechanism

A selection model estimates how strongly an outcome decides whether it is recorded — the quantity two identical datasets showed no statistic can see — and it does so by assuming the outcome is normal. Where that holds and the outcome does decide, it repairs a slope complete cases put at 0.4318 to 0.5795. Where the missingness is at random and the residual is merely skewed, it reports selection that is not there, moves the slope from 0.5971 to 1.0319, and rejects missingness at random in 72.5% of studies.

7 figures · Missingness, part 7

Stopping rules

A p-value is defined relative to a sampling plan, so two experiments with identical data and different stopping rules have different p-values. That sounds philosophical and it is arithmetical: testing five times at the nominal level rejects a true null 14% of the time.

Testing at 0.05 every time the data is looked at. The null is true in every one of these trials and the test is correct every time it is run. Looking once rejects 4.9% of the time, as it should; looking ten times rejects 19.2% of the time. Nothing changed except permission to look.

When the looking happens

A p-value is defined relative to a sampling plan, so the same data means different things under different stopping rules. Testing five times at the nominal level rejects a true null 14% of the time, and no observation in the dataset changed.

6 figures · Stopping, part 1
Four boundaries for 5 looks, all spending 5% in total. test at 0.05 every look: 1.96, 1.96, 1.96, 1.96, 1.96. Pocock — a constant, higher boundary: 2.41, 2.41, 2.41, 2.41, 2.41. O'Brien–Fleming — strict early, nearly nominal at the end: 4.55, 3.22, 2.63, 2.27, 2.03. Bonferroni across looks: 2.58, 2.58, 2.58, 2.58, 2.58. Every one except the first spends the same total error rate; they differ in when they spend it.

Spending the error rate

The repair for interim testing is to spend 5% across the looks rather than at each one. The boundaries are solvable rather than quotable, and a trial that can stop early uses 298 observations where a fixed design uses 400 — at a cost of half a point of power.

6 figures · Stopping, part 2
Forty O'Brien–Fleming trials at a true effect of 0.16, with the boundary written as an effect. The dashed line is the smallest effect a trial can report and still stop at each look: 0.510 at 80 observations, 0.255 at 160 observations, 0.170 at 240 observations, 0.128 at 320 observations, 0.102 at 400 observations. The true effect is 0.16, so at 3 of the five looks a trial cannot stop without reporting more than it. 29 of these forty trials stop before the last look, each marked where it stopped.

The effect a stopped trial reports

An O'Brien–Fleming trial at 88.45% power holds its error rate exactly and reports an effect 9.6% too large on average. The 11.39% of trials that stop at the second look report 1.83 times the truth, the ones that cross at the last look report 0.80 times it, and pooling every trial by its size gives the truth back to the last digit.

7 figures · Stopping, part 11
Ordered stagewise: the outcomes at least as extreme as stopping at 160 observations with z = 3.3. Each column is one look of an O'Brien–Fleming trial; above the boundary a trial stops there. Highlighted are the outcomes that count as at least as extreme as the observed one when outcomes are ordered stagewise: at 80, z ≥ 4.56 (probability 2.54 × 10⁻⁶ with no effect); at 160, z ≥ 3.30 (probability 4.82 × 10⁻⁴ with no effect); at 240, none; at 320, none; at 400, none. The two-sided p-value is 9.69 × 10⁻⁴.

The outcomes a trial could have stopped with

A trial that stops at its second look with z = 3.3 has a two-sided p-value of 0.000969, 0.000987, 0.00187 or 0.0421, depending on how the outcomes it could have stopped with are ordered. One of the four orderings does not change when the looks the trial never reached are replanned, and the same one gives a trial that ran to the end with z = 6 a p-value of 0.0256.

7 figures · Stopping, part 12
Forty trials at a true effect of 0.16, under the rule "power at the trend < 10%". The upper line is the benefit boundary (4.56, 3.23, 2.63, 2.28, 2.04); the lower line is where the rule stops a trial for futility (0.40 at 80, 0.66 at 160, 0.95 at 240, 1.31 at 320). Of forty trials with a real effect, 29 cross for benefit and 11 are stopped for futility.

A boundary for giving up

Adding "stop if z is below zero" to an O'Brien–Fleming trial costs 5.20 points of power at the effect it was designed for and halves the observations a trial with no effect uses. Stopping when conditional power at the observed trend falls under 10% costs 13.23 points and stops 21.28% of trials with a real effect. Making that rule binding lowers the benefit boundary from 2.040 to 1.901, and a binding rule that is then ignored rejects a true null 3.523% of the time instead of 2.5%.

8 figures · Stopping, part 13

FieldsThreadsSeriesConceptsFigure librarySearch