The collection

Every essay — page 18

Essays 409 to 432 of 436, in the same order.

Two tests, a threshold, and the rate they are read against

A predictive value is computed from a sensitivity, a specificity and a prevalence, and each of the three is softer than it is quoted as being. Two positives from a 90/95 test on a one-in-a-thousand condition are worth 24.49% if the tests' errors are independent, 10.16% at a correlation of 0.1, and 3.16% at 0.5 — barely more than the 1.77% one positive was worth. The sensitivity and specificity are not two properties of a test but one property read at a threshold somebody chose, and the threshold that minimises harm moves by three standard deviations of the score across the prevalences such a test is used at. The prevalence itself is usually estimated from the same test's positive rate, which at one in a thousand reads 5.09%.

What a diagnostic plot is showing

A quantile plot is read by eye against a band drawn for one point at a time, and a genuinely normal sample of forty has forty chances to leave it: 45.0% of them do. The band a reader is actually using is that one widened by a factor of 1.502, and nothing draws it. The quantity is wrong as well as the width — a t interval needs the sampling distribution of the mean to be normal rather than the data, and the source that leaves the band on every sample of forty covers 94.80% while one that leaves it on 57% covers 95.73%. And what is plotted are not the errors: a design whose leverages run from 0.045 to 0.663 gives residuals whose spreads differ by a factor of 1.68 with the model exactly right.

Past the first term of the normal approximation

The Berry–Esseen bound guarantees how far a standardised sum can be from the normal, and on an exponential source it is 8.62 times the real worst error at every sample size — a constant set by a fair coin, whose error is a half-step, and at a hundred draws larger than the 2.5% tail it would be asked to certify. One Edgeworth term takes the normal's error at two standard deviations from 38% to 8% on ten draws, and turns negative in the short tail at every sample size, beyond a threshold that recedes only as the sixth root of n. The saddlepoint approximation is built at the threshold rather than at the mean, and reads a tail six standard deviations out to within 0.19% on five draws and within 2.2% on one.

A proportion's interval near the boundary, and the coin

Wilson's interval for a proportion wobbles a point or two around 95% away from the boundary and has a hole near it: its worst coverage is 83.50% at ten trials and 83.81% at a thousand, at an expected count of 0.1765 each time, because below that count the interval built on one success lies wholly above the truth. The intervals that never fall below 95% pay in average coverage, and the two of them pay for different promises — Clopper–Pearson holds each side under 2.5% and is 10.4% wider than Wilson at thirty trials, Blaker holds only the total and is 4.6% wider. Only an interval that adds a uniform random draw to the count covers exactly 95% at every proportion, at 0.9% more width than Wilson, and it is not used because two analysts with the same data would report different intervals.

Coverage of the Wilson interval against the expected count, at 10, 30, 100 and 1,000 trials. Read against the expected number of successes the four sample sizes draw the same curve near the boundary. The worst coverage is 83.50% at n = 10, 83.71% at n = 30, 83.79% at n = 100, 83.81% at n = 1000, each at an expected count near 0.177, and the limiting depth is e^(−0.1765) = 83.82%.

A hole no sample size fills

Wilson's interval is the recommended repair for a proportion, and away from the boundary it wobbles a point or two around 95%. Near zero it has a hole: at an expected count of 0.1765 its coverage is 83.50% at ten trials, 83.79% at a hundred and 83.81% at a thousand, and it never climbs past e to the minus 0.1765, which is 83.82%. The hole is where the interval built on one success stops containing the truth, and it belongs to the count rather than to the sample size.

5 figures · Oscillation, part 3
Worst and average coverage of six 95% intervals for a proportion, 30 trials. The worst coverage over every proportion beside the average over a uniform one, with the average expected width. Clopper–Pearson: worst 95.05%, average 97.34%, width 0.299. Blaker: worst 95.00%, average 96.31%, width 0.283. Wilson: worst 83.71%, average 95.24%, width 0.271.

What a guaranteed minimum costs

Clopper–Pearson's interval never covers less than 95%, and at thirty trials it averages 97.34% and is 10.4% wider than Wilson's. Blaker's interval keeps the same guarantee, averages 96.31% and is 4.6% wider. The difference is not waste: Clopper–Pearson guarantees each side separately, holding both below 2.5%, and Blaker guarantees only their sum — so at ten trials and a proportion of 0.15 it misses on one side 5.00% of the time.

5 figures · Oscillation, part 4
Clopper–Pearson, mid-p and the randomised interval: coverage across the proportion, 20 trials. Clopper–Pearson never falls below 95% and runs up to 99.80%. The mid-p interval, which is the randomised interval with its coin fixed at one half, runs from 92.94% to 99.80%. The randomised interval covers 95% at every proportion, to within the 0.043% of the numerical integration over the coin.

The coin that makes it exact

Every interval for a proportion either covers less than 95% somewhere or more than 95% on average, because a count is discrete. One construction covers exactly 95% at every proportion: it adds a uniform random draw to the count. At thirty trials it is 0.9% wider than Wilson's interval and narrower than both exact ones — and two analysts with the same data report different intervals, and one study in forty that sees nothing reports an empty one.

5 figures · Oscillation, part 5

An interval read beside something else

A 95% interval is almost never read alone. Read against a replication's estimate it captures an equal-sized one 83.42% of the time, because both estimates are uncertain, and ten replications of one study miss together: none of them lands outside 31.40% of the time against 16.32% if they were independent. Read against a second interval, two that just touch mark a difference with p = 0.0056 rather than 0.05, standard-error bars that touch mark p = 0.157, and in a figure of ten identical groups some pair of 95% intervals fails to overlap one time in seven. Read against a second verdict, two studies of one effect at 50% power disagree about significance half the time, and when they do the difference between them is significant in 9.75% of cases.

Where a replication's estimate lands against a 95% interval, replication the same size. The chance that a 95% interval contains a replication's estimate is 95.00% when the original landed on the truth, 82.99% one standard error away and 48.40% two away. Averaged over where originals land it is 83.42%, and 5.00% of originals capture a replication less than half the time.

Five times in six

A 95% interval is read as a 95% chance that a replication's estimate will land inside it. With the spread known and a replication of the same size, the chance is 83.42% — five times in six — because both estimates are uncertain. An original that landed two standard errors from the truth captures a replication 48.40% of the time; a replication a tenth the size lands inside 44.54% of the time; and among significant originals from studies with 17% power, 66.94%.

7 figures · Repetition, part 2
Two 95% intervals 2.772 standard errors of the difference apart, standard errors in the ratio 1. The intervals are separated, and the test of the difference gives p = 0.0056. Two 95% intervals with equal standard errors just touch at p = 0.0056.

Two intervals that overlap

Two 95% intervals that just touch are read as a difference at the edge of significance. With equal standard errors their difference has p = 0.0056, not 0.05; two intervals can overlap by 58.6% of an arm and still differ at exactly 5%; standard-error bars that just touch mark p = 0.157; and when the two estimates are correlated at 0.8, touching intervals conceal a difference of 6.2 standard errors. Read as a test, non-overlap needs 1.66 times the sample for the same power.

6 figures · Repetition, part 3
Two studies of the same effect, z statistics with mean 1.96: where one is significant and the other is not. Five hundred pairs. With the true effect identical in both, exactly one of the two is significant in 50.0% of pairs; among those, the difference between the two is significant in 9.7%. The dashed lines mark 1.96 on each axis; the diagonal lines mark a significant difference.

Significant in one, not in the other

Two studies of exactly the same effect, each with 50% power, disagree about significance half the time — and when they do, the test of the difference between them is significant in 9.75% of cases. A p of 0.01 beside a p of 0.20 is a difference with p = 0.36. Among four subgroups sharing one effect, at least one significant and one not happens 87.5% of the time, and the test that would tell a real difference apart needs four times the sample the effect itself needed.

5 figures · Repetition, part 4

What partial pooling does to one group, to the set, and to a ranking

Partial pooling minimises the total squared error across groups, and three uses of its estimates are not the total. One group: with each group's standard error equal to the population's spread, every group truly more than 1.73 widths from the centre — 8.33% of a normal population — is estimated worse than by its own mean, and among eight groups with the spread estimated the truly most extreme one is worse off in 61.6% of datasets; capping each group's shift at one standard error improves the total, the truly extreme group and the observed extreme group at once. The set: posterior means spread 0.707 as widely as the truth and count 0.234% of groups beyond two widths against a true 2.28%. The ranking: in a league table of a hundred groups of unequal size, raw means fill 62.1% of the top ten with small groups, posterior means 13.4%, and no ranking recovers more than 5.47 of the true ten.

Which groups partial pooling serves, standard error 1 population width. Pooling's expected squared error for a group, divided by its own mean's, against how far the group truly sits from the centre. It is ×0.25 at the centre and crosses ×1 at 1.732 population widths, beyond which 8.33% of a normal population lies; capping the shift at one standard error holds every group under ×2.

A group from the population's own tail

Partial pooling halves the total squared error when a group's own standard error equals the spread between groups. Every group whose true effect sits more than 1.73 population widths from the centre — 8.33% of a perfectly normal population — does worse than it would have with its own mean, and its loss grows without bound. Among eight groups with the spread estimated, the most extreme is worse off in 61.6% of datasets. Capping the shift at one standard error keeps the total at 0.528 of the unpooled error and holds every group under twice it.

6 figures · Shrinkage, part 4
How the truth, the raw means, the posterior means and the constrained estimates spread, standard error 1. Beyond two population widths above the centre lie 2.28% of the true values, 7.86% of the raw means, 0.234% of the posterior means, and 2.28% of the constrained estimates.

Estimates that are too alike

Posterior means give each group its least-error estimate, and as a set they are too alike: with each group's standard error equal to the population's spread, they spread 0.707 as widely as the truth. Beyond two population widths lie 2.28% of the true effects, 7.86% of the groups' own means, and 0.234% of the posterior means — a tenth of the truth. Rescaling the estimates to the right spread counts the tail exactly and costs 17% more squared error; summing each group's posterior chance of being beyond the line counts it without changing any estimate.

5 figures · Shrinkage, part 5
How many of a league table's top ten are small groups, ranked four ways. A hundred groups with sizes from 4 to 400, of which 36% have twenty units or fewer. Small groups make up 36.3% of the true top ten, 62.1% of the top ten by raw means, 13.4% by posterior means and 22.9% by the posterior chance of being in the top ten. The three rankings recover 4.43, 5.38 and 5.47 of the true top ten.

A league table of a hundred

A hundred groups with sizes from 4 to 400, and a top ten to publish. Ranked by their own means, small groups fill 62.1% of the top ten against their 36.3% share of the true top ten. Ranked by posterior means they fill 13.4%. The ranking built from each group's chance of being in the top ten recovers 5.47 of the true ten, the best of three and barely half; and the group ranked first could hold any rank from 1 to 31.

5 figures · Shrinkage, part 6

What a sample-size calculation was given

A sample-size calculation takes a standard deviation, an effect and an outcome as given, and each is usually a choice or an estimate. Sized from a pilot of ten's standard deviation, 55.9% of trials have less than their planned 80% power and 11.1% less than half, although the planned sample is right on average; sizing from the pilot's 80% upper confidence limit leaves 19.8% short at 1.65 times the patients. With the effect believed to be half a standard deviation give or take a quarter, a trial of sixty-four per arm succeeds 69.2% of the time, and no trial of any size can succeed more often than the planners believe the effect is positive. Cutting the outcome at a threshold keeps at most 63.7% of its information, and a responder threshold one standard deviation out needs twice the patients.

The power trials actually have when sized for 80% from a pilot of 10. Four thousand pilots of 10 observations, each sizing a trial for 80% power at half a standard deviation from its own standard deviation. 55.9% of the trials have less than 80% power and 11.1% less than 50%; the median trial has 76.8%.

The spread a pilot supplies

A trial sized for 80% power from a pilot's standard deviation is sized from an estimate that is too small more often than not. With a pilot of ten, 55.9% of the trials it sizes have less than 80% power and 11.1% less than 50%, although the planned sample is right on average. Sizing from the pilot's 80% upper confidence limit instead leaves 19.8% short, at 1.65 times the sample; from its 90% limit, 10.0% short at 2.12 times.

5 figures · Power, part 4
The chance a trial succeeds against its size, when the expected effect of 0.5 is uncertain by four amounts. With the effect known, 80% is reached at 63 per arm. With the effect uncertain by 0.25 standard deviations it takes 113; by 0.5, 1268; by 0.75, no sample size at all, because the chance can never exceed the 74.8% prior probability that the effect is positive.

The chance a trial succeeds

A trial of sixty-four per arm has 80% power at an effect of half a standard deviation. If the effect is only believed to be about half a standard deviation, give or take a quarter, the chance the trial reaches significance is 69.2%; give or take a half, 61.4%. Reaching 80% then takes 113 per arm, or 1,268 — and when the belief is uncertain by three quarters of a standard deviation no number of patients reaches 80%, because the chance can never exceed the 74.8% probability that the effect is positive at all.

5 figures · Power, part 5
How much of a normal outcome's information survives cutting it into two, by where the cut is. For a small shift, a cut at the mean keeps 63.7% of the information, so the trial needs 1.57 times the sample. A cut at the top tenth keeps 34.2% and needs 2.92 times; a cut two standard deviations out keeps 13.1%.

An outcome cut in two

Replacing a measured outcome with whether it crossed a threshold keeps 63.7% of the information when the cut is at the mean, 34.2% at the top tenth and 13.1% two standard deviations out. A trial that needs 63 patients per arm on the measured outcome needs 102 cut at the mean and 185 cut at one and a half standard deviations. The responder rates that result read as a share of patients who respond — 50.0% against 69.1% — when every patient moved by the same amount; and a cut chosen after looking turns a 5% test into a 17.7% one.

6 figures · Power, part 6

What makes it checkable

A stated seed for every random draw, so a simulated figure comes out the same every time it is drawn. A closed form beside every simulation, so there are two routes to each number. And the simulation's own error counted as carefully as the method's: a table of correct cells that flags one of them most of the time, two methods whose comparison is sharper or noisier for sharing their draws, a run that stops when it looks done and reports what it was looking for, and draws aimed at a tail that look settled while their variance is infinite.

FieldsThreadsSeriesConceptsFigure librarySearch