The collection

Every essay — page 20

Essays 457 to 470 of 470, in the same order.

What partial pooling does to one group, to the set, and to a ranking

Partial pooling minimises the total squared error across groups, and three uses of its estimates are not the total. One group: with each group's standard error equal to the population's spread, every group truly more than 1.73 widths from the centre — 8.33% of a normal population — is estimated worse than by its own mean, and among eight groups with the spread estimated the truly most extreme one is worse off in 61.6% of datasets; capping each group's shift at one standard error improves the total, the truly extreme group and the observed extreme group at once. The set: posterior means spread 0.707 as widely as the truth and count 0.234% of groups beyond two widths against a true 2.28%. The ranking: in a league table of a hundred groups of unequal size, raw means fill 62.1% of the top ten with small groups, posterior means 13.4%, and no ranking recovers more than 5.47 of the true ten.

What a sample-size calculation was given

A sample-size calculation takes a standard deviation, an effect and an outcome as given, and each is usually a choice or an estimate. Sized from a pilot of ten's standard deviation, 55.9% of trials have less than their planned 80% power and 11.1% less than half, although the planned sample is right on average; sizing from the pilot's 80% upper confidence limit leaves 19.8% short at 1.65 times the patients. With the effect believed to be half a standard deviation give or take a quarter, a trial of sixty-four per arm succeeds 69.2% of the time, and no trial of any size can succeed more often than the planners believe the effect is positive. Cutting the outcome at a threshold keeps at most 63.7% of its information, and a responder threshold one standard deviation out needs twice the patients.

The power trials actually have when sized for 80% from a pilot of 10. Four thousand pilots of 10 observations, each sizing a trial for 80% power at half a standard deviation from its own standard deviation. 55.9% of the trials have less than 80% power and 11.1% less than 50%; the median trial has 76.8%.

The spread a pilot supplies

A trial sized for 80% power from a pilot's standard deviation is sized from an estimate that is too small more often than not. With a pilot of ten, 55.9% of the trials it sizes have less than 80% power and 11.1% less than 50%, although the planned sample is right on average. Sizing from the pilot's 80% upper confidence limit instead leaves 19.8% short, at 1.65 times the sample; from its 90% limit, 10.0% short at 2.12 times.

5 figures · Power, part 4
The chance a trial succeeds against its size, when the expected effect of 0.5 is uncertain by four amounts. With the effect known, 80% is reached at 63 per arm. With the effect uncertain by 0.25 standard deviations it takes 113; by 0.5, 1268; by 0.75, no sample size at all, because the chance can never exceed the 74.8% prior probability that the effect is positive.

The chance a trial succeeds

A trial of sixty-four per arm has 80% power at an effect of half a standard deviation. If the effect is only believed to be about half a standard deviation, give or take a quarter, the chance the trial reaches significance is 69.2%; give or take a half, 61.4%. Reaching 80% then takes 113 per arm, or 1,268 — and when the belief is uncertain by three quarters of a standard deviation no number of patients reaches 80%, because the chance can never exceed the 74.8% probability that the effect is positive at all.

5 figures · Power, part 5
How much of a normal outcome's information survives cutting it into two, by where the cut is. For a small shift, a cut at the mean keeps 63.7% of the information, so the trial needs 1.57 times the sample. A cut at the top tenth keeps 34.2% and needs 2.92 times; a cut two standard deviations out keeps 13.1%.

An outcome cut in two

Replacing a measured outcome with whether it crossed a threshold keeps 63.7% of the information when the cut is at the mean, 34.2% at the top tenth and 13.1% two standard deviations out. A trial that needs 63 patients per arm on the measured outcome needs 102 cut at the mean and 185 cut at one and a half standard deviations. The responder rates that result read as a share of patients who respond — 50.0% against 69.1% — when every patient moved by the same amount; and a cut chosen after looking turns a 5% test into a 17.7% one.

6 figures · Power, part 6
The sample a trial needs for the same power, as a share of the unadjusted analysis's, by how strongly a baseline covariate predicts the outcome. Adjusting for the covariate as measured needs one minus the squared correlation of the unadjusted sample; cut at the median, one minus 2/π of the squared correlation. At a correlation of 0.7 those are 0.510 and 0.688; cut in three, 0.611; the change from baseline, 0.600. The change score needs fewer patients than the median-cut adjustment once the correlation passes 0.624.

A baseline cut in two

Adjusting a trial's result for a baseline measurement that predicts the outcome shrinks the sample it needs to 1 − ρ² of the unadjusted one. Adjusting for whether that measurement was above its median shrinks it only to 1 − (2/π)ρ² — the cut keeps 63.7% of what the covariate could remove, exactly the fraction a cut outcome keeps. In patients the loss grows with the covariate's strength: a cut costs 12% more patients at a correlation of 0.5, 35% at 0.7 and 2.55 times as many at 0.9, where a simple change from baseline would have done better than the cut adjustment.

5 figures · Power, part 7
What four adjusted analyses report about their own precision and what they have, 50 patients an arm, baseline correlated 0.7. Reported against actual variance of the treatment estimate, as shares of the unadjusted analysis's, over 20,000 trials of 50 patients an arm. Adjusted for the baseline as measured: 0.529 reported, 0.518 actual. Median cut fixed in advance: 0.716 and 0.700. The decile cut with the best fit: 0.671 and 0.736. The decile cut with the smallest p-value: 0.762 and 1.615.

The cut that fitted best

A trial that adjusts for its baseline cut at whichever of the sample's nine deciles fits the outcomes best — without ever looking at the treatment difference — reports a variance 0.671 of the unadjusted analysis's where its estimates actually have 0.736, which is worse than the 0.700 a median fixed in advance delivers. The sample says the chosen cut kept 72.4% of what the baseline could remove; the population says it kept less than a median. A true null is rejected 5.87% of the time at fifty patients an arm and 9.31% at ten, and a cut chosen for its p-value rejects 14.22%.

6 figures · Power, part 8
The variance of the baseline-adjusted estimate for four arrangements of visits, against the correlation between visits. At a correlation of 0.5 between visits, the adjusted analysis has 0.750 of the unadjusted variance with one visit before treatment and one after, 0.667 with two before, 0.500 with two after and 0.417 with two of each.

A visit before or a visit after

A trial that can afford one more measurement per patient can take it before treatment, to sharpen the baseline adjustment, or after it, to average the outcome. With every pair of visits correlated 0.5, the second visit after treatment removes a quarter of the variance and the second visit before it removes a twelfth, and it is never the other way round: the baseline's contribution is ρ²(1 − ρ)/(1 + ρ), which peaks at ρ⁵ = 0.090 when ρ is the golden ratio's reciprocal. Below a correlation of one half, a single extra follow-up beats any number of baselines. When the correlation fades with the time between visits, averaging an earlier baseline into the adjustment makes the trial less precise, not more.

5 figures · Power, part 9

What makes it checkable

A stated seed for every random draw, so a simulated figure comes out the same every time it is drawn. A closed form beside every simulation, so there are two routes to each number. And the simulation's own error counted as carefully as the method's: a table of correct cells that flags one of them most of the time, two methods whose comparison is sharper or noisier for sharing their draws, a run that stops when it looks done and reports what it was looking for, and draws aimed at a tail that look settled while their variance is infinite.

Twenty samples of 40, every one of them genuinely normal. Each panel is a quantile-quantile plot of 40 draws from a normal distribution. The worst point in the worst panel sits 0.87 standard deviations off the line. Anything a reader would reject here would be a false alarm.

The seed is part of the figure

Every other site in this fleet draws from a deterministic rule, so a figure either is or is not what it claims. Here the figures are samples, and a sample can be right by luck. That changes what a figure has to carry.

5 figures · Seeds, part 1
Coverage of four nominal 95% intervals, n = 30. Computed exactly by summing over all 31 possible counts, not simulated. The Wald interval drops to 26.0% and is jagged everywhere; Clopper–Pearson never falls below 95% and pays for it in width.

Two routes to every number

A site about probability that only simulates has one route to each answer and no way to tell a right one from a plausible one. Every important number here is computed twice, by arithmetic that shares nothing, and the two are required to agree.

4 figures · Routes, part 2
Twenty cells of an interval that is exactly 95%, 1,000 replications each. The t interval covers exactly 95% in every cell. Estimated at 1,000 replications its cells read 93.9% to 96.5%, and 2 of the twenty are flagged by their own ±1.96 standard errors.

A coverage table with its own error

Twenty cells estimating the coverage of an interval that is exactly 95%, at a thousand replications each, read from 93.9% to 96.5% — and a table like that flags at least one of its correct cells on 69.9% of honest runs. Ten times the replications does not repair it: at ten thousand the same table still flags one 63.3% of the time.

6 figures · Seeds, part 2
Which samples Wilson and Clopper–Pearson each cover, n = 50, p = 0.2. Each bar is the probability of one count, shaded by which interval built on that count contains 0.2. Both cover 95.1% of samples, only Wilson 0.0%, only Clopper–Pearson 1.6%, neither 3.3%. The correlation between their hits is 0.810, so on shared draws the variance of their difference is 4.891 times smaller than on independent ones.

The same draws for both methods

Two intervals computed on the same simulated datasets give a difference in coverage whose variance can be 4.891 times smaller than on separate datasets — or, for a pair that covers different samples, 1.164 times larger. Which one a comparison gets is an exact sum over the counts each interval covers, and a standard error that ignores the sharing covers 100.00% for one pair and 93.07% for the other.

7 figures · Seeds, part 3
Twenty runs simulating an exactly 95% interval, checked every 250 replications. Each line is one run's running estimate; the dashed band is where the Wilson interval of the running estimate still contains 95%, and a run stops, marked, the first time it leaves the band. 8 of these twenty stop before 10,000 replications. The exact probability of stopping, from the recursion over the count, is 29.54%.

A simulation that stops when it looks settled

A simulation of an interval that covers exactly 95%, checked every 250 replications for a significant departure and stopped when it finds one, flags that correct interval on 29.54% of runs. Stopped instead as soon as its estimate reaches 95%, it reports an interval that covers 94% as meeting its level on 37.21% of runs. Stopped when the estimate stops moving, it reports the right number — and has quietly chosen to run about fifteen hundred replications.

6 figures · Seeds, part 4
Estimating P(Z > 5) = 2.8665×10⁻⁷ with plain draws and with four proposals. One seed each. Plain simulation draws nothing past 5 in 100,000 and estimates zero throughout. At 100,000 draws the proposal N(5, 1) reads 1.009 of the truth, N(9, 1) 1.041, N(4.5, 0.25²) 1.039 and N(5, 0.3²) 1.008. Values above 2.2 are drawn at the top edge.

The draws aimed at the tail

The chance a standard normal exceeds 5 is 2.8665×10⁻⁷, and a plain simulation needs 349 million draws to estimate it to within ten per cent. Draws aimed at the tail and weighted back need 565. Aimed slightly too narrowly, the same method has an infinite variance, an interval that covers 86.0% and gets worse with more draws, and an effective sample size that reads healthier than a proposal that works.

5 figures · Seeds, part 5
Two companions on one simulation, two hundredfold apart. How many times as many draws each companion is worth, on the same 4,000 simulated samples of 40 observations. The coverage of the interval is estimated with the observed count as its companion, whose expectation is 12 exactly; they correlate at 0.2665 and the companion is worth 1.08 times the draws. The expected width is estimated with p̂(1 − p̂) as its companion, whose expectation is 0.20475 exactly; they correlate at 0.9977 because the width is a monotone function of it, and the companion is worth 214 times the draws — 856 thousand simulated samples' worth of precision from four thousand.

The check worth more than the check

The same exactly known companion that verifies a simulation can sharpen it. On one set of four thousand draws, one companion is worth 1.08 times the draws and another is worth 214 times them, and the factor is 1 − ρ² with nothing else in it.

3 figures · Routes, part 4

FieldsThreadsSeriesConceptsFigure librarySearch