Every essay — page 22
What partial pooling does to one group, to the set, and to a ranking
Partial pooling minimises the total squared error across groups, and three uses of its estimates are not the total. One group: with each group's standard error equal to the population's spread, every group truly more than 1.73 widths from the centre — 8.33% of a normal population — is estimated worse than by its own mean, and among eight groups with the spread estimated the truly most extreme one is worse off in 61.6% of datasets; capping each group's shift at one standard error improves the total, the truly extreme group and the observed extreme group at once. The set: posterior means spread 0.707 as widely as the truth and count 0.234% of groups beyond two widths against a true 2.28%. The ranking: in a league table of a hundred groups of unequal size, raw means fill 62.1% of the top ten with small groups, posterior means 13.4%, and no ranking recovers more than 5.47 of the true ten.
What a sample-size calculation was given
A sample-size calculation takes a standard deviation, an effect and an outcome as given, and each is usually a choice or an estimate. Sized from a pilot of ten's standard deviation, 55.9% of trials have less than their planned 80% power and 11.1% less than half, although the planned sample is right on average; sizing from the pilot's 80% upper confidence limit leaves 19.8% short at 1.65 times the patients. With the effect believed to be half a standard deviation give or take a quarter, a trial of sixty-four per arm succeeds 69.2% of the time, and no trial of any size can succeed more often than the planners believe the effect is positive. Cutting the outcome at a threshold keeps at most 63.7% of its information, and a responder threshold one standard deviation out needs twice the patients.
The spread a pilot supplies
A trial sized for 80% power from a pilot's standard deviation is sized from an estimate that is too small more often than not. With a pilot of ten, 55.9% of the trials it sizes have less than 80% power and 11.1% less than 50%, although the planned sample is right on average. Sizing from the pilot's 80% upper confidence limit instead leaves 19.8% short, at 1.65 times the sample; from its 90% limit, 10.0% short at 2.12 times.
The chance a trial succeeds
A trial of sixty-four per arm has 80% power at an effect of half a standard deviation. If the effect is only believed to be about half a standard deviation, give or take a quarter, the chance the trial reaches significance is 69.2%; give or take a half, 61.4%. Reaching 80% then takes 113 per arm, or 1,268 — and when the belief is uncertain by three quarters of a standard deviation no number of patients reaches 80%, because the chance can never exceed the 74.8% probability that the effect is positive at all.
An outcome cut in two
Replacing a measured outcome with whether it crossed a threshold keeps 63.7% of the information when the cut is at the mean, 34.2% at the top tenth and 13.1% two standard deviations out. A trial that needs 63 patients per arm on the measured outcome needs 102 cut at the mean and 185 cut at one and a half standard deviations. The responder rates that result read as a share of patients who respond — 50.0% against 69.1% — when every patient moved by the same amount; and a cut chosen after looking turns a 5% test into a 17.7% one.
A baseline cut in two
Adjusting a trial's result for a baseline measurement that predicts the outcome shrinks the sample it needs to 1 − ρ² of the unadjusted one. Adjusting for whether that measurement was above its median shrinks it only to 1 − (2/π)ρ² — the cut keeps 63.7% of what the covariate could remove, exactly the fraction a cut outcome keeps. In patients the loss grows with the covariate's strength: a cut costs 12% more patients at a correlation of 0.5, 35% at 0.7 and 2.55 times as many at 0.9, where a simple change from baseline would have done better than the cut adjustment.
The cut that fitted best
A trial that adjusts for its baseline cut at whichever of the sample's nine deciles fits the outcomes best — without ever looking at the treatment difference — reports a variance 0.671 of the unadjusted analysis's where its estimates actually have 0.736, which is worse than the 0.700 a median fixed in advance delivers. The sample says the chosen cut kept 72.4% of what the baseline could remove; the population says it kept less than a median. A true null is rejected 5.87% of the time at fifty patients an arm and 9.31% at ten, and a cut chosen for its p-value rejects 14.22%.
A visit before or a visit after
A trial that can afford one more measurement per patient can take it before treatment, to sharpen the baseline adjustment, or after it, to average the outcome. With every pair of visits correlated 0.5, the second visit after treatment removes a quarter of the variance and the second visit before it removes a twelfth, and it is never the other way round: the baseline's contribution is ρ²(1 − ρ)/(1 + ρ), which peaks at ρ⁵ = 0.090 when ρ is the golden ratio's reciprocal. Below a correlation of one half, a single extra follow-up beats any number of baselines. When the correlation fades with the time between visits, averaging an earlier baseline into the adjustment makes the trial less precise, not more.
The correlation a pilot can see
Two models of repeated visits agree that adjacent visits correlate 0.7 and disagree about how many follow-ups a trial should buy: three under equal correlation, one when the correlation fades. Planning on the wrong one costs 28.0% or 9.9%. Two follow-ups cost at most 3.2% more than the best under either, and at most 11.1% under any curve with that adjacent correlation. A pilot of twenty patients at four visits chooses worse on average than that rule when the correlation fades; it takes about forty to draw level with it, and at an effect of half a standard deviation those forty patients' extra visits cost fifty times what the better schedule saves.
A schedule the trial chooses for itself
A separate pilot cannot pay for itself in choosing a trial's follow-up schedule, because its patients are recruited, measured and discarded. A trial can instead measure its own first patients on the schedule that needs no pilot — two follow-ups — fit the correlation curve to them, and give everyone after them the schedule that curve makes cheapest. Those first patients are counted in the analysis and cost no extra visit. Forty of them cut the trial's worst-case excess cost across every curve with adjacent correlation 0.7 from 11.1% to 5.9% in a trial of two hundred an arm, and eighty cut it to 3.6%; twenty buy almost nothing. And because the curve is read from within-arm deviations, the choice cannot see the effect: the pilot's effect estimate is the same whichever schedule it chose.
What makes it checkable
A stated seed for every random draw, so a simulated figure comes out the same every time it is drawn. A closed form beside every simulation, so there are two routes to each number. And the simulation's own error counted as carefully as the method's: a table of correct cells that flags one of them most of the time, two methods whose comparison is sharper or noisier for sharing their draws, a run that stops when it looks done and reports what it was looking for, and draws aimed at a tail that look settled while their variance is infinite.
The seed is part of the figure
Every other site in this fleet draws from a deterministic rule, so a figure either is or is not what it claims. Here the figures are samples, and a sample can be right by luck. That changes what a figure has to carry.
Two routes to every number
A site about probability that only simulates has one route to each answer and no way to tell a right one from a plausible one. Every important number here is computed twice, by arithmetic that shares nothing, and the two are required to agree.
A coverage table with its own error
Twenty cells estimating the coverage of an interval that is exactly 95%, at a thousand replications each, read from 93.9% to 96.5% — and a table like that flags at least one of its correct cells on 69.9% of honest runs. Ten times the replications does not repair it: at ten thousand the same table still flags one 63.3% of the time.
The same draws for both methods
Two intervals computed on the same simulated datasets give a difference in coverage whose variance can be 4.891 times smaller than on separate datasets — or, for a pair that covers different samples, 1.164 times larger. Which one a comparison gets is an exact sum over the counts each interval covers, and a standard error that ignores the sharing covers 100.00% for one pair and 93.07% for the other.
A simulation that stops when it looks settled
A simulation of an interval that covers exactly 95%, checked every 250 replications for a significant departure and stopped when it finds one, flags that correct interval on 29.54% of runs. Stopped instead as soon as its estimate reaches 95%, it reports an interval that covers 94% as meeting its level on 37.21% of runs. Stopped when the estimate stops moving, it reports the right number — and has quietly chosen to run about fifteen hundred replications.
The draws aimed at the tail
The chance a standard normal exceeds 5 is 2.8665×10⁻⁷, and a plain simulation needs 349 million draws to estimate it to within ten per cent. Draws aimed at the tail and weighted back need 565. Aimed slightly too narrowly, the same method has an infinite variance, an interval that covers 86.0% and gets worse with more draws, and an effective sample size that reads healthier than a proposal that works.
The check worth more than the check
The same exactly known companion that verifies a simulation can sharpen it. On one set of four thousand draws, one companion is worth 1.08 times the draws and another is worth 214 times them, and the factor is 1 − ρ² with nothing else in it.
A proposal fitted to its own draws
The cross-entropy method fits an importance-sampling proposal from its own pilot draws, and asked to fit a normal's mean and spread to P(Z > 5) it fails before the question of variance arises: the spread halves at every stage, the level stalls near 3.7, and 95% of runs never reach the threshold, so the interval covers 1.1% of the time. Its ideal end point, the normal closest to Z conditioned past 5, is N(5.187, 0.181²), which has an infinite variance inside its own draws' reach and covers 93.7% at a thousand draws and 90.8% at ten thousand. Fitting the mean alone lands within a tenth of the best shift and covers 94.5%.