Every essay — page 18
Two tests, a threshold, and the rate they are read against
A predictive value is computed from a sensitivity, a specificity and a prevalence, and each of the three is softer than it is quoted as being. Two positives from a 90/95 test on a one-in-a-thousand condition are worth 24.49% if the tests' errors are independent, 10.16% at a correlation of 0.1, and 3.16% at 0.5 — barely more than the 1.77% one positive was worth. The sensitivity and specificity are not two properties of a test but one property read at a threshold somebody chose, and the threshold that minimises harm moves by three standard deviations of the score across the prevalences such a test is used at. The prevalence itself is usually estimated from the same test's positive rate, which at one in a thousand reads 5.09%.
The second test that is not a second opinion
Two positives from a 90/95 test on a one-in-a-thousand condition give a 24.49% chance of disease if the tests are independent. At a correlation of 0.1 between their errors it is 10.16%, and at 0.5 it is 3.16% — barely more than the 1.77% one positive was worth.
The test is a point somebody chose
A test reported as 90% sensitive and 95% specific is not two properties of a test. It is one property read at a threshold, and the threshold that minimises harm runs from 3.05 standard deviations of the score at a prevalence of one in ten thousand to −0.12 at one in two — 45% of cases detected at one end and 99.9% at the other.
The prevalence the test has to estimate
Every predictive value takes a prevalence as given, and the prevalence is usually estimated from the same test's positive rate. At a true prevalence of one in a thousand that rate reads 5.09% — fifty times the truth — and the correction that inverts it is unbiased, 18% more variable, and negative on 48.6% of samples of a thousand.
What a diagnostic plot is showing
A quantile plot is read by eye against a band drawn for one point at a time, and a genuinely normal sample of forty has forty chances to leave it: 45.0% of them do. The band a reader is actually using is that one widened by a factor of 1.502, and nothing draws it. The quantity is wrong as well as the width — a t interval needs the sampling distribution of the mean to be normal rather than the data, and the source that leaves the band on every sample of forty covers 94.80% while one that leaves it on 57% covers 95.73%. And what is plotted are not the errors: a design whose leverages run from 0.045 to 0.663 gives residuals whose spreads differ by a factor of 1.68 with the model exactly right.
The band the eye was standing in for
The confidence band software draws on a quantile plot holds each point at 95%, and a genuinely normal sample of forty has forty chances to leave it — so 45.0% of them do. The band a reader is actually using is that one widened by a factor of 1.502, and nothing draws it.
The plot is about the wrong quantity
A t interval needs the sampling distribution of the mean to be normal, not the data. A two-lump source leaves its quantile band on 100% of samples of forty and its interval covers 94.80%; a t on three degrees of freedom leaves it on 57% and covers 95.73%, the best of five sources.
Residuals are not the errors
A residual's standard deviation is σ√(1 − hᵢᵢ), so a design whose leverages run from 0.045 to 0.663 produces residuals whose spreads differ by a factor of 1.68 with the model exactly right. On the samples where the high-leverage point really did have the largest error, a raw residual plot shows it as the largest on 0.0% of them.
Past the first term of the normal approximation
The Berry–Esseen bound guarantees how far a standardised sum can be from the normal, and on an exponential source it is 8.62 times the real worst error at every sample size — a constant set by a fair coin, whose error is a half-step, and at a hundred draws larger than the 2.5% tail it would be asked to certify. One Edgeworth term takes the normal's error at two standard deviations from 38% to 8% on ten draws, and turns negative in the short tail at every sample size, beyond a threshold that recedes only as the sixth root of n. The saddlepoint approximation is built at the threshold rather than at the mean, and reads a tail six standard deviations out to within 0.19% on five draws and within 2.2% on one.
A bound written for a coin
The Berry–Esseen theorem guarantees how far a standardised sum can be from the normal, and the guarantee is true. On an exponential source it is 8.62 times the real worst error at every sample size, the worst error sits at the centre rather than in a tail, and at a hundred draws the bound is larger than the 2.5% tail it would be asked to vouch for.
A correction that goes below zero
One Edgeworth term takes the normal approximation's error at two standard deviations from 38% to 8% on ten exponential draws, and stretches the range within 10% of the truth from 1.66 to 3.09 standard deviations at a hundred. It also turns negative in the short tail at every sample size — past 3.13 standard deviations at a hundred draws and 9.83 at a hundred thousand — because the region recedes only as the sixth root of n.
An approximation built at the threshold
The saddlepoint approximation reads the tail of a sum of five exponential draws to within 0.19% six standard deviations out, where the normal is short by a factor of more than sixty thousand. It is within 2.2% out to ten standard deviations on a single draw, where there is nothing to average, and within 1.1% on a binomial whose expected count is one. It works because it is built where the tail is read rather than at the mean.
A proportion's interval near the boundary, and the coin
Wilson's interval for a proportion wobbles a point or two around 95% away from the boundary and has a hole near it: its worst coverage is 83.50% at ten trials and 83.81% at a thousand, at an expected count of 0.1765 each time, because below that count the interval built on one success lies wholly above the truth. The intervals that never fall below 95% pay in average coverage, and the two of them pay for different promises — Clopper–Pearson holds each side under 2.5% and is 10.4% wider than Wilson at thirty trials, Blaker holds only the total and is 4.6% wider. Only an interval that adds a uniform random draw to the count covers exactly 95% at every proportion, at 0.9% more width than Wilson, and it is not used because two analysts with the same data would report different intervals.
A hole no sample size fills
Wilson's interval is the recommended repair for a proportion, and away from the boundary it wobbles a point or two around 95%. Near zero it has a hole: at an expected count of 0.1765 its coverage is 83.50% at ten trials, 83.79% at a hundred and 83.81% at a thousand, and it never climbs past e to the minus 0.1765, which is 83.82%. The hole is where the interval built on one success stops containing the truth, and it belongs to the count rather than to the sample size.
What a guaranteed minimum costs
Clopper–Pearson's interval never covers less than 95%, and at thirty trials it averages 97.34% and is 10.4% wider than Wilson's. Blaker's interval keeps the same guarantee, averages 96.31% and is 4.6% wider. The difference is not waste: Clopper–Pearson guarantees each side separately, holding both below 2.5%, and Blaker guarantees only their sum — so at ten trials and a proportion of 0.15 it misses on one side 5.00% of the time.
The coin that makes it exact
Every interval for a proportion either covers less than 95% somewhere or more than 95% on average, because a count is discrete. One construction covers exactly 95% at every proportion: it adds a uniform random draw to the count. At thirty trials it is 0.9% wider than Wilson's interval and narrower than both exact ones — and two analysts with the same data report different intervals, and one study in forty that sees nothing reports an empty one.
An interval read beside something else
A 95% interval is almost never read alone. Read against a replication's estimate it captures an equal-sized one 83.42% of the time, because both estimates are uncertain, and ten replications of one study miss together: none of them lands outside 31.40% of the time against 16.32% if they were independent. Read against a second interval, two that just touch mark a difference with p = 0.0056 rather than 0.05, standard-error bars that touch mark p = 0.157, and in a figure of ten identical groups some pair of 95% intervals fails to overlap one time in seven. Read against a second verdict, two studies of one effect at 50% power disagree about significance half the time, and when they do the difference between them is significant in 9.75% of cases.
Five times in six
A 95% interval is read as a 95% chance that a replication's estimate will land inside it. With the spread known and a replication of the same size, the chance is 83.42% — five times in six — because both estimates are uncertain. An original that landed two standard errors from the truth captures a replication 48.40% of the time; a replication a tenth the size lands inside 44.54% of the time; and among significant originals from studies with 17% power, 66.94%.
Two intervals that overlap
Two 95% intervals that just touch are read as a difference at the edge of significance. With equal standard errors their difference has p = 0.0056, not 0.05; two intervals can overlap by 58.6% of an arm and still differ at exactly 5%; standard-error bars that just touch mark p = 0.157; and when the two estimates are correlated at 0.8, touching intervals conceal a difference of 6.2 standard errors. Read as a test, non-overlap needs 1.66 times the sample for the same power.
Significant in one, not in the other
Two studies of exactly the same effect, each with 50% power, disagree about significance half the time — and when they do, the test of the difference between them is significant in 9.75% of cases. A p of 0.01 beside a p of 0.20 is a difference with p = 0.36. Among four subgroups sharing one effect, at least one significant and one not happens 87.5% of the time, and the test that would tell a real difference apart needs four times the sample the effect itself needed.
What partial pooling does to one group, to the set, and to a ranking
Partial pooling minimises the total squared error across groups, and three uses of its estimates are not the total. One group: with each group's standard error equal to the population's spread, every group truly more than 1.73 widths from the centre — 8.33% of a normal population — is estimated worse than by its own mean, and among eight groups with the spread estimated the truly most extreme one is worse off in 61.6% of datasets; capping each group's shift at one standard error improves the total, the truly extreme group and the observed extreme group at once. The set: posterior means spread 0.707 as widely as the truth and count 0.234% of groups beyond two widths against a true 2.28%. The ranking: in a league table of a hundred groups of unequal size, raw means fill 62.1% of the top ten with small groups, posterior means 13.4%, and no ranking recovers more than 5.47 of the true ten.
A group from the population's own tail
Partial pooling halves the total squared error when a group's own standard error equals the spread between groups. Every group whose true effect sits more than 1.73 population widths from the centre — 8.33% of a perfectly normal population — does worse than it would have with its own mean, and its loss grows without bound. Among eight groups with the spread estimated, the most extreme is worse off in 61.6% of datasets. Capping the shift at one standard error keeps the total at 0.528 of the unpooled error and holds every group under twice it.
Estimates that are too alike
Posterior means give each group its least-error estimate, and as a set they are too alike: with each group's standard error equal to the population's spread, they spread 0.707 as widely as the truth. Beyond two population widths lie 2.28% of the true effects, 7.86% of the groups' own means, and 0.234% of the posterior means — a tenth of the truth. Rescaling the estimates to the right spread counts the tail exactly and costs 17% more squared error; summing each group's posterior chance of being beyond the line counts it without changing any estimate.
A league table of a hundred
A hundred groups with sizes from 4 to 400, and a top ten to publish. Ranked by their own means, small groups fill 62.1% of the top ten against their 36.3% share of the true top ten. Ranked by posterior means they fill 13.4%. The ranking built from each group's chance of being in the top ten recovers 5.47 of the true ten, the best of three and barely half; and the group ranked first could hold any rank from 1 to 31.
What a sample-size calculation was given
A sample-size calculation takes a standard deviation, an effect and an outcome as given, and each is usually a choice or an estimate. Sized from a pilot of ten's standard deviation, 55.9% of trials have less than their planned 80% power and 11.1% less than half, although the planned sample is right on average; sizing from the pilot's 80% upper confidence limit leaves 19.8% short at 1.65 times the patients. With the effect believed to be half a standard deviation give or take a quarter, a trial of sixty-four per arm succeeds 69.2% of the time, and no trial of any size can succeed more often than the planners believe the effect is positive. Cutting the outcome at a threshold keeps at most 63.7% of its information, and a responder threshold one standard deviation out needs twice the patients.
The spread a pilot supplies
A trial sized for 80% power from a pilot's standard deviation is sized from an estimate that is too small more often than not. With a pilot of ten, 55.9% of the trials it sizes have less than 80% power and 11.1% less than 50%, although the planned sample is right on average. Sizing from the pilot's 80% upper confidence limit instead leaves 19.8% short, at 1.65 times the sample; from its 90% limit, 10.0% short at 2.12 times.
The chance a trial succeeds
A trial of sixty-four per arm has 80% power at an effect of half a standard deviation. If the effect is only believed to be about half a standard deviation, give or take a quarter, the chance the trial reaches significance is 69.2%; give or take a half, 61.4%. Reaching 80% then takes 113 per arm, or 1,268 — and when the belief is uncertain by three quarters of a standard deviation no number of patients reaches 80%, because the chance can never exceed the 74.8% probability that the effect is positive at all.
An outcome cut in two
Replacing a measured outcome with whether it crossed a threshold keeps 63.7% of the information when the cut is at the mean, 34.2% at the top tenth and 13.1% two standard deviations out. A trial that needs 63 patients per arm on the measured outcome needs 102 cut at the mean and 185 cut at one and a half standard deviations. The responder rates that result read as a share of patients who respond — 50.0% against 69.1% — when every patient moved by the same amount; and a cut chosen after looking turns a 5% test into a 17.7% one.
What makes it checkable
A stated seed for every random draw, so a simulated figure comes out the same every time it is drawn. A closed form beside every simulation, so there are two routes to each number. And the simulation's own error counted as carefully as the method's: a table of correct cells that flags one of them most of the time, two methods whose comparison is sharper or noisier for sharing their draws, a run that stops when it looks done and reports what it was looking for, and draws aimed at a tail that look settled while their variance is infinite.
The seed is part of the figure
Every other site in this fleet draws from a deterministic rule, so a figure either is or is not what it claims. Here the figures are samples, and a sample can be right by luck. That changes what a figure has to carry.
Two routes to every number
A site about probability that only simulates has one route to each answer and no way to tell a right one from a plausible one. Every important number here is computed twice, by arithmetic that shares nothing, and the two are required to agree.
A coverage table with its own error
Twenty cells estimating the coverage of an interval that is exactly 95%, at a thousand replications each, read from 93.9% to 96.5% — and a table like that flags at least one of its correct cells on 69.9% of honest runs. Ten times the replications does not repair it: at ten thousand the same table still flags one 63.3% of the time.