an-interval-with-its-coverage-counted

Coverage of four nominal 95% intervals, n = 30

Computed exactly by summing over all 31 possible counts, not simulated. The Wald interval drops to 26.0% and is jagged everywhere; Clopper–Pearson never falls below 95% and pays for it in width.

Intervals, countedslider: sample size, 7 positionswide47 views

What else it draws

The same object, drawn to answer the other questions the essays put to it.

How many times as many draws each companion is worth, on the same 4,000 simulated samples of 40 observations. The coverage of the interval is estimated with the observed count as its companion, whose expectation is 12 exactly; they correlate at 0.2665 and the companion is worth 1.08 times the draws. The expected width is estimated with p̂(1 − p̂) as its companion, whose expectation is 0.20475 exactly; they correlate at 0.9977 because the width is a monotone function of it, and the companion is worth 214 times the draws — 856 thousand simulated samples' worth of precision from four thousand.

Computed exactly by summing over all 31 possible counts, not simulated. The Wald interval drops to 26.0% and is jagged everywhere; Clopper–Pearson never falls below 95% and pays for it in width.

Coverage does not improve monotonically. n = 19 covers 93.8% while the larger n = 20 covers 81.9%. The sample space is discrete, so the endpoints jump as n changes.

2 of the twenty miss the true value. The 95% is a property of the procedure across repetitions — no single interval has a 95% chance of anything, because it either contains 0.35 or it does not.

The Wald interval is the shortest and covers 94.2%. Clopper–Pearson covers 98.3% and is 13% wider. Shortness is not a virtue on its own — an interval of zero width is the shortest of all.

Measured over 20,000 samples. The t interval covers 94.9% and the z interval 90.7%. The difference is the price of pretending the standard deviation was known.

Uniform data on [0, 1]. For the mean the percentile bootstrap covers 93.5%. For the maximum it covers 0.0%, because a resample can never contain a value larger than the largest one observed, so the interval cannot reach above it.

The t interval covers exactly 95% in every cell. Estimated at 1,000 replications its cells read 93.9% to 96.5%, and 2 of the twenty are flagged by their own ±1.96 standard errors.

Bars: cells of 1,000 t intervals on five normal observations, each from its own seeds. Line: Binomial(1000, 0.95)/1000, computed. The counted spread is 1.038 times the binomial variance, and the chi-square across the binomial's eight bins is 5.83 on 7 degrees of freedom.

Computed exactly from the binomial: one correct cell of 1,000 replications is flagged with probability 5.83%, so twenty cells flag at least one 69.9% of the time, and a hundred 99.8%. More replications do not repair it — at 10,000 a cell twenty flag 63.3% — while reading each cell at Bonferroni's level holds twenty to 7.6%. The point with its bar is a hundred counted tables: 72%.

A cell's true coverage below 95% by δ points is detected 80% of the time, at a two-sided 5%, after ((1.96·√(0.95·0.05) + 0.84·√(p₁(1 − p₁)))/δ)² replications. A thousand replications see 2.03 points; ten thousand see 0.62. The marked points are checked by exact binomial summation.

Each bar is the probability of one count, shaded by which interval built on that count contains 0.2. Both cover 95.1% of samples, only Wilson 0.0%, only Clopper–Pearson 1.6%, neither 3.3%. The correlation between their hits is 0.810, so on shared draws the variance of their difference is 4.891 times smaller than on independent ones.

Each dot is one simulation's estimate of the difference in coverage, whose exact value is -1.636 points. On shared draws the estimates spread with standard deviation 0.398 points (exact 0.401); on independent draws 0.880 (exact 0.887). The bracket under each row is ±1 exact standard deviation.

Each point is one pair of the Wald, Wilson, Agresti–Coull and Clopper–Pearson intervals at one sample size (20 or 50) and one true proportion (0.05 to 0.50), computed exactly. 44 of the 120 settings are omitted because the two intervals cover exactly the same counts. The ratio ranges from 0.859 to 5.619; 5 settings sit below one, where shared draws make the comparison noisier. The curve is 1/(1 − ρ), which the ratio would equal if the two coverages had the same variance.

Dots are counted over 400 simulations with ±2 standard-error bars; ticks are exact. Wilson–Clopper–Pearson, paired, shared draws: 92.5% (exact 92.65%); Wilson–Clopper–Pearson, unpaired, shared draws: 100.0% (exact 100.00%); Wilson–Clopper–Pearson, unpaired, independent draws: 93.5%; Wilson–Wald, paired, shared draws: 93.3% (exact 94.97%); Wilson–Wald, unpaired, shared draws: 90.0% (exact 93.07%); Wilson–Wald, unpaired, independent draws: 96.8%.

Each line is one run's running estimate; the dashed band is where the Wilson interval of the running estimate still contains 95%, and a run stops, marked, the first time it leaves the band. 8 of these twenty stop before 10,000 replications. The exact probability of stopping, from the recursion over the count, is 29.54%.

Exact, from the recursion over the binomial count. Checked every 100 replications a correct 95% interval has been flagged 35.61% of the time by 10,000; every 250, 29.54%; every 1,000, 18.83%. The point at 10,000 is 4,000 counted runs checked every 250: 29.13%.

Exact. Stopping at the first of forty looks where the estimate reads at least 95%, an interval whose true coverage is 94% is reported as reaching 95% 37.21% of the time, and one at 93% 12.28%. Run to a fixed 10,000 replications, the 94% interval reads 95% with probability 0.0009%. The point at 94% is 4,000 counted runs: 37.43%.

Bars: the chance the run ends reporting a coverage of at least 95%. a fixed 10,000 replications: 0.0009%, reporting 94.00% on average (exact); a fixed 1,500: 5.45%, reporting 94.00% on average (exact); stopped when it settles: 6.15%, reporting 93.99% on average (counted, 1496 used on average); stopped when it reaches 95%: 37.21%, reporting 94.59% on average (exact).

Exact. Reading the Wilson interval at z = 2.7628 — 0.573% a look — holds forty looks to 4.990% for a correct interval, where 1.96 flags 29.54%. Against an interval covering 94.5% the boundary detects 46.84%, a single analysis at 10,000 detects 62.67%, and the naive monitor 78.97% — most of it the same false alarms.

One seed each. Plain simulation draws nothing past 5 in 100,000 and estimates zero throughout. At 100,000 draws the proposal N(5, 1) reads 1.009 of the truth, N(9, 1) 1.041, N(4.5, 0.25²) 1.039 and N(5, 0.3²) 1.008. Values above 2.2 are drawn at the top edge.

The relative error per draw of importance sampling for P(Z > 5) with a N(m, 1) proposal, √(exp(m²)·Q(5 + m)/p² − 1), in closed form. It is smallest, 2.376, at m = 5.097, where 565 draws reach a 10% relative error. At m = 4 it is 3.347, at 7 6.343, at 8 21.52 and at 9 119.5. Plain simulation, m = 0, is 1868. The point at 5 is counted from 600 runs of 10,000 draws.

Each curve integrates φ(x)²/q(x) from 5 to the typical largest of R draws, exactly, and reports the relative error per draw it implies. N(4.5, 0.25²) and N(5, 0.3²) both have infinite second moments. The first one's visible error rises from 5.456 at 1,000 draws to 9.488 at 100,000; the second one's moves from 1.135 to 1.148, and does not begin to climb until well past any run counted here. N(3, 1) and N(5, 1) are finite and settle at 7.768 and 2.383. Points: the same quantity read off each run's own sample variance, the median over runs — N(4.5, 0.25²) at 1,000: 6.81, N(4.5, 0.25²) at 10,000: 8.67, N(4.5, 0.25²) at 100,000: 10.83, N(5, 0.3²) at 1,000: 1.15, N(5, 0.3²) at 10,000: 1.15, N(5, 0.3²) at 100,000: 1.15.

Counted over 2,000 runs at 1,000 draws, 600 at 10,000 and 150 at 100,000, with ±2 standard-error bars. N(5, 1): 94.3%, 94.7%; N(3, 1): 92.7%, 93.3%; N(4.5, 0.25²): 85.2%, 86.0%, 81.3%; N(5, 0.3²): 94.3%, 95.0%, 95.3%; N(9, 1): 19.3%, 54.5%. The last is below the plotted range.

For each proposal: the share of its draws landing past 5, the effective sample size of its weights as a share of all draws and of the landed draws, and the counted coverage of its nominal interval. N(5, 1): 50.00% landed, 14.96% and 29.9%, coverage 94.7%; N(5, 0.3²): 50.00% landed, 43.02% and 86.3%, coverage 95.0%; N(3, 1): 2.28% landed, 1.62% and 71.6%, coverage 93.3%; N(4.5, 0.25²): 2.28% landed, 1.20% and 52.9%, coverage 86.0%; N(9, 1): 100.00% landed, 0.02% and 0.0%, coverage 54.5%.

Read against the expected number of successes the four sample sizes draw the same curve near the boundary. The worst coverage is 83.50% at n = 10, 83.71% at n = 30, 83.79% at n = 100, 83.81% at n = 1000, each at an expected count near 0.177, and the limiting depth is e^(−0.1765) = 83.82%.

The coverage is the probability of the counts whose intervals contain the true proportion. Just below an expected count of 0.1765, 0 covers, and the coverage is 84.10%.

More trials do not bring the worst coverage of the first three to 95%. At 1,000 trials: Wilson 83.81%, Jeffreys 89.77%, Wilson, repaired near zero 92.72%, Agresti–Coull 94.36%. Wilson's approaches e^(−0.1765) = 83.82% and stops there.

The worst coverage over every proportion beside the average over a uniform one, with the average expected width. Clopper–Pearson: worst 95.05%, average 97.34%, width 0.299. Blaker: worst 95.00%, average 96.31%, width 0.283. Wilson: worst 83.71%, average 95.24%, width 0.271.

Both exact intervals stay at or above 95% everywhere; Clopper–Pearson sits higher. Wilson crosses the line in both directions. Coverage below 90% is drawn at the floor of the axis.

The chance the whole interval lies below the truth and the chance it lies above, at every proportion. The worst below is 2.50% and the worst above is 2.50%. Clopper–Pearson guarantees each side separately.

Clopper–Pearson never falls below 95% and runs up to 99.80%. The mid-p interval, which is the randomised interval with its coin fixed at one half, runs from 92.94% to 99.80%. The randomised interval covers 95% at every proportion, to within the 0.043% of the numerical integration over the coin.

The same data give a lower end anywhere from 0.0564 to 0.0323 and an upper end from 0.3783 to 0.3187, depending on a uniform draw nobody observed. The coin at one half gives the mid-p interval, 0.0396 to 0.3561. Clopper–Pearson's 0.0321 to 0.3789 contains every one of them.

Averaged over a uniform proportion. Wilson 0.2708, worst coverage 83.71%; randomised 0.2732, worst coverage 95.00%; mid-p 0.2759, worst coverage 92.45%; Blaker 0.2833, worst coverage 95.00%; Clopper–Pearson 0.2990, worst coverage 95.05%. The randomised interval covers exactly 95% at every proportion and is 0.9% wider than Wilson's.

The chance that a 95% interval contains a replication's estimate is 95.00% when the original landed on the truth, 82.99% one standard error away and 48.40% two away. Averaged over where originals land it is 83.42%, and 5.00% of originals capture a replication less than half the time.

A replication a tenth of the size lands inside 44.54% of the time, one of the same size 83.42%, and one ten times the size 93.83%. Only an infinitely large replication reaches 95%.

Only originals that reached p < 0.05 are kept. At 8% power their intervals capture a replication's estimate 52.49% of the time; at 52% power, 83.78%. Unselected intervals capture 83.42% whatever the power.

The replications share the original, so they miss together. None of 10 lands outside with probability 31.40%, against 16.32% if they were independent; at least half land outside with probability 8.72%, against 1.51%.

The intervals are separated, and the test of the difference gives p = 0.0056. Two 95% intervals with equal standard errors just touch at p = 0.0056.

Two 95% intervals with equal standard errors touch at p = 0.0056, and at a ratio of ten at p = 0.0319. Two ±1 standard error bars touch at p = 0.157 and 0.274. Neither kind of bar marks 5%.

For independent estimates with equal standard errors, touching intervals mark a difference of 2.77 standard errors. At a correlation of 0.5, 3.92; at 0.9, 8.77 — a p-value below 6e-10 already at 0.8.

At no true difference the ordinary test rejects 5.00% of the time and non-overlap occurs 0.56% of the time. At a true difference of 2.8 standard errors — 80% power for the ordinary test — the intervals fail to overlap 51.1% of the time.

Every group has the same true mean and the same standard error. With ten groups, some pair of 95% intervals fails to overlap 14.61% of the time, some pair of 83.4% intervals 62.74%, and some pair of ±1 standard error bars 92.32%.

Five hundred pairs. With the true effect identical in both, exactly one of the two is significant in 50.0% of pairs; among those, the difference between the two is significant in 9.7%. The dashed lines mark 1.96 on each axis; the diagonal lines mark a significant difference.

Two studies split — one significant, one not — with probability 2π(1 − π), peaking at 50% at half power. Among splits the difference is significant 9.8% of the time near that peak. With 2 subgroups of the same effect, at least one is significant and at least one is not with probability up to 49.9%.

At the sample that gives 80% power for the overall effect, a difference between the halves as large as the overall effect is detected 28.8% of the time, and one half as large 10.8%. Detecting them at 80% needs 4 and 16 times the sample.

Where it is used

32 essays draw this figure, each at the numbers its own argument is about, so the same picture answers 32 different questions.

All 80 figures

FieldsThreadsSeriesConceptsAll essaysSearch