Concept

Posterior — where it appears

The distribution of the parameters given the data, which is what Bayes' rule produces from a prior and a likelihood. Its shape is what a credible interval is read off, and it depends on the prior by an amount that shrinks with the data at a rate worth measuring rather than assuming.

Named by 16 essays across 7 fields — each of them below, with the objects they name alongside it.

weakly informative — Beta(2, 2), updated by 5 of 20. The prior is worth 4 observations. With 20 observations the posterior mean is 0.292, against a data proportion of 0.250 and a prior mean of 0.500.

What a prior is worth

A prior is not a philosophical position, it is a component with a stated size. For a proportion it is worth exactly a + b observations, which turns "how much does the prior matter" from an argument into a subtraction.

bayes · Prior
What eight groups say about τ, when the truth is 1. The posterior density for the population spread after eight groups whose standard errors run from 0.5 to 2.1. The shaded band is the central 95% interval, from 1.02 to 4.44; the posterior median is 1.99 and the mean 2.18. The vertical mark at 1.07 is the moment estimate that empirical Bayes substitutes and then treats as known.

What the plug-in forgets

The shrinkage weight needs a population spread, and the population spread has to be estimated from eight numbers. Empirical Bayes estimates it, substitutes it, and proceeds as though it were known — and the interval that comes out covers 79% rather than the 95% it claims.

fullbayes · Shrinkage
Testing at 0.05 every time the data is looked at. The null is true in every one of these trials and the test is correct every time it is run. Looking once rejects 4.9% of the time, as it should; looking ten times rejects 19.2% of the time. Nothing changed except permission to look.

When the looking happens

A p-value is defined relative to a sampling plan, so the same data means different things under different stopping rules. Testing five times at the nominal level rejects a true null 14% of the time, and no observation in the dataset changed.

sequential · Stopping
What a second positive is worth, prevalence 0.10%. A 90% sensitive, 95% specific test. One positive gives 1.77%. Two independent positives give 24.49%, which is what multiplying the likelihood ratios says. At a correlation of 0.1 between the tests' errors it is 10.16%, and at 0.5 it is 3.16%.

The second test that is not a second opinion

Two positives from a 90/95 test on a one-in-a-thousand condition give a 24.49% chance of disease if the tests are independent. At a correlation of 0.1 between their errors it is 10.16%, and at 0.5 it is 3.16% — barely more than the 1.77% one positive was worth.

screening · Baserate
Three priors on the spread, at a true τ of 0.5. The posterior for τ under a flat prior (mean 1.66), a half-Cauchy of scale 1 (1.32) and one of scale 0.25 (1.15). The three answers differ by 30% of the widest. The prior does visible work when eight groups cannot separate a small spread from none, and almost none when they can.

A prior on the spread

Integrating over the population spread means putting a prior on it, which sounds like the objection rather than the repair. The prior's effect is measurable, it is invisible where the groups are clearly different, and the reflex choice for a scale parameter turns out not to have a posterior at all.

fullbayes · Prior
Ten groups of 10, pooled on the log-odds scale. Each row is a group. The hollow circle is its own proportion, the filled one is the estimate after pooling, and the vertical rule is the pooled population proportion of 31.3%. One group saw no events at all, and its raw proportion of zero becomes 21.2% — an estimate the group's own data cannot produce and the population's can. The arrows are not the same length, and none of the groups differs in size.

Pooling a proportion

A proportion cannot be shrunk on its own scale — an estimate would leave the interval, and how much information a count carries depends on where it sits. Move to log-odds and the approximation works, at the price of a group that saw nothing having no estimate at all until the correction supplies one.

multilevel · Levels
The weight on the population, σ = 3. Each curve is one population spread τ. A group's estimate moves B = se²/(se² + τ²) of the way to the population mean, where se = σ/√n is what the group's own mean does not know. At τ = 1 a group of 9 observations sits halfway.

The weight that decides

B = se²/(se² + τ²) is not a compromise between two answers. It is exactly the posterior mean's weight, it agrees with a numerical integration to ten digits, and an argument that mentions no population at all arrives at almost the same estimator.

hierarchical · Shrinkage
What a 95% credible interval covers, n = 20. Computed by summing over all 21 possible counts rather than by simulating them. Jeffreys' prior covers close to 95% across the range; a confident prior centred in the wrong place covers almost nothing where the truth is far from it.

What a credible interval covers

A credible interval makes the statement everyone wants and does not claim to have a coverage. It has one anyway, it can be summed over the sample space exactly, and on a reasonable prior it beats the interval taught first.

bayes · Credible
How often "there is no spread between the groups" is reported about data that has one. Every dataset here was generated with a real population spread of 1. The moment estimator is the difference between the observed spread and what noise alone would produce, clamped at zero, and the difference comes out negative often: at eight groups it reports exactly zero on 32.6% of datasets, which is an instruction to pool completely and give all eight groups the same estimate. The rate falls to 4.2% at 48 groups.

When the spread estimates to zero

The usual estimate of a population spread is a difference of two positive quantities, clamped at zero. On a third of eight-group datasets with a real spread in them the difference comes out negative, the estimate is exactly zero, and every group is pooled completely on data that said no such thing.

fullbayes · Pooling
A normal mean with a flat prior — one interval, two readings. Both intervals are [1.878, 3.446]. The frequentist reading is that the procedure captures the truth 95% of the time; the Bayesian reading is that the parameter is in this interval with probability 0.95. The endpoints are identical to machine precision.

Where the two schools agree

With a flat prior on a normal mean, the credible interval and the confidence interval are the same interval, endpoint for endpoint. Knowing exactly when that stops being true is more useful than either camp's general argument.

bayes · Credible
What a positive test means, sensitivity 90%, specificity 95%. At a prevalence of one in a thousand, 98 of every hundred positives are false. At one in 10, 33 are. The test has not changed.

The base rate was always Bayes

The screening arithmetic everybody finds counter-intuitive is a posterior update with a prior of one in a thousand. Naming it that way turns a famous puzzle into an instance of a rule, and makes the sequential version obvious.

bayes · Baserate
Counted coverage of two 95% intervals, over 2,000 datasets. Each point is one of the eight groups, at its own standard error. The integrated interval covers 95.2% overall against its stated 95%; the plug-in covers 78.8%, and its shortfall grows with the group's standard error — from 86.1% at se 0.5 to 77.0% at se 2.1. The mean widths are 3.48 and 2.66.

The interval that integrates

A credible interval for one group in a hierarchy has to average over every value the population spread might take. That averaging is what makes it cover — 95.2% against the plug-in's 78.8% — and it costs 31% more width, a heavier tail, and a mixture rather than a normal.

fullbayes · Credible
The p-value of a study with 80% power, twenty thousand times. Twenty thousand two-sided z-tests, each on 25 observations whose true mean is 0.5603 standard deviations from the null, a noncentrality of 2.802. The bars are the counted share of p-values in bins a quarter of a power of ten wide, with the leftmost bin holding everything smaller; the line is the closed form. The middle eighty per cent of the p-values runs from 4.4×10⁻⁵ to 0.13, 3.46 orders of magnitude, the median is 0.0051, and 80.0% fall below 0.05, which is what the power means.

The p-value a replication gets

Under a true null a p-value is flat. Under a real effect its distribution is closed form and wide — a study with 80% power returns anything from 4.4×10⁻⁵ to 0.13 in eight runs of ten — and the chance that an exact replication of a p = 0.05 result is significant again is exactly one half, under both of the models people use without naming them.

testing · Uniformity
Two 95% intervals for 2 of 20, under Jeffreys — Beta(½, ½). The equal-tailed interval runs from 0.0214 to 0.2839 and is 0.2625 wide; the shortest interval runs from 0.0093 to 0.2540 and is 0.2447 wide. Both hold 95% of the posterior, and the shorter one buys its 6.8% by moving its lower endpoint towards the denser side.

The shortest interval, and the one that does not move

Two 95% intervals come out of every posterior and they are not the same set. The shorter one is shorter by 4.86% on average and 22.41% at its best, it covers 86.72% where the other covers 95.68%, and it is not even the shortest once the parameter is written a different way.

bayes · Credible
Four 95% intervals for the odds after 2 of 20. credible, transformed: 0.0218 to 0.3964. Wald, transformed: -0.0305 to 0.3012. delta method on the odds: -0.0512 to 0.2734. delta method on the log-odds: 0.0258 to 0.4789. The first two are the same intervals for the proportion with their endpoints put through the odds; the last two are fresh approximations made on the new scale.

An interval for something else

An interval for the odds is free — put the endpoints through the odds and the coverage does not move, exactly, for any interval at all. The method everyone uses instead computes a new standard error on the new scale, and at twenty trials that costs four points of coverage, produces negative odds, and has no value at all when nothing was observed.

bayes · Credible
A prior worth 35 observations, moved across the range — truth 0.1, n = 20. The same prior weight centred at each of 33 places. Its interval covers 100.0% where the centre is near the truth and 0.0% at its worst, while the mean width where it covers least is 0.221 against a flat prior's 0.263 on the same data.

When the prior is confident and wrong

A prior worth thirty-five observations, centred in the wrong place, produces a 95% interval that covers nothing at all — and reports a width 5% narrower than an honest one. It takes seventeen thousand observations to repair, not thirty-five, and the worst study to run is the one whose sample size equals the prior's weight, exactly.

bayes · Credible

Named alongside it

The objects these essays reach for when they reach for this one.

PriorCoverageFlat priorCredible intervalShrinkageHierarchical modelPosterior meanEmpirical BayesPartial poolingDiscretenessJeffreys' priorLog-odds

All concepts