What a prior is worth
The objection to Bayesian methods is that the prior is arbitrary and lets an analyst put a thumb on the scale. The objection is answerable, and the answer is not a philosophical argument. A prior has a size, the size is a number of observations, and the number can be stated.
The update, in closed form
For a proportion the arithmetic is small enough to write down completely, which is why this field starts here rather than with a general statement.
A prior for a proportion is a Beta distribution with two parameters, a and b. Observing k successes in n trials turns it into another Beta distribution:
That is the entire update. The successes add to the first parameter, the failures add to the second, and no integration is required. A prior with this property is called conjugate, and its existence is what makes the whole field computable exactly rather than by simulation.
The posterior mean follows immediately:
The sentence that answers the objection
Look at what that expression is. It is the number of successes plus a, over the number of trials plus a + b. The prior enters exactly as though a successes and b failures had already been observed.
So a prior is worth a + b observations. Not metaphorically — the arithmetic is identical to having seen that many.
That converts an argument into a measurement. A Beta(1, 1) prior is worth two observations and will be swamped by any real dataset. A Beta(20, 20) prior is worth forty, so a study of twenty subjects is being outvoted two to one by an assumption. Anyone can compute which situation they are in, and the computation is an addition.
The posterior mean is a weighted average
The same expression, rearranged, says something sharper. The posterior mean is the data proportion and the prior mean, mixed in proportion to their weights:
Two consequences, and both are checkable rather than assertable.
The posterior always lies between the prior mean and the data. It cannot be more extreme than either. This is asserted by the figure above: the posterior mean is on the data’s side of the prior mean and never beyond the data. A method that produced an estimate outside that range would have a bug, and the assertion is there to catch one.
The prior’s share of the answer is a + b over a + b + n, which is a number anyone can compute before collecting anything. With a prior worth forty and a sample of ten, the prior supplies 80% of the answer. With the same prior and a sample of a thousand, it supplies 3.8%.
The three non-informative priors are worth zero, one and two observations
The measurement the section above installs prices the standard defaults immediately, and the range it gives them is narrower than the debate about them.
Haldane’s Beta(0, 0) is worth nothing — and is improper, which is the price of that. Jeffreys’ Beta(½, ½) is worth one observation. The flat Beta(1, 1) is worth two.
So the whole disagreement between the two usable non-informative priors is a single pseudo-observation. At twenty trials that is one part in twenty-one, and the prior’s weight in the posterior mean is for the flat prior and for Jeffreys’.
Against that, a Beta(20, 20) is worth forty observations and carries of the posterior mean on a study of twenty. Two thirds of the answer is the assumption.
Where one observation is not a small thing
The weight in the posterior mean is the wrong denominator when the successes are rare, and this is where the prior choice stops being a formality.
Take twenty trials at a true proportion of 0.05. The most likely count is zero, and the second most likely is one. The three priors then give posterior means of
- Haldane:
- Jeffreys:
- flat:
Three priors all described as non-informative, giving point estimates of 0, 0.024 and 0.045 for the same data — a spread as large as the proportion being estimated.
Nothing has gone wrong. The prior is worth one or two observations either way, and one or two observations is a large share of what a sample containing no successes actually knows. The denominator that matters is not the twenty trials; it is the amount of information about the quantity in question, and near a boundary that is much smaller than n.
The rate the weight decays at
One more consequence of , with m the prior’s worth in observations, is worth stating because it is slower than people expect.
The weight falls like 1/n, not exponentially. A prior worth m observations still carries a tenth of the posterior mean at and a hundredth at . Once n is comfortably past m, doubling the data halves the prior’s weight and no more.
So a Beta(20, 20) is down to a third of the answer at sixty observations, a tenth at three hundred and sixty, and a hundredth at four thousand. That is a long tail for an assumption to have, and it is the best argument for stating a+b in the paper: a reader who disagrees with the prior can compute how many observations they would need before the disagreement stopped mattering.
How fast the prior stops mattering
The pull toward the prior falls as 1/n, and measuring it makes the rate concrete. For a prior worth forty observations against a truth of 0.2:
- at n = 10, the posterior mean sits 0.240 away from the data
- at n = 100, 0.086 away
- at n = 1,000, 0.012 away
Ten times the data, about a tenth the pull. That is the rate the closed form predicts and the site’s gate requires the measurement to match it, because a demonstration that the prior washes out is worth much less than a demonstration that it washes out at a stated rate.
The practical reading: with a few hundred observations, any prior anyone would defend in public has stopped mattering. The disagreements that persist at that sample size are disagreements about the model, not about the prior — and they get attributed to the prior because the prior is the part with a name.
Where the weight is not the whole story
The a + b arithmetic is exact for a Beta prior on a proportion, and it is the right intuition in general. Two situations break it, and both matter.
A prior that excludes the truth is not overwhelmed by anything. Weight measures how much evidence is needed to move a prior, and it assumes the prior puts some mass where the truth is. A prior that assigns zero probability to a region can never produce a posterior there, however much data arrives — no amount of evidence multiplies zero into something. This is the failure that matters, and it does not show up in the weight at all.
For a Beta prior it cannot happen: every Beta with positive parameters is positive on the whole interval. It happens constantly in more elaborate models, where a prior is a bounded range and the truth is outside it.
And the weight is per-parameter. A model with forty parameters and a modest prior on each is carrying a great deal of prior information in total, and the per-parameter weight will look reassuring. The quantity that matters is how much the prior contributes to the thing being estimated, which in a hierarchical model can be much more than the sum of its parts suggests.
Both of those are arguments for stating the prior rather than against having one, which is the position this field takes throughout.
The prior nobody admits to
The comparison that makes the objection collapse is with what happens when no prior is stated.
An analyst choosing a model has made a hundred decisions with the same character as a prior: which variables to include, what functional form to assume, which observations are outliers, whether the effect is additive or multiplicative. Each constrains the answer, none is derived from the data, and none carries a weight anybody can compute.
The prior is the one such decision with an explicit representation and an arithmetic for how much it counts. That makes it the most auditable assumption in a typical analysis, not the least.
That is not a defence of any particular prior. A confident prior centred in the wrong place produces intervals that essentially never cover, and the size of that failure is measured in the next essay. The point is narrower: the prior’s contribution is computable, which is a property the rest of the modelling decisions do not have.
What to state, and what a reader can check
The field’s practical output is short.
State a and b, not a description. “A weakly informative prior” is unfalsifiable. Beta(2, 2) is worth four observations and anyone can see whether that is reasonable against the sample size.
State the prior’s weight beside the sample size, because that ratio is the whole question and it is a division.
Show the posterior under at least one other prior. If the conclusion survives a flat prior and a sceptical one, the prior is not carrying the result. If it does not survive, that is the finding, and it is a more interesting one than the original analysis.
The last is the one that changes practice, and it costs almost nothing when the update is conjugate: recomputing under a different prior is changing two numbers and adding.
That last figure carries a warning worth making explicit. A prior that happens to be centred near the truth converges immediately and looks excellent. The same prior against a different truth is the disaster in the coverage figure above. Nothing about how well a prior performs on one dataset says whether it was a good prior, because a prior is a statement made before the data and its quality is a property of the whole range of truths it might have faced.
Which is the same argument this site makes about intervals, tests and estimators, arriving from a new direction: a procedure is judged by what it does across the cases that could have occurred, not by what it did in the case that did.
Why conjugacy is not a cheat
One methodological note, since everything above rests on the closed form and closed forms invite suspicion that the easy case was chosen.
It was chosen, and it was chosen because the whole field can then be computed exactly rather than sampled. Every coverage number in this field is a finite sum over the sample space, the same sum the frequentist intervals on this site are measured with, so the two schools are compared on one axis with no simulation noise in either.
But conjugacy is a computational convenience and not a modelling assumption, so it needs checking that it has not quietly changed the answer. The site’s gate does that directly: the posterior mean from the closed form is compared with one computed by multiplying prior by likelihood on a fine grid and normalising — brute-force numerical integration, sharing no arithmetic with the conjugate shortcut beyond the Beta density itself.
They agree to ten digits. So the conjugate update is the posterior, rather than something convenient that resembles it, and the exactness the rest of the field depends on is earned.
The priors this field uses, and what each is for
Four named priors run through these essays. Naming them here saves repeating the description, and setting them beside each other shows that the choice is a small, discrete decision rather than an open field.
Beta(1, 1), the flat prior. Every proportion equally likely beforehand. Worth two observations. It is the obvious candidate for “no information” and it is not neutral: flat on the proportion means sharply informative about the log-odds, which is the reparameterisation problem in its simplest form.
Beta(½, ½), Jeffreys’ prior. Constructed to be invariant under reparameterisation, so the answer does not depend on which coordinate was written down. Worth one observation — lighter than flat, which surprises people. It piles a little extra mass near zero and one, which is exactly the correction needed where the flat prior’s intervals drift.
Beta(2, 2), weakly informative. A gentle statement that extreme proportions are less likely than middling ones. Worth four observations. Used when there is genuine background knowledge that the answer is not 0 or 1 and no more than that.
Beta(20, 20), confident. Worth forty observations, centred at a half. Included in this field not as a recommendation but as an object of study: it is the prior that shows what a heavy prior does, and it is heavier than most datasets it will meet.
The instructive comparison is Jeffreys against flat. The lighter prior is the one with the better frequentist behaviour, which cuts against the intuition that “less informative” means “flatter”. A prior’s weight and a prior’s neutrality are different properties, and the flat prior has more of the first and less of the second than its name suggests.
What a prior cannot do
The weight arithmetic can encourage a picture of the prior as a dial running from “no influence” to “total influence”, and two things fall outside that picture.
A prior cannot be overwhelmed where it assigns no probability. Weight measures resistance to evidence, and it assumes there is somewhere for the posterior to move to. A prior that is zero on a region produces a posterior that is zero there for any amount of data, because the update multiplies and zero times anything is zero. For a Beta prior with positive parameters this cannot arise, which is one reason the field is built on it. In models where a parameter is given a bounded range, it arises whenever the truth is outside the range, and it is undetectable from inside the analysis.
And a prior is not the only thing standing between the data and the answer. The likelihood encodes a model — that the trials are independent, that the probability is constant, that the sampling was what it is claimed to be. Those assumptions are not weighted in observations and cannot be washed out by more data; a thousand observations from a mis-specified model give a confident wrong posterior rather than an uncertain one.
The order of worry is therefore the same as it is for coverage under a misspecified model: is the model right, then is the prior reasonable, then how heavy is it. The last question is the one with tidy arithmetic, and it is the least important of the three.
Why a stated prior beats an unstated one
The argument that closes this essay is comparative rather than absolute, and it does not require believing that any particular prior is correct.
Every analysis embeds assumptions that constrain the answer without deriving from the data. Which covariates to adjust for. Which observations to exclude. Whether the effect is additive on the outcome or on its logarithm. Whether a borderline observation is a data-entry error. Each of those moves the result, each is a judgement, and none has a weight that can be computed.
The prior is the one assumption in that list with an explicit representation, an arithmetic for how much it contributes, and a straightforward sensitivity check — recompute under a different one. It is the best-behaved assumption in a typical analysis, not the worst.
So the objection that Bayesian methods introduce subjectivity has the sign wrong. The subjectivity is already there in any analysis that involves choices. What the prior does is make one piece of it explicit and quantifiable, and the discomfort it causes is discomfort at seeing something that was previously distributed invisibly across a dozen unlabelled decisions.
That is the case for the field, and it is why every essay here reports the prior’s weight beside the sample size. The number is a division, it takes a second, and it answers the only question anybody actually has about a prior: how much of this answer is the data?
One number to take away
If a single quantity survives from this essay it should be the ratio a + b to n: the prior’s weight against the sample size. It is a division, it takes a moment, and it answers the question every argument about priors is a proxy for.
Below about a tenth, the prior is decoration and the dispute is empty. Above about one, the prior is the analysis and the data is a correction to it. Between those, both matter and both should be reported — which is nearly always the interesting case, and nearly never the one that gets described.
That last case is where the sensitivity analysis stops being a formality and becomes the result: two defensible priors giving materially different answers is a finding about how much the data actually determines, and it is worth more than either answer reported alone.
Reporting both is not hedging. It is the honest description of what a small sample and a real prior jointly support.
What links here
Computed from the collection, not written here: the essays that point at this one.
Reads more easily once this is understood
Essays that name this one as worth reading first.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- What the plug-in forgets — both name coverage, flat prior, posterior, posterior mean, prior
- The interval that integrates — both name coverage, flat prior, posterior, prior
- Ninety-three observations, and nothing assumed — both name beta distribution, coverage, sample size
- A block size that changes — both name coverage, sample size
- A coverage table with its own error — both name coverage, sample size
- A league table of a hundred — both name posterior mean, sample size
Named objects
A flat tag is an object no other essay names yet.
Beta distributionConjugacyCoverageFlat priorPosteriorPosterior meanPriorPrior weightSample size