The prior, doing visible work

An interval for something else

An interval for the odds is free — put the endpoints through the odds and the coverage does not move, exactly, for any interval at all. The method everyone uses instead computes a new standard error on the new scale, and at twenty trials that costs four points of coverage, produces negative odds, and has no value at all when nothing was observed.

Worth reading first: What a credible interval covers.

Two successes in twenty is a proportion of a tenth, and a 95% interval for it is routine. The question a reader usually has is about the odds — the proportion divided by one minus it, the quantity a logistic regression reports and an epidemiologist compares — and the answer takes one line.

Put the endpoints of the interval for the proportion through the odds. That is the interval for the odds.

Four 95% intervals for the odds after 2 of 20credible, transformed: 0.0218 to 0.3964. Wald, transformed: -0.0305 to 0.3012. delta method on the odds: -0.0512 to 0.2734. delta method on the log-odds: 0.0258 to 0.4789. The first two are the same intervals for the proportion with their endpoints put through the odds; the last two are fresh approximations made on the new scale.00.2000.400the oddsone row per intervalzero — the odds cannot go below itcredible, transformedWald, transformeddelta method on the oddsdelta method on the log-odds2 of 20, Jeffreys — Beta(½, ½)two of these are the same interval in other units
Fig. 1 Four 95% intervals for the odds after two of twenty. The top two are intervals for the proportion with their endpoints transformed; the bottom two are fresh approximations computed on the new scale. One of the bottom two starts at −0.0512, which is a value the odds cannot take.

Why the one line is exact

The argument is about monotone functions and has no statistics in it.

A 95% interval [L,U][L, U] for pp is a statement of the form the truth lies between these two numbers. If ff is increasing, then pp lies between LL and UU exactly when f(p)f(p) lies between f(L)f(L) and f(U)f(U) — the two statements are the same statement, written with different symbols.

So whatever probability attached to the first attaches to the second. For a credible interval that probability is the posterior probability, and it stays 95%. For a confidence interval it is the coverage, and it stays whatever it was. The transformation preserves what was true and creates nothing.

This is worth stating carefully because the conclusion cuts both ways, and the half that gets forgotten is that it is not a Bayesian result. The Wald interval for a proportion covers 87.60% at twenty trials and a truth of a tenth; its endpoints put through the odds cover the true odds 87.60% of the time — the same number to every digit the arithmetic carries, because the interval covers on exactly the same counts.

Which counts, rather than how often

That last clause is the mechanism and it is checkable directly.

Which samples Wilson and Wald each cover, n = 20, p = 0.1. Each bar is the probability of one count, shaded by which interval built on that count contains 0.1. Both cover 83.5% of samples, only Wilson 12.2%, only Wald 4.1%, neither 0.2%. The correlation between their hits is -0.044, so on shared draws the variance of their difference is 1.040 times larger than on independent ones.
Fig. 2 Coverage as a set of outcomes rather than a number: each bar is one possible count, shaded by whether the interval built on that count contains the truth. Coverage is the total height of the shaded bars, which is why two procedures with the same shading have the same coverage however the axis is relabelled.

An interval procedure assigns, to each of the twenty-one possible counts, a verdict: covered or not. Coverage is the binomial probability of the covered set. A change of variable relabels the axis and leaves that set alone, so it leaves the sum alone.

Which means the question what is the coverage of the transformed interval has an answer before anything is computed, and the answer is the same. There is no measurement to make.

What gets computed instead

The interval nobody actually builds that way is the one this essay is about.

Standard practice for a function of an estimate is the delta method: approximate the function as locally linear, multiply the standard error by the derivative, and put a normal interval around the transformed estimate. For the odds the derivative is 1/(1p)21/(1-p)^2; for the log-odds it is 1/(p(1p))1/(p(1-p)), and the interval is built there and exponentiated.

Both are perfectly respectable approximations. Neither is the transformed interval, and the difference is not a matter of rounding.

At two of twenty the transformed Wald interval runs from −0.0305 to 0.3012 — negative already, because the Wald interval for the proportion started below zero — and the delta-method interval on the odds runs from −0.0512 to 0.2734. The log-odds version runs from 0.0258 to 0.4789, which is positive by construction and 1.40 times as long as the delta-method interval for the same quantity from the same data.

Two approximations of one interval, differing by three quarters of their own length. Nothing in the data chose between them; the choice was a scale, made before the arithmetic started.

Counted

Since these are new intervals rather than old ones rewritten, their coverage is a new question, and this site answers those by summing over the sample space.

What a 95% interval for the odds covers, n = 20. Summed over all 21 counts at each of 97 true proportions. Transforming an interval's endpoints leaves its coverage exactly where it was; computing a fresh standard error on the new scale does not, and the log-odds version is the worst of the four at 27.2%.
Fig. 3 The exact coverage of each of the four intervals for the odds, at ninety-seven true proportions. The two transformed intervals are the proportion-scale curves unchanged. The two delta-method intervals are new curves, and the log-odds one is the lowest of the four almost everywhere on the left.

At a true proportion of a tenth the four cover 95.68%, 87.60%, 87.84% and 83.52% — credible transformed, Wald transformed, delta on the odds, delta on the log-odds. The last is four points below the interval it was meant to improve on, which is the wrong direction for a device chosen for its better behaviour near a boundary.

The worst cases, over true proportions from 0.02 to 0.6, separate them further: 89.34% for the transformed credible interval, 33.18% for the transformed Wald one, 33.24% for the delta method on the odds, and 27.25% for the delta method on the log-odds. All three of the last are catastrophic at the same place — a true proportion of 0.02, where twenty trials usually produce a count of zero — and the transformed credible interval is not.

Zero, where the whole apparatus stops

The count of zero is where the difference stops being a matter of percentage points.

Four 95% intervals for the odds after 0 of 20. credible, transformed: 0.0000 to 0.1320. Wald, transformed: 0.0000 to 0.0000. delta method on the odds: 0.0000 to 0.0000. delta method on the log-odds: undefined at this count. The first two are the same intervals for the proportion with their endpoints put through the odds; the last two are fresh approximations made on the new scale.
Fig. 4 Nothing observed in twenty trials. The estimated proportion is zero, so the estimated standard error is zero, so the delta-method interval on the odds is the single point 0 and the transformed Wald interval is the single point 0. The delta-method interval on the log-odds does not exist at all: its standard error is a division by zero.

None of that is a failure of the delta method as an approximation. It is the approximation being handed an estimate at the boundary of its parameter space, where the estimated standard error is a function of the estimate and the estimate is zero. Every interval built from p^(1p^)\hat p(1-\hat p) has this problem and the usual repairs — add a half to each cell, add two successes and two failures — are priors that nobody calls priors.

The credible interval has no such case. After nothing in twenty trials, Jeffreys’ prior gives a posterior with mass spread over the small proportions, and the interval for the odds runs from 0.0000 to 0.1320. The lower endpoint is very small and it is not zero, which is the correct statement: a study that saw nothing has not established that the odds are zero.

At a true proportion of 0.02 the count is zero on 66.8% of studies, so this is not the tail of the problem. It is most of it.

Where the delta method is fine, which is most of the axis

The two figures so far were drawn at counts of two and zero, and an argument built on them alone would be an argument about a corner.

Four 95% intervals for the odds after 7 of 20. credible, transformed: 0.2081 to 1.3136. Wald, transformed: 0.1641 to 1.2678. delta method on the odds: 0.0437 to 1.0332. delta method on the log-odds: 0.2148 to 1.3496. The first two are the same intervals for the proportion with their endpoints put through the odds; the last two are fresh approximations made on the new scale.
Fig. 5 Seven of twenty, where nothing is wrong. All four intervals exist, all four are entirely positive, and they agree closely enough that the choice between them would not change a sentence anybody wrote. The delta-method interval on the odds runs from 0.0437 to 1.0332.

At an estimated proportion of 0.35 the normal approximation is good, the linearisation is good, and the four intervals are four slightly different answers to the same question. The delta method is not a bad method; it is a method whose failures are concentrated where the estimate is near a boundary, and at seven of twenty it is nowhere near one.

That matters for how the earlier numbers should be read. The coverage figures are averages over the whole sample space at a given truth, so they mix counts like this one — where every method agrees — with counts like zero, where two of the four are degenerate. At a true proportion of a tenth the count is zero or one on 39.2% of studies, which is why a coverage that looks like a modest shortfall is really a good answer most of the time and no answer the rest of it.

It also says where a comparison drawn at a single count would have gone wrong. “The delta-method interval runs below zero” is true at two of twenty and false at seven, so it is a fact about a setting and not about the family, which is the commonest way a claim like this goes wrong. It is the reason the claim that the delta interval runs below zero on 86.72% of the sample space is counted across every count rather than read off one picture.

At four times the data, the ordering changes and the worst case does not

The natural expectation is that all of this is a small-sample complaint, and half of it is.

What a 95% interval for the odds covers, n = 80. Summed over all 81 counts at each of 97 true proportions. Transforming an interval's endpoints leaves its coverage exactly where it was; computing a fresh standard error on the new scale does not, and the log-odds version is the worst of the four at 77.9%.
Fig. 6 The same four intervals at eighty trials. At a true proportion of a tenth the delta method on the log-odds now covers 96.26% against the transformed credible interval’s 93.79% — it has gone from four points below to two and a half above — while the worst cases over the range have barely moved.

At a tenth the ordering has reversed. The log-odds approximation, which was the worst of the four at twenty trials, is the best of them at eighty, and the transformed credible interval has dropped below its nominal level. That is the normal approximation coming good: at eighty trials and a proportion of a tenth there are eight successes on average, the log-odds estimate is close to normal, and the interval built on it is close to right.

The worst cases tell the other half. Over the same range of true proportions they are 91.19% for the transformed credible interval, 80.02% for the delta method on the odds and 77.90% for the log-odds version — against 89.34%, 33.24% and 27.25% at twenty trials. The credible interval’s worst case has moved by less than two points; the other two have improved enormously and are still ten to thirteen points short.

The worst case is at a smaller true proportion now, and there is always a smaller one. Quadrupling the data moves the place where the approximation fails; it does not remove it, because the failure is driven by the expected count rather than by the sample size, and at any sample size there is a proportion small enough that the expected count is near zero. That is the same structure the interval field’s first essay found in the Wald interval and the reason a worst case is worth reporting beside an average.

The delta method is not solving this problem

It would be easy to read the last three sections as an argument that the delta method is a bad device, and that is not what they say. It is a good device aimed somewhere else.

The delta method exists because a function of an estimate usually has no distribution anybody can write down. Where the function is monotone and of one parameter, that difficulty does not arise: the interval transforms, exactly, and there is nothing left to approximate. The delta method earns its place where endpoint transformation gives nothing — a function of several parameters at once, a function that is not monotone, a ratio whose denominator changes sign.

So the honest summary is narrower than “use transformation instead”. It is that for a monotone function of one parameter, an approximation is being made where an identity was available, and the approximation’s cost is a new coverage that has to be counted rather than inherited.

That the cost is four points at one setting and sixty at another is the part worth carrying. A device that is a good idea in general can be a poor one in a case it was never needed for, and the case where it is not needed is exactly the one most often met.

What the four intervals are actually made of

Stepping back from the numbers, the four objects in the hero figure are made of three different kinds of ingredient, and separating them says why two of them behave and two do not.

The transformed credible interval is a posterior quantile, put through a function. It needs the Beta–binomial conjugacy, which is exact, and monotonicity, which is a fact about the odds. No approximation enters at any point, which is why its coverage is the proportion-scale coverage and why it exists at every count.

The transformed Wald interval is a normal approximation to the sampling distribution of p^\hat p, put through the same function. One approximation, made on the proportion scale, and the function applied after.

The two delta-method intervals are a normal approximation made on the transformed scale, using a derivative evaluated at the estimate. Two approximations rather than one: the normal law, and the linearisation. At an estimate of zero the second is exact and useless — the derivative is finite but the standard error it multiplies is zero.

Written that way the results stop being surprising. One approximation is better than two; zero is better than one; and the place all of them fail is the place where the estimate sits on a boundary of its own range, which is where an estimate at the edge of its range has already been shown to break everything computed from it in a completely different setting.

What does not transform

Everything so far has been about intervals. The estimate at the middle of one behaves differently, and the difference is sharper than the interval’s.

Three summaries of the same posterior odds, n = 20. The posterior median of the odds is the odds of the posterior median, exactly, because a median is a quantile. The posterior mean of the odds is not the odds of the posterior mean — it is larger at every count — and on 1 of the 21 counts it is infinite, because the integral diverges.
Fig. 7 Three summaries of the same posterior for the odds. The median of the odds is the odds of the median, exactly, because a median is a quantile and quantiles transform. The mean of the odds is not the odds of the mean — it is larger at every count — and past a count it is not a number at all.

The posterior median transforms for the reason the equal-tailed interval does: it is a quantile, and the argument of the second section applies unchanged.

The posterior mean does not. The odds is a convex function, so the average of the odds exceeds the odds of the average, at every count. After four of twenty the posterior mean of the odds is 0.2903 against an odds of the posterior mean of 0.2727 — six and a half per cent apart, which is small and is not zero.

And after twenty of twenty it is infinite. The posterior is Beta(20.5, 0.5), the integral of p/(1p)p/(1-p) against it diverges, and the quantity a reader would have asked for does not exist. The median of the odds at that count is 88.5 — large, because twenty successes in twenty is strong evidence that the proportion is near one, but finite and computable. Its 95% interval runs from 7.57 to 41242, which is enormous and is the honest answer: it is what the data leaves open. The mean is not large. It does not exist.

An interval survives a change of variable and a mean does not, which reverses the usual ordering of how solid the two feel. A point estimate looks like the simpler object and it is the one that breaks.

The maximum-likelihood estimate is the exception on the frequentist side and is worth naming so the symmetry is not overstated: it is invariant, odds^=p^/(1p^)\widehat{\text{odds}} = \hat p/(1-\hat p) by definition, which is the whole content of the invariance property. So the two schools each have one summary that transforms — the mode and the median — and one that does not, and it is the posterior mean, the summary Bayesian output reports by default.

What is claimed here, and what is not

The claim is what a change of variable does to four intervals and three point estimates for a proportion: that transforming endpoints preserves both the posterior probability and the frequentist coverage exactly, for any interval and not only a Bayesian one; that a delta-method interval computed on the new scale is a different interval whose coverage has to be counted again and is between four and sixty points worse depending on where the truth is; and that the posterior mean is the one summary here that fails to transform and sometimes fails to exist.

Everything is an exact sum over the twenty-one possible counts. The only approximations anywhere in the essay are the ones being measured.

What stays out: functions of more than one parameter, which is where the delta method is not replaceable by anything as simple and where a ratio whose interval has to be the whole line takes the argument up; non-monotone functions, where an interval for pp maps to a set that need not be an interval; profile-likelihood intervals for a transformed quantity, which are invariant by construction and are a third route nobody here has counted; and the small-sample corrections — Haldane’s, Agresti’s — that repair the boundary case by adding pseudo-counts, which is a prior arriving under a different name and deserves its own measurement rather than a mention.

One thing deliberately not claimed: that the transformed credible interval is good. It is exact about its posterior probability and it inherits whatever coverage it had, and what it had was the coverage the essay that first counted it found — good under Jeffreys’ prior and arbitrarily bad under a bad one. Transformation is a conservation law, not an improvement, and an interval that was wrong on the proportion scale is exactly as wrong on the odds scale. The comparisons here are between procedures given the same prior; they say nothing about the prior.

Still open: a prior in the wrong place

Every number above uses Jeffreys’ prior, and the transformed credible interval wins every comparison here on the strength of it. That is not a property of Bayesian inference; it is a property of a prior that happens to be reasonable, and the coverage a credible interval actually has has already been measured under a confident prior in the wrong place, at one setting.

What has never been measured here is the shape of that failure: how coverage depends on the distance between the prior and the truth, how much data repairs it, and — the question a reader of one analysis actually faces — whether anything on the page says it has happened. The answer to the last is no, and the sample size that repairs a prior worth thirty-five observations is not thirty-five or a few hundred but seventeen thousand. That is what a confident prior in the wrong place costs.

The check, and the refusal

Two claims are gated, at different strengths on purpose.

The identity is required to hold to 101410^{-14}: the transformed Wald interval’s coverage must equal the untransformed Wald interval’s, and the transformed credible interval’s must equal the proportion’s, at four true proportions each. Those are the same finite sum reached by two routes, so anything weaker than machine precision would be hiding an arithmetic difference rather than tolerating one.

The comparison is stated as an inequality, because it is one: the delta-method interval must cover at least five points less than the transformed credible interval at a tenth, and the log-odds version must differ from the transformed interval by more than two points. An inequality is what a claim about two different procedures can support; an equality would be a claim about a coincidence.

The refusal is the essay’s own distinction, required to fail: the delta-method interval must not be the transformed interval. It is required to run below zero on more than half the sample space, to be a single point where nothing was observed, and to disagree with its own log-scale version by more than two points of coverage. If it passed, the distinction the whole essay is built on would be a distinction between two names for one object, and every comparison above would be measuring rounding error — which is the failure a check that cannot reject always is.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

CoverageCredible intervalDelta methodJeffreys' priorLog-oddsMonotone transformationOddsPoint estimatePosteriorPosterior meanWald interval