Hierarchy past one number

What a two-unit study should report

The fixed-effect interval covers the mean of the two sites in hand 96.37% of the time and the population mean 54.77%. The random-effects interval covers the population mean 94.96% — exactly its level, from one degree of freedom — and is 11.6 times wider.

Worth reading first: Eight groups, one population · The prior the data estimates.

The essay before this one established that a variance estimated from two units is uninformative: a scaled chi-square on one degree of freedom, with an interquartile range spanning a factor of thirteen and a ten-to-ninety range spanning a factor of a hundred and seventy-one, coming out exactly zero more than a quarter of the time.

That is a statement about a nuisance parameter. What a two-site study actually prints is an interval for its overall mean, and three things can be done about the variance on the way to one: treat the sites as the population and ignore it, estimate it from the single degree of freedom available, or take it from outside the study.

They give different intervals, and — the part that is easy to miss — two of them are intervals for different quantities.

Each interval covers one question and not the other. Coverage of each interval for the overall mean, scored against both estimands, over 20,000 two-site studies of 10 observations apiece. The fixed-effect interval covers the mean of the two sites in hand 96.37% of the time and the population mean 54.77%. The random-effects interval covers the population mean 94.96% — exactly its level, from one degree of freedom — and over-covers the two sites in hand at 98.25%. Both are correct; they are answers to different questions printed in the same place.
Fig. 1 Coverage of each interval for the overall mean, scored against both estimands, over twenty thousand two-site studies of ten observations apiece. The fixed-effect interval covers the mean of the two sites in hand 96.37% and the population mean 54.77%.

Neither number is a mistake. They are answers to different questions printed in the same place.

The two questions

A two-site trial can be asked about two things and the words used to report it do not distinguish them.

These two sites. The estimand is the average of the two sites that were actually studied, and the only uncertainty is the within-site sampling error. There is no between-site variance in the question at all, because the sites are not a sample of anything.

Sites in general. The estimand is the mean of the population the two sites were drawn from, and the uncertainty has two parts: the within-site error, and the fact that two sites are a very small sample of sites.

The first is a statement about a study and the second is a statement about the world, and the second is almost always what a reader takes away — a trial run at two hospitals is read as evidence about the treatment, not about those two hospitals. The fixed-effect interval covers that reading 54.77% of the time.

The confusion is easy to make and hard to notice because the point estimate is the same. Both analyses report the average of the two site means, and they differ only in the width around it and in what the width is a statement about. A reader comparing two papers with the same estimate and very different intervals will reach for a difference in sample size or in analysis quality, and the difference is in the question.

The interval that is exact, and why it is

The random-effects interval at two units has an unexpectedly clean property: it is exact, not approximate.

With two units the between-unit standard error is half the gap between the two site means, on one degree of freedom, and the interval yˉ±t(1)yˉ1yˉ2/2\bar y \pm t(1)\,|\bar y_1 - \bar y_2| / 2 is an ordinary t interval with one degree of freedom. It covers the population mean 94.96% of the time, which is its level, and it does so at every combination of within- and between-site variance — the t argument does not care how the two are divided, because both enter the site means the same way.

So the previous essay’s finding about the variance being uninformative does not make the interval invalid. It makes it wide.

What one degree of freedom costs in width. The median width of each interval for the overall mean, over 20,000 two-site studies of 10 observations apiece. The fixed-effect interval is 1.127 wide and is an interval about these two sites. The exact random-effects interval is 13.02 — eleven times wider — because it multiplies half the gap between the two site means by a t on one degree of freedom, which is 12.706. Its tenth and ninetieth percentiles are 2.39 and 31.7, a factor of thirteen, which is a between-unit variance from one degree of freedom arriving as a width.
Fig. 2 The median width of each interval. The fixed-effect one is 1.127 and the exact random-effects one is 13.02 — eleven and a half times wider — because it multiplies half the gap by a t on one degree of freedom, which is 12.706.

The multiplier is 12.706, against 1.96 for a normal. Its tenth and ninetieth percentiles of width are 2.39 and 31.67, a factor of thirteen — the previous essay’s chi-square arriving as a width rather than as a variance.

That is an honest interval and it is nearly useless. It covers what it says it covers, and a study reporting one has told a reader that the data supports almost nothing about sites in general — which is true, and is the finding.

The one interval that over-covers, and why that is not a fourth option

The table has a fifth row worth reading: the exact random-effects interval, scored against the two sites in hand, covers 98.25%.

That is over-coverage and it is what a reader would expect — an interval built to account for between-site variation contains the average of the two studied sites more often than 95% of the time, because part of its width is doing a job the narrower question does not need done.

It matters because it rules out the obvious compromise. A reader tempted to report the exact interval and interpret it as a statement about these two sites gets an interval that is correct in the weak sense of covering at least its level, and is eleven times wider than the interval that question actually needs. Paying for a between-site variance and then not making the claim it buys is the worst available arrangement: the width of the general claim and the content of the specific one.

Where the cost actually sits

The two-unit case reads as pathological and the sweep says it is the end of a curve.

The whole cost sits at two units and three. The median width of each interval against the number of top-level units, at 10 observations in each. The interval that estimates the between-unit variance is 13.02 wide at 2 units, 4.42 at three and 0.983 at 20. The interval handed that variance is 2.965 and 0.938 at the same two ends. So what knowing the variance is worth runs from a factor of 4.4 to almost nothing, and the two-unit case is the end of a curve rather than a different problem.
Fig. 3 The median width of each interval against the number of top-level units. The interval that estimates the between-unit variance is 13.02 wide at two units, 4.42 at three and 0.98 at twenty. The one handed that variance is 2.97 and 0.94 at the same two ends.

Knowing the variance rather than estimating it is worth a factor of 4.4 at two units, 1.8 at three, 1.4 at four and 1.05 at twenty. By ten units the two intervals are within a fifth of each other and by twenty the question has stopped mattering.

So the whole of what this essay is about lives at two units and three. That is not a small corner: two-site trials, two-arm cluster studies, before-and-after designs with two periods and multi-country studies with two countries are all in it, and they are in it because of a count nothing in their output reports.

A second reading of the same curve is worth having, because it is the one a design meeting can act on. The exact interval’s width at K units is proportional to t(K1)/Kt(K-1)/\sqrt{K}, and that product falls from 8.98 at two units to 2.48 at three, 1.59 at four and 0.47 at twenty. Almost all of the fall is in the first step: going from two units to three cuts the interval to a third, where going from ten to twenty cuts it by a quarter.

The reason is that the t multiplier and the square root are both improving at once, and the t multiplier improves fastest where the degrees of freedom are fewest — from 12.706 to 4.303 in one step. A study contemplating a second site and a third is not making the same kind of decision twice.

The variance supplied from outside

The third option is to take the between-unit variance from somewhere else — an earlier study, a meta-analysis, a regulator’s assumption — and build a normal interval with it. It covers 95.14% when the value is right, at a width of 2.97, which is between the other two.

What it has to be right about is the question.

What a supplied variance has to be right about. Coverage of the interval built from a between-unit standard deviation taken from outside the study, against the factor that value is wrong by. At the truth it covers 95.14% in 20,000 studies. At a quarter of the truth it covers 59.74%; at half, 75.16%; at four times, 100.00% with an interval 3.8 times as wide. The error is one-sided in consequence — too small breaks the promise and too large only costs width — which is the same shape as every other conservative choice.
Fig. 4 Coverage of the interval built from a supplied between-site standard deviation, against the factor that value is wrong by. A quarter of the truth covers 59.74%; half covers 75.16%; four times covers 100.00% at 3.8 times the width.

The error is one-sided in consequence. Too small breaks the promise badly — 59.74% at a quarter of the truth — and too large costs only width, 99.98% at twice the truth and 100.00% at four times.

That shape has a practical reading: a supplied variance should be supplied generously, and the sensible thing to supply is an upper bound rather than a best estimate. A bound from a previous literature is usually available where a point value is contested, and the loss function says the bound is the right object.

It also means the third option is not a free lunch. It converts a wide but honest interval into a narrow one whose validity depends on a number the study did not measure, and there is no diagnostic inside a two-site study that could check it — one degree of freedom is one degree of freedom whatever the interval is built from.

The three options as one table

Collecting them gives the choice a two-site study actually faces, with the price of each.

A fixed effect for the unit costs nothing and answers the narrow question. Width 1.127, coverage 96.37% of the mean of these two sites, 54.77% of the population mean. It needs no input the study does not have.

The estimated between-unit variance answers the broad question and is exact. Width 13.02, coverage 94.96% of the population mean. It needs no input either, and it says the data supports almost nothing.

A supplied between-unit variance answers the broad question at width 2.97 and coverage 95.14%, and it needs a number from outside whose error is not symmetric in consequence: a quarter of the truth takes the coverage to 59.74%.

There is no fourth column with a short interval, the broad question, and no external input in it, and the reason is not that nobody has built one. Two units is one degree of freedom, and an interval for a population from a sample of two either uses a t on one degree of freedom or uses something the sample did not supply. The table is a complete account of the options because the arithmetic is.

What a defensible report looks like

Add units before adding observations. The curve says the return to a third site is enormous and the return to a twenty-first is nothing, so a study with a fixed budget and a choice between more observations per site and more sites has an answer at the small end that is not close. This is the same arithmetic a design effect makes visible: observations inside a unit repeat each other and units do not.

Say which estimand the interval is for. “The overall mean” is ambiguous between two quantities whose intervals differ by a factor of eleven, and both analyses are routinely described with the same phrase. One sentence — these two sites or sites of this kind — settles it.

Report the number of units, not just the number of observations. A study with two hundred observations at two sites and one with two hundred at twenty sites are the same size by every count an ordinary methods section makes, and their intervals for sites in general differ by a factor of thirteen. That is the same complaint a cluster-robust interval read against the wrong reference makes about rows and clusters, arriving one level up.

Report the gap between the two sites. It is the entire between-site evidence the study contains — with two units, half the gap is the between-unit standard error — so printing it lets a reader construct any of the three intervals above. It also makes the design’s limitation visible in a way a width does not, because a reader can see that one number is carrying the whole claim. That is the same discipline a second number beside a promise asks for everywhere here.

Prefer the wide honest interval to a narrow assumed one, unless the assumption is external and stated. The exact interval at two units is enormous and it is the only one of the three that is right without an input the study did not produce. Where a supplied variance is used it should be named, sourced, and chosen from the generous end.

Do not read a narrow interval from a two-unit study as reassuring. Of the three options, the narrowest by far is the one that answers the question a reader is least likely to be asking, and the second narrowest depends on an unverifiable input. A two-site study with a tight interval about sites in general has either assumed a variance or reported the wrong estimand, and which of the two is not visible in the interval.

Do not let the estimate’s stability stand in for the interval’s. The point estimate of the overall mean is the same under all three analyses and is perfectly well behaved; it is the width that carries the design’s limitation. A study reporting a stable estimate across analyses has demonstrated that its estimate is stable, which is the property that was never in doubt. It is the same confusion a coverage that says nothing is about: one good property is not a summary of a procedure.

And consider whether the question can be narrowed. A study that cannot say anything useful about sites in general may be able to say something useful about these two, and the fixed-effect interval is a perfectly good interval for that — 1.127 wide and covering 96.37% of its own estimand. Narrowing the claim to what the design supports is a better response than widening it to a factor of thirteen or assuming a number.

The count that is doing the work

One reading collects everything above, and it is the same reading two other fields arrive at from their own directions.

The quantity that governs an interval about a population is the number of units drawn from that population, and it is not the number of observations. Two hundred observations at two sites give an interval about sites in general on one degree of freedom; the same two hundred at twenty sites give one on nineteen. The multiplier falls from 12.706 to 2.093 and the square root from 1/21/\sqrt{2} to 1/201/\sqrt{20}, and between them they account for the factor of thirteen.

That is the same arithmetic a cluster-robust standard error runs into, where the consistency is in the number of clusters and a study with five has five of whatever the limit is in. It is the same arithmetic the design effect makes visible, where observations inside a unit repeat each other and units do not. And it is the reason a variance from two units is uninformative rather than merely imprecise: one degree of freedom is not a small amount of information about a variance, it is the smallest amount that is not zero.

A methods section that reports n and not K has withheld the number the reader needs, and no amount of care in the analysis puts it back.

What is claimed here and what is not

Twenty thousand studies at each setting. Every coverage figure is a count over that many simulated two-site studies, so a reading of 94.96% carries a standard error of about 0.15 points and the difference between 94.96% and 95% is not a finding. What is a finding is 54.77% against 96.37%, which is four hundred standard errors.

Ten observations per site, one variance ratio. The sweeps use a between-site standard deviation of 1 against a within-site one of 1.2 with ten observations each, so the site means carry about equal contributions from the two sources. A design where within-site error dominates makes the fixed-effect interval’s failure less severe and the exact interval’s width less extreme; one where between-site variance dominates makes both worse.

The exact interval’s exactness is a t argument, not a simulation finding. yˉ±t(K1)s/K\bar y \pm t(K-1)\,s/\sqrt{K} covers at its level whenever the unit means are normal, whatever the two variance components are, and the measurement at 94.96% is a check on the implementation rather than evidence for the property. What the simulation supplies that the theory does not is the width and its spread, which is what this essay is about.

The fixed-effect reading is scored against the realised site effects. Its estimand is the average of the two site effects that were actually drawn, so it is a moving target across replications — which is correct: that is what “these two sites” means, and a fixed-effect analysis is a statement about whichever units were studied.

Normal site effects throughout. The exact interval’s exactness rests on the two site means being normal, and with two of them there is no way to check. A heavier-tailed distribution of site effects would make the t interval’s coverage depart from its level, and the departure would be undetectable by any diagnostic the study could run — which is a limitation of the design rather than of the interval, and is the reason a two-unit study’s assumptions are load-bearing in a way a twenty-unit study’s are not.

And the supplied-variance sweep assumes the supplied value is the only thing wrong. A study taking a variance from elsewhere is also assuming that elsewhere’s units resemble its own, which is an exchangeability assumption with no test available at two units. The sweep prices getting a number wrong; it does not price the assumption that the number is about the same thing.

Still open: what an integrated variance buys

Three responses were named at the previous essay and this one measures two of them properly plus a third that is between the two. The fourth — putting a prior on the between-unit variance and integrating it out — is the one a Bayesian analysis of a two-site study actually does, and it is not measured here.

There is a specific reason to expect it to be interesting rather than a compromise. A prior on the variance is a supplied input like the third option, but it supplies a distribution rather than a value, so its interval should sit between the supplied one and the exact one — and where exactly it sits is entirely a property of the prior’s tail, since with one degree of freedom the data cannot pull the posterior away from it. The uncomfortable possibility is that the resulting interval is narrow because the prior is confident and reads as though the data had spoken.

What the coverage of such an interval is, against both estimands, under a prior that is right and under one that is not, and whether there is a prior whose interval is as short as the supplied one and as honest as the exact one, is the measurement this field has not made.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Conservative intervalCoverageDegrees of freedomEffective sample sizeEstimandHierarchical modelInterval widthMonte CarloPartial poolingPrior sensitivityRandom-effectsStudy designT intervalVariance components