An interval read beside something else

One characteristic that really matters

A trial of four hundred tabulates eight characteristics and one of them genuinely changes the treatment's effect — the whole effect on one side, none on the other. Tested alone at 5% the real interaction is found 81.0% of the time; at the Bonferroni threshold that protects the table, 53.6%. It is the table's largest interaction in 82.4% of trials when the characteristics are independent and in 54.4% when they overlap like nested age bands, and in that case another characteristic crosses the Bonferroni line in half of all trials and is the top finding in a quarter. A family-wise threshold counts false rows; it cannot say which of two overlapping rows is the cause.

Worth reading first: What the 95% refers to.

Sixteen subgroups and one effect drew a trial whose treatment effect was the same for every patient and read its subgroup table. Eight characteristics gave eight interaction tests, some one of them reached 5% in 34.6% of trials, and a Bonferroni reading over the eight held the family near 5% whatever the characteristics’ overlap. Every interaction in those tables was false, so every number there was a rate of error.

A table is worth reading only if something in it might be real, and the essay ended on that case: one characteristic genuinely modifies the effect and seven do not. A family-wise threshold then protects against the seven and costs power on the one. What that trade looks like — how much power the one keeps, whether it stands out from the seven, and how often a trial with one real modifier reports a different one — is what turns a protocol’s list of subgroups into a design with known properties.

How often one genuine interaction in a table of eight is found, by its size, at latent correlation 0A trial of 400 with 80% power for its main effect, eight characteristics, the first modifying the effect. At h = 0, 0.25, 0.5, 0.75, 1 the real interaction crosses 5% in 4.2%, 10.3%, 29.2%, 56.6%, 81.0% of trials; Bonferroni's 2.734 in 0.4%, 1.8%, 8.7%, 26.6%, 53.6%; the exact threshold for this overlap, 2.727, in 0.4%, 1.9%, 8.8%, 27.0%, 53.8%.00.2500.5000.750100.2500.5000.7501h: the effect is δ(1 + h) on one side of the real characteristic, δ(1 − h) on the othershare of trials in which the real interaction crosses5% on its ownBonferroni over eight, 2.73exact for the overlap, 2.732,000 trials a point, latent correlation 0protecting eight rows costs the real one
Fig. 1 How often the one genuine interaction in a table of eight crosses 5% on its own, the Bonferroni threshold over eight, and the exact threshold for the characteristics’ overlap, against the interaction’s size. The slider sets how much the characteristics overlap.

A modifier with a size

The trial is the one the earlier table used: four hundred patients, half to each arm, a normal outcome, and an average treatment effect that gives the whole trial 80% power. Eight binary characteristics, each the side of the median a latent measurement falls on, with latents that share a common factor at a correlation of 0, 0.5 or 0.9. Now the first characteristic modifies the effect: on one of its sides the effect is δ(1+h)\delta(1+h) and on the other δ(1−h)\delta(1-h), so the average effect, and the whole trial’s power, are unchanged.

The size of the modification is hh. At h=0.5h = 0.5 the effect is three times as large on one side as on the other, and the difference between them equals the average effect itself. At h=1h = 1 the treatment works fully on one side and not at all on the other — the largest modification a subgroup claim usually has in mind. Because each side holds half the patients, the interaction statistic has an expected value of hh times the whole trial’s expected z of 2.80: 1.40 at h=0.5h = 0.5 and 2.80 at h=1h = 1. An interaction as large as the main effect is, to a trial sized for the main effect, a test with about a third of its power.

What the real row keeps

Tested on its own at 5% — as if the protocol had named it in advance as the one characteristic expected to matter — the real interaction is found in 29.2% of trials at h=0.5h = 0.5 and 81.0% at h=1h = 1, with independent characteristics. That is the power its expected z gives it, and it is the most any analysis of this trial can have for this question.

Read as one row of a table protected by Bonferroni, it must cross 2.732.73 rather than 1.96. At h=0.5h = 0.5 it does so in 8.7% of trials; at h=1h = 1, in 53.6%. The protection costs the real row two thirds of its chance at a modification as large as the main effect, and a third of its chance at the largest modification there is.

The exact threshold for the overlap — the level at which some one of eight equicorrelated interaction statistics crosses 5% of the time — buys little of that back. With independent characteristics it is 2.727, all but Bonferroni’s own. At a latent correlation of 0.5 it is 2.697; at 0.9, 2.545, and there it lifts the real row’s power at h=1h = 1 from 54.1% to 61.2%. Overlap makes the family smaller and the exact threshold lower, but a table of eight nested age bands is still a family of about four, and a family of four still charges its real member for the other three.

Without the protection

The alternative a table’s authors often take is to read every row at 5% and let the reader discount. With one real modifier that reading finds more, and it finds more of everything.

At h=1h = 1 with independent characteristics, the unprotected table reports the real interaction in 81.0% of trials. It reports it alone in 54.3%; in the other 26.7% a null row crosses 5% beside it, and in 11.1% of all trials a null row has a larger interaction than any real finding and heads the list. At h=0.5h = 0.5 the real row crosses in 29.2% of trials and some null row in 31.1% — a table whose one real modification is as large as the main effect reports a false interaction about as often as the true one, and a reader has nothing in the table to tell them apart.

That is the arithmetic what the correction corrects described for any family and the price of control put a rate of exchange on: each row a table adds is a further chance at a false finding that competes with the true one for the reader’s attention, and the protection that removes those chances removes most of the true one’s too. The table cannot be read at a threshold that keeps the real row’s power and drops the false rows, because at these sizes their statistics overlap.

Is the real row the largest?

A reader of a table does not only ask which rows cross a line. The row with the largest interaction is the one that is discussed, and the natural hope is that a real modifier will at least be the table’s top row.

How often the one genuine interaction is the largest of eight, by its size and the characteristics' overlap. Over 2,000 trials a point. At h = 0, 0.25, 0.5, 0.75, 1: latent correlation 0, 12.3%, 20.6%, 39.3%, 63.6%, 82.4%; latent correlation 0.5, 11.7%, 19.7%, 35.6%, 56.3%, 76.1%; latent correlation 0.9, 12.3%, 16.6%, 28.5%, 41.1%, 54.4%. With no real interaction one characteristic in eight is the largest by chance.
Fig. 2 How often the one genuine interaction is the largest of the eight in the table, against its size, for three overlaps between the characteristics. With nothing real, each of eight rows is the largest one time in eight.

With independent characteristics it is the largest in 39.3% of trials at h=0.5h = 0.5 and 82.4% at h=1h = 1. With a genuine modification as large as the main effect, then, the real row is out-ranked by one of seven null rows three trials in five. A table’s top row is evidence about which characteristic matters in proportion to how large the modification is, and at the sizes a trial sized for its main effect can hope to see, that proportion is modest.

Overlap makes it worse, for a reason that is the centre of this essay. At a latent correlation of 0.9 the real row is the largest in 28.5% of trials at h=0.5h = 0.5 and 54.4% at h=1h = 1 — little better than a coin, at the largest modification the design allows.

Why a neighbour inherits the signal

A characteristic that overlaps the real modifier is not a null characteristic in the table’s sense. Its two subgroups differ in how many patients from each side of the real modifier they contain, so its interaction contrast picks up a share of the real difference. For median splits the share is the correlation between the two characteristics’ indicators, (2/π)arcsin⁡ρ(2/\pi)\arcsin\rho: 0.333 at a latent correlation of 0.5 and 0.713 at 0.9. At h=1h = 1, where the real interaction’s expected z is 2.80, every neighbour at 0.9 has an expected z of about 2.0 — more than enough to cross 5% on its own, and seven of them sit beside the real row with their own noise.

One trial's eight interaction statistics when one characteristic genuinely modifies the effect and the characteristics overlap. Latent correlation 0.9, the first characteristic modifying the effect at h = 1. Interaction |z| by characteristic: 1: 2.67, 2: 2.16, 3: 1.55, 4: 2.49, 5: 1.99, 6: 1.33, 7: 2.58, 8: 2.88. The Bonferroni threshold is 2.73; the 5% threshold is 1.96.
Fig. 3 One trial with characteristics overlapping at a latent correlation of 0.9 and the first genuinely modifying the effect at h = 1: each characteristic’s interaction statistic, with the 5% and Bonferroni thresholds. The real row stops short of the Bonferroni line; a neighbour crosses it.

The trial drawn above is ordinary for that setting. The real row’s statistic is 2.67 and falls just short of the Bonferroni threshold; characteristic eight, which modifies nothing, reaches 2.88 and crosses it, and characteristics four and seven reach 2.49 and 2.58. A protocol that reads the table correctly by every rule of multiple testing reports characteristic eight as the one subgroup with a significant interaction.

Nothing in that report is an error of the kind family-wise control is built to prevent. Characteristic eight’s interaction is not zero; its subgroups genuinely differ in their average effect, because they differ in their mix of the real modifier’s two sides. The family-wise rate counts rows whose contrast is truly null, and at this overlap there are none.

The inheritance is proportional, which is why moderate overlap already matters. At a latent correlation of 0.5 the indicators correlate 0.333, and a neighbour of a real modifier at h=1h = 1 has an expected interaction z of about 0.93 — nowhere near a threshold on its own. But there are seven of them, each a draw centred at 0.93 rather than at zero, and seven chances at a statistic whose centre has moved a third of the way to 2.8 is a family whose false rows are not false. The Bonferroni threshold was set for seven rows centred at zero, and it is being asked to hold seven rows centred elsewhere.

The same arithmetic runs the other way when the real modifier is weak. At h=0.25h = 0.25 its own expected z is 0.70 and its neighbours’ at a latent correlation of 0.9 is 0.50: the whole table has shifted a little, nothing stands out, and the real row is the largest in 16.6% of trials against the one in eight that chance gives any row. A modest real modification in an overlapping table does not look like one row rising; it looks like the table leaning.

What the family reports

The consequence can be counted directly, as how often a Bonferroni reading of the table reports a characteristic that is not the real one.

What a Bonferroni reading of eight interactions reports when one is real, latent correlation 0.9. Over 2,000 trials a point, at h = 0, 0.25, 0.5, 0.75, 1: another characteristic crosses in 3.1%, 6.2%, 14.3%, 29.8%, 50.2%; the largest crossing is another characteristic in 2.9%, 5.7%, 11.9%, 20.4%, 25.7%; only the real one crosses in 0.1%, 0.6%, 2.9%, 8.7%, 11.8%.
Fig. 4 When one characteristic genuinely modifies the effect: how often some other characteristic crosses the Bonferroni threshold, how often the largest crossing row is another characteristic, and how often only the real row crosses, against the modification’s size.

With independent characteristics the family behaves as advertised. Some other row crosses Bonferroni’s threshold in 4.3% to 5.5% of trials at every size of modification, the largest crossing row is another characteristic in at most 4.8%, and at h=1h = 1 the real row crosses alone in 50.7% of trials.

At a latent correlation of 0.9 it does not. Some other characteristic crosses in 14.3% of trials at h=0.5h = 0.5 and 50.2% at h=1h = 1. The table’s largest crossing row is another characteristic in 11.9% and 25.7%. The real row crosses alone — the one outcome that identifies it — in 2.9% and 11.8%. The exact threshold for the overlap makes the trade sharper: at h=1h = 1 it lifts the real row to 61.2% and lets another row lead the table in 29.4% of trials.

At a latent correlation of 0.5, a moderate overlap, the numbers sit between: another row crosses in 22.0% of trials at h=1h = 1 and leads the table in 9.2%, and the real row crosses alone in 37.0%.

What protecting eight rows costs in patients

The power figures translate into sample sizes directly, because the interaction’s expected z grows with the square root of the trial’s size. For the real row to have 80% power on its own at 5%, its expected z must reach 1.96+0.84=2.801.96 + 0.84 = 2.80; as a row of eight under Bonferroni, 2.73+0.84=3.582.73 + 0.84 = 3.58.

At h=1h = 1, where a trial of four hundred already has an expected interaction z of 2.80, the prespecified test needs exactly those four hundred patients and the protected table needs 652. At h=0.5h = 0.5, where the expected z at four hundred is 1.40, the prespecified test needs 1,600 patients and the protected table 2,607 — six and a half times the trial that was sized for the main effect.

So a subgroup question has a price that can be stated before the trial: four times the patients to answer one interaction as large as the main effect, six and a half to answer it as a row of eight. A protocol that lists eight characteristics and sizes for the main effect has bought none of it, and its table should be read as the exploratory family it is. How many analyses there really were found that an analyst’s own family is usually larger than the one reported; a subgroup table is the case where the family is on the page and its price can be paid in the design rather than discovered in the reading.

A cause and its correlates

That is the general shape of the problem, and it is not specific to subgroups. A table of overlapping characteristics measures associations with the effect’s size, and a characteristic correlated with the real modifier is associated with it. Two intervals that overlap and significant in one, not in the other showed that comparing verdicts row by row misreads the evidence; the measurement here shows that even the right comparison, the interaction test read as a family, cannot say which of two correlated characteristics carries the modification.

Multiple testing controls how often a table reports a row whose contrast is zero. It says nothing about attribution, because attribution is a question about causes and the table contains only contrasts. A subgroup inside its own trial found that a subgroup is correlated f\sqrt f with its whole and should be tested against the rest; the same arithmetic applied between two subgroups says that an interaction found on one of two strongly overlapping characteristics has been found on both, and the table’s choice between them is made by noise.

The one design that separates them is the one that breaks the overlap: characteristics tabulated jointly, so the real modifier’s interaction is estimated holding its neighbour fixed. That is a regression with both interaction terms, and it pays for the separation with variance, since two strongly correlated interaction terms are estimated far less precisely together than either alone — the multivariable cost a point ordinary on every axis measured as variance inflation. At a latent correlation of 0.9, a trial of four hundred cannot afford it.

What a protocol can do with this

Name the characteristic in advance, and test it alone. A single prespecified interaction keeps its full power — 81.0% at the largest modification, 29.2% at one as large as the main effect — and the table’s other rows become the exploratory family they are. Naming a handful in advance priced the list: each additional named characteristic costs the others power, and a list of one keeps it all.

Expect the real row to lose rank. At a modification as large as the main effect, the real row is the table’s largest two times in five with independent characteristics and less than three times in ten with nested ones. A protocol that plans to follow up “the largest interaction” has planned to follow up a null row most of the time.

Treat overlapping characteristics as one question. Two age thresholds, or a comorbidity and the drug that treats it, are one modifier measured twice. Report a single interaction for the pair, or report both with the statement that the table cannot choose between them; the Bonferroni line does not make that choice either.

Size the trial for the interaction if the interaction is the question. An interaction test between halves needs four times the patients of the main comparison to have the same power at a difference as large as the main effect, which the verdicts watch the threshold and the earlier subgroup essays priced from several directions. A trial of four hundred sized for its average effect has 29.2% power for that interaction on its own and 8.7% as a row of eight.

What is exact and what is counted

Exact: the interaction statistic of a characteristic that modifies the effect by hh has expected value hh times the whole trial’s expected z; a characteristic overlapping it inherits the share (2/π)arcsin⁡ρ(2/\pi)\arcsin\rho of that; and the family’s exact threshold at the overlap, from the equicorrelated integral, is 2.727, 2.697 and 2.545 at latent correlations of 0, 0.5 and 0.9.

Counted, over two thousand trials of four hundred at each setting: the real row’s power of 81.0% at 5% and 53.6% at Bonferroni’s threshold at h=1h = 1 with independent characteristics; its rank as the largest in 82.4% of those trials and 54.4% at a latent correlation of 0.9; and, at that overlap, another characteristic crossing Bonferroni’s threshold in 50.2% of trials and leading the table in 25.7%.

Not claimed: that real modifiers come one at a time, or that they act on median splits. A modifier acting on a continuous measurement is diluted by any binary cut of it, which lowers every power here; two genuine modifiers make the attribution question harder still. Not claimed either that the overlap is known exactly in practice — it is estimable from baseline data, as the earlier table essay noted, and the exact threshold depends on it.

Still open: a table read by its shape

Every reading here looks at rows one at a time against a threshold. A table in which one characteristic genuinely modifies the effect has a shape a row-by-row reading ignores: the real row and its neighbours are raised together, in proportion to their overlap with the real one, while characteristics unrelated to it are not. A modifier leaves a pattern across the table, and noise does not leave that pattern.

A reading that fits the pattern — each candidate modifier predicting every row’s interaction through the overlap matrix, and the table scored against each candidate — would use all eight rows to answer the attribution question the threshold cannot. Whether such a reading picks the real modifier more often than the largest row does, how much of the overlap it needs to know, and what it reports when nothing is real, have not been measured here.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

BonferroniCorrelationEffective number of testsFamilywise error rateInteractionMultiple comparisonsStatistical powerSubgroup analysis