An interval read beside something else

The pattern a cause leaves on its neighbours

A characteristic that modifies a treatment's effect raises every characteristic that overlaps it, in proportion to the overlap, and fitting that pattern was supposed to say which one is the cause. Fitted with the noise's own correlation it names the largest row on every trial, because the noise draws the same pattern — so the 54.4% at which the largest row is right among nested characteristics is the most any reading of the table can manage, and the honest answer is a set of about five.

Worth reading first: What the 95% refers to.

One characteristic that really matters drew a trial of four hundred whose treatment effect was changed by one of eight baseline characteristics and nothing else, and read its table of interaction tests the way tables are read: row by row, against a threshold. Among eight characteristics that overlap like nested age bands, the real modifier was the table’s largest interaction in 54.4% of trials at the largest modification there is, and another characteristic crossed the Bonferroni line in half of all trials. A family-wise threshold counts rows whose contrast is zero, and at that overlap no row’s contrast is zero.

That essay ended on a hope. A real modifier does not raise one row; it raises its neighbours too, each in proportion to how much it overlaps the real one, and leaves unrelated characteristics where they were. That is a shape across the table, and noise was supposed not to have it. A reading that fitted each candidate’s shape to all eight rows, rather than looking at rows one at a time, might use the table’s structure to answer the question a threshold cannot: which of these characteristics is the cause.

The shape is real, and fitting it is the right idea. What the fit returns is the surprise.

What a modifier does to the rows beside it

The trial is the earlier one. Four hundred patients, half to each arm, a normal outcome, an average effect that gives the whole trial 80% power, so the main comparison’s expected z is 2.80. Each of eight binary characteristics is the side of the median a latent measurement falls on. Characteristic one modifies the effect: δ(1+h)\delta(1+h) on one of its sides and δ(1−h)\delta(1-h) on the other, so the average effect is unchanged and its own interaction statistic has expected value 2.80h2.80h.

Every other characteristic inherits a share. Its two subgroups differ in how many patients they hold from each side of the real modifier, and for median splits the difference in their average effects is exactly the correlation between the two split indicators times the real modifier’s own contrast. So the expected interaction of characteristic kk is

E zk=2.80 h  Ck1,\mathbb{E}\,z_k = 2.80\,h\;C_{k1},

where CC is the matrix of correlations between the split indicators — the table’s overlap, readable from the baseline data before any outcome is seen. Under the hypothesis that characteristic jj is the modifier, the table’s expected shape is the jj-th column of CC scaled by the modification.

The shape one real modifier leaves on a table of eight characteristics, and one trial's table. Expected interaction z when characteristic 1 modifies the effect at h = 1: 2.80, 2.00, 2.00, 2.00, 0.00, 0.00, 0.00, 0.00 — the real row 2.80, its three nested neighbours 0.713 of that, the four unrelated characteristics nothing. One seeded trial of 400 reads 1.91, 1.90, 2.06, 1.95, 0.14, 0.33, 0.39, 2.08.
Fig. 1 Bars: each characteristic’s expected interaction when characteristic 1 modifies the effect at h = 1, in a table whose first four characteristics are nested bands and whose last four are unrelated to anything. Dots: one seeded trial’s eight statistics.

The figure draws that shape for a table built to show it plainly: four characteristics nested at a latent correlation of 0.9, whose split indicators correlate 0.713, and four unrelated to anything. With the first as the real modifier at h=1h = 1, the real row expects 2.80, its three nested neighbours 2.00 each, and the four unrelated characteristics nothing. A cause leaves a footprint, and the footprint is a block.

The noise draws the same block

The dots are one trial, and they are the essay in miniature. The first four rows read 1.91, 1.90, 2.06 and 1.95; the four unrelated rows read 0.14, 0.33, 0.39 and 2.08. The block is plainly raised. The table’s largest row is characteristic eight, which modifies nothing and whose interaction points the other way.

The eye wants to say that four rows at about two, all of them overlapping one another, is better evidence than one row at 2.08, and that the block names the first four even if it cannot say which of them. The second half of that is right. The first half is the mistake the shape reading was built on.

The four nested rows do not carry four pieces of evidence. Their statistics are computed from overlapping groups of the same patients, so their noise is correlated by the same matrix CC that sets their expected values: two nested interaction statistics correlate about 0.71 whether or not anything is real. A trial in which the shared part of the block’s noise happened to come out high raises all four rows together, and it looks exactly like a modifier inside the block. A modifier leaves the pattern CejC e_j, and the noise in a table with nothing in it leaves draws whose covariance is CC — the same pattern, column by column.

That is what the effective count of a subgroup table found from the other side: eight nested characteristics are worth about four independent tests, because their statistics move together. A block of four raised rows is about one raised row’s worth of evidence, and the question is which single row of the table is best explained as the cause.

Fitting the shape gives back the largest row

The reading the earlier essay proposed can be written down exactly. Under candidate jj, the table’s statistics are normal with mean μ Cej\mu\,C e_j for some unknown modification μ\mu and covariance CC. Fitting μ\mu by generalised least squares and scoring each candidate by how well its shape fits gives the statistic

Sj=∣mj⊤C−1z∣mj⊤C−1mj,mj=Cej.S_j = \frac{\lvert m_j^{\top} C^{-1} z\rvert}{\sqrt{m_j^{\top} C^{-1} m_j}}, \qquad m_j = C e_j .

And C−1mj=C−1Cej=ejC^{-1} m_j = C^{-1} C e_j = e_j. The numerator is ∣zj∣\lvert z_j\rvert and the denominator is Cjj=1\sqrt{C_{jj}} = 1. The fitted shape’s score for candidate jj is candidate jj’s own interaction statistic, and the best-fitting shape is the largest row.

That holds for every overlap, not just the equal one, and it holds without knowing CC at all: the matrix cancels before any number is computed from it. Run on two thousand trials of each structure in the figures below, the shape fitted this way names the largest row on every trial, with no exceptions to count.

It is also the best attribution there is. With the eight characteristics equally likely to be the modifier and the modification of size μ\mu, the posterior chance that characteristic jj is the one is proportional to exp⁡(μzj−μ2/2)\exp(\mu z_j - \mu^2/2), and if the direction is not known, to cosh⁡(μzj)\cosh(\mu z_j). Both rise with ∣zj∣\lvert z_j\rvert alone. Whatever the modification’s size, the characteristic most likely to be the cause is the one with the largest interaction, and the rule that names it is right as often as any rule can be. The 54.4% is not the weakness of a row-by-row reading that a cleverer reading could repair. Among eight nested characteristics at a modification as large as the main effect, it is the ceiling.

Two readings that use more of the table and get less

There are two ways to use the overlap that do not cancel, and both are natural enough to be what an analysis would actually do.

The first fits the shape as though the rows were independent: each candidate’s column of CC correlated against the table, with no allowance for the rows’ shared noise. It is the literal form of “score the table against each candidate’s pattern”. The second is the one the earlier essay offered as the design that separates overlapping characteristics — every interaction fitted in one model, each estimated holding the others fixed, which in this notation is C−1zC^{-1} z standardised.

How often four readings of a table of eight name the one real modifier, at a modification as large as the main effect. Over 2,000 trials of 400 a structure at h = 1. eight nested: largest row 54.4%, pattern as if independent 53.6%, joint fit 45.6%; four nested + four unrelated, real one nested: largest row 63.5%, pattern as if independent 63.8%, joint fit 50.1%; four nested + four unrelated, real one unrelated: largest row 83.4%, pattern as if independent 74.3%, joint fit 82.2%; eight unrelated: largest row 82.4%, pattern as if independent 81.6%, joint fit 82.4%. The shape fitted with the overlap's own correlation names the largest row on every trial.
Fig. 2 How often each reading names the one real modifier at h = 1, over two thousand trials of four hundred for each of four structures: the largest row, which the shape fitted with the overlap’s own correlation always equals; the shape fitted as though the rows were independent; and every interaction fitted jointly.

The pattern fitted as if independent does no harm where the overlap is even. Among eight nested characteristics it names the real one in 53.6% of trials against the largest row’s 54.4%, since with every pair overlapping alike each candidate’s pattern differs from the next only in its own row. Where the overlap is uneven, it is drawn to the block. With the real modifier among the four unrelated characteristics and four nested ones beside it, the largest row names it in 83.4% of trials and the independent-pattern fit in 74.3% — and the trials it loses go to the nested block, which it names 17.5% of the time against the largest row’s 7.4%. A block of correlated noise looks to it like four agreeing witnesses.

The joint fit pays the price a point ordinary on every axis measured as variance inflation. Holding three nested neighbours fixed leaves each nested interaction estimated from the small share of patients on whom the neighbours disagree, and the real modifier’s adjusted statistic shrinks with it. Among eight nested characteristics the joint fit names the real one in 45.6% of trials against 54.4%; with the real one inside a block of four, in 50.1% against 63.5%. Where nothing overlaps, holding the others fixed changes nothing and the two readings agree at 82.4%.

So both are worse, and for opposite reasons. The independent-pattern fit counts correlated rows as agreeing evidence and is attracted to them; the joint fit throws away the shared part of each row and with it most of the real modifier’s signal. The likelihood keeps the shared part and discounts it by exactly the right amount, and what it keeps is the row.

The ceiling, by size and by structure

The largest row is therefore the curve to read, and it can be read as the most any analysis of this table can do.

The most often any reading can name the real modifier, by its size and the table's overlap. Over 2,000 trials of 400 a point, at h = 0, 0.25, 0.5, 0.75, 1: eight nested, 12.3%, 16.6%, 28.5%, 41.1%, 54.4%; four nested, four unrelated, 9.4%, 14.4%, 30.3%, 47.3%, 63.5%; eight unrelated, 12.3%, 20.6%, 39.3%, 63.6%, 82.4%. One in eight is chance.
Fig. 3 How often the largest row is the real modifier, against the modification’s size h, for eight nested characteristics, four nested beside four unrelated with the real one nested, and eight unrelated. No reading names it more often at any size.

With eight unrelated characteristics the real modifier is the largest row in 39.3% of trials at h=0.5h = 0.5 and 82.4% at h=1h = 1. Nesting it among seven overlapping neighbours cuts those to 28.5% and 54.4%, the numbers the earlier essay counted. Nesting it among three, with four unrelated characteristics beside them, gives 30.3% and 63.5%: fewer neighbours to inherit the signal, so a higher ceiling, and the loss is still most of the way to the fully nested table.

The shape of these curves is the attribution question’s whole answer at this trial size. At a modification as large as the main effect, h=0.5h = 0.5, no reading names the real characteristic among nested neighbours even three times in ten, because its expected statistic of 1.40 is not far above its neighbours’ 1.00 and every row’s noise is a unit wide. A trial of four hundred cannot tell which of two strongly overlapping characteristics carries a modification of that size, and no method of reading the table it produced can make it.

What a table with nothing in it says

A reading that always names a characteristic needs something in front of it that says whether to name one at all. The attribution step answers “which”, and it answers it on a table with nothing in it too.

With no real modifier, the largest row is spread nearly evenly over eight nested characteristics — 12.3% of trials name the first, one in eight being chance. In the table of four nested and four unrelated, it is not even: the unrelated characteristics are the largest row 14.2% to 16.5% of the time and the nested ones 9.4% to 10.0%. Four correlated statistics have a smaller maximum than four independent ones, so a table with nothing in it tends to point away from its overlapping block, and a table with something in the block has to overcome that before its block is named. The arithmetic is the same one that thinned the subgroup table’s family: correlated rows are fewer chances.

So the family’s question comes first and the shape’s second. Whether any characteristic modifies the effect is a question about the largest row against a threshold set for the family, the one what the correction corrects describes, and the overlap lowers that threshold by an amount the baseline data supply — the same effective count how many analyses there really were computed for an analyst’s own family. Only after it is crossed does “which one” mean anything — and the answer to that second question is not a row.

A set rather than a winner

If the largest row is the best single answer and is wrong half the time, the honest answer is the set of characteristics the table cannot tell apart. The posterior above gives it directly. With the eight equally likely beforehand and a normal prior on the modification’s expected z with a scale of 2.80 — a modification up to about the size of the main effect — the posterior weight on characteristic jj is proportional to exp⁡ ⁣(s2zj2/2(1+s2))\exp\!\big(s^2 z_j^2 / 2(1+s^2)\big) with s=2.80s = 2.80, and the smallest set holding 90% of the posterior is the report. It is the same normal-prior arithmetic the weight that decides worked through for a group’s mean, applied to a choice among eight rather than to a number.

How many characteristics an honest attribution has to name, by the modification's size and the table's overlap. The smallest set holding 90% of the posterior over which characteristic is the modifier. Average size at h = 0, 0.25, 0.5, 0.75, 1: eight nested, 7.22, 7.06, 6.63, 5.89, 4.96; four nested, four unrelated, 6.84, 6.67, 6.16, 5.27, 4.10; eight unrelated, 6.75, 6.63, 6.22, 5.30, 3.86. It holds the real modifier in 93.0% and 95.5%, 92.8% and 97.4%, 93.5% and 98.9% of trials at h = 0.5 and 1.
Fig. 4 The average number of characteristics in the smallest set holding 90% of the posterior over which one is the modifier, against the modification’s size, for the three structures. With nothing real the set holds nearly all eight.

For the trial in the first figure the set is five characteristics: eight, three, four, one and two, with posterior weights from 0.213 down to 0.155. It holds the real modifier and the noise row that led the table, and it leaves out the three unrelated rows that are visibly flat. That is the right report of that table. It says that the modification is in the nested block or on characteristic eight, and that four hundred patients cannot say which.

Over trials the set behaves as a set should. At h=1h = 1 it holds the real modifier in 95.5% of trials among eight nested characteristics, 97.4% with the real one inside a block of four, and 98.9% among eight unrelated ones — above its 90% because the prior spreads its weight over smaller modifications than the one drawn. Its average size is 4.96 characteristics when all eight are nested, 4.10 for the block of four and 3.86 when nothing overlaps; at h=0.5h = 0.5 it holds 6.63, 6.16 and 6.22. With nothing real it holds 7.22 of the eight nested characteristics, which is the table saying, correctly, that it has nothing to attribute.

That set is wider than a reader of a subgroup table expects, and it is the width the data support. The ceiling figure says the same thing as a hit rate. The set says it in the form a report can carry: the characteristics compatible with the table, rather than a winner that is wrong as often as it is right. It is the same move intervals for the findings made for effect sizes after selection — report the spread of what the data allow, not the point the selection landed on.

What a protocol can take from it

Read the largest row, and know that it is the ceiling. No re-weighting of a subgroup table names the real modifier more often than its largest interaction, provided one characteristic is the modifier and none is favoured beforehand. Effort spent on a cleverer reading of the same table is effort spent on nothing; the overlap’s effect is in the data, not in the analysis.

Do not score candidates against the table as though its rows were independent. That reading is the one a spreadsheet invites, and it moves attributions into whichever block of characteristics overlaps most, at a cost of nine points when the real modifier is outside it.

Do not fit every interaction jointly to decide which one matters. The joint fit answers a different question — which characteristic modifies the effect holding the others fixed — and among nested characteristics it answers it with a fraction of the information. It names the real one less often than the largest row at every structure with overlap.

Report a set. The posterior set costs one line of arithmetic per characteristic and uses nothing but the interaction statistics. A subgroup finding reported as “one of characteristics one to four, or eight” is less quotable than “characteristic eight”, and it is true. Naming a handful in advance remains the one design change that shrinks the set rather than describing it, because a characteristic not tabulated cannot be in it.

The algebra, and what was counted

Exact: under a single modifier with split indicators correlated CC, the table’s expected shape is the modifier’s column of CC; the generalised-least-squares score of candidate jj is ∣zj∣\lvert z_j\rvert for every CC; and with equal prior weight on each characteristic the posterior is increasing in ∣zj∣\lvert z_j\rvert, so the largest row maximises the chance of a correct attribution at every modification size.

Counted, over two thousand trials of four hundred for each structure and size: the largest row names the real modifier in 54.4% of trials among eight nested characteristics at h=1h = 1, 63.5% inside a block of four and 82.4% among unrelated ones; the independent-pattern fit in 53.6%, 63.8% and 81.6%, and 74.3% against 83.4% when the real one sits outside a nested block; the joint fit in 45.6%, 50.1% and 82.4%; and the 90% posterior set holds the real one in 95.5% of nested trials with 4.96 members on average.

Not claimed: that a real modifier comes alone. Everything above assumes one characteristic carries the modification, and the equivalence between the fitted shape and the largest row rests on it. Not claimed either that the prior’s scale is right — a different scale changes the set’s size, not which characteristic heads it.

Still open: two modifiers at once

The equivalence holds because the hypothesis is one cause. With two real modifiers — age and a comorbidity, say, each changing the effect — the table’s expected shape is a combination of two columns of CC, and the largest row is no longer the likelihood’s answer: a pair of moderately raised rows whose overlaps explain the whole table can be better supported than any single row. That is where fitting the shape stops collapsing to the obvious reading, and where the joint fit’s lost information might be bought back by the pairs it can separate.

How often a two-modifier fit names both, how often it invents a second where there is one, and how large the set of plausible pairs is among eight characteristics — twenty-eight pairs rather than eight rows — have not been measured here. They decide whether a subgroup table can ever report more than one characteristic with a straight face.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

CorrelationFamilywise error rateGeneralised least squaresInteractionLikelihoodMultiple comparisonsPosteriorPriorSubgroup analysisVariance inflation