The pattern a cause leaves on its neighbours
Worth reading first: What the 95% refers to.
One characteristic that really matters drew a trial of four hundred whose treatment effect was changed by one of eight baseline characteristics and nothing else, and read its table of interaction tests the way tables are read: row by row, against a threshold. Among eight characteristics that overlap like nested age bands, the real modifier was the table’s largest interaction in 54.4% of trials at the largest modification there is, and another characteristic crossed the Bonferroni line in half of all trials. A family-wise threshold counts rows whose contrast is zero, and at that overlap no row’s contrast is zero.
That essay ended on a hope. A real modifier does not raise one row; it raises its neighbours too, each in proportion to how much it overlaps the real one, and leaves unrelated characteristics where they were. That is a shape across the table, and noise was supposed not to have it. A reading that fitted each candidate’s shape to all eight rows, rather than looking at rows one at a time, might use the table’s structure to answer the question a threshold cannot: which of these characteristics is the cause.
The shape is real, and fitting it is the right idea. What the fit returns is the surprise.
What a modifier does to the rows beside it
The trial is the earlier one. Four hundred patients, half to each arm, a normal outcome, an average effect that gives the whole trial 80% power, so the main comparison’s expected z is 2.80. Each of eight binary characteristics is the side of the median a latent measurement falls on. Characteristic one modifies the effect: on one of its sides and on the other, so the average effect is unchanged and its own interaction statistic has expected value .
Every other characteristic inherits a share. Its two subgroups differ in how many patients they hold from each side of the real modifier, and for median splits the difference in their average effects is exactly the correlation between the two split indicators times the real modifier’s own contrast. So the expected interaction of characteristic is
where is the matrix of correlations between the split indicators — the table’s overlap, readable from the baseline data before any outcome is seen. Under the hypothesis that characteristic is the modifier, the table’s expected shape is the -th column of scaled by the modification.
The figure draws that shape for a table built to show it plainly: four characteristics nested at a latent correlation of 0.9, whose split indicators correlate 0.713, and four unrelated to anything. With the first as the real modifier at , the real row expects 2.80, its three nested neighbours 2.00 each, and the four unrelated characteristics nothing. A cause leaves a footprint, and the footprint is a block.
The noise draws the same block
The dots are one trial, and they are the essay in miniature. The first four rows read 1.91, 1.90, 2.06 and 1.95; the four unrelated rows read 0.14, 0.33, 0.39 and 2.08. The block is plainly raised. The table’s largest row is characteristic eight, which modifies nothing and whose interaction points the other way.
The eye wants to say that four rows at about two, all of them overlapping one another, is better evidence than one row at 2.08, and that the block names the first four even if it cannot say which of them. The second half of that is right. The first half is the mistake the shape reading was built on.
The four nested rows do not carry four pieces of evidence. Their statistics are computed from overlapping groups of the same patients, so their noise is correlated by the same matrix that sets their expected values: two nested interaction statistics correlate about 0.71 whether or not anything is real. A trial in which the shared part of the block’s noise happened to come out high raises all four rows together, and it looks exactly like a modifier inside the block. A modifier leaves the pattern , and the noise in a table with nothing in it leaves draws whose covariance is — the same pattern, column by column.
That is what the effective count of a subgroup table found from the other side: eight nested characteristics are worth about four independent tests, because their statistics move together. A block of four raised rows is about one raised row’s worth of evidence, and the question is which single row of the table is best explained as the cause.
Fitting the shape gives back the largest row
The reading the earlier essay proposed can be written down exactly. Under candidate , the table’s statistics are normal with mean for some unknown modification and covariance . Fitting by generalised least squares and scoring each candidate by how well its shape fits gives the statistic
And . The numerator is and the denominator is . The fitted shape’s score for candidate is candidate ’s own interaction statistic, and the best-fitting shape is the largest row.
That holds for every overlap, not just the equal one, and it holds without knowing at all: the matrix cancels before any number is computed from it. Run on two thousand trials of each structure in the figures below, the shape fitted this way names the largest row on every trial, with no exceptions to count.
It is also the best attribution there is. With the eight characteristics equally likely to be the modifier and the modification of size , the posterior chance that characteristic is the one is proportional to , and if the direction is not known, to . Both rise with alone. Whatever the modification’s size, the characteristic most likely to be the cause is the one with the largest interaction, and the rule that names it is right as often as any rule can be. The 54.4% is not the weakness of a row-by-row reading that a cleverer reading could repair. Among eight nested characteristics at a modification as large as the main effect, it is the ceiling.
Two readings that use more of the table and get less
There are two ways to use the overlap that do not cancel, and both are natural enough to be what an analysis would actually do.
The first fits the shape as though the rows were independent: each candidate’s column of correlated against the table, with no allowance for the rows’ shared noise. It is the literal form of “score the table against each candidate’s pattern”. The second is the one the earlier essay offered as the design that separates overlapping characteristics — every interaction fitted in one model, each estimated holding the others fixed, which in this notation is standardised.
The pattern fitted as if independent does no harm where the overlap is even. Among eight nested characteristics it names the real one in 53.6% of trials against the largest row’s 54.4%, since with every pair overlapping alike each candidate’s pattern differs from the next only in its own row. Where the overlap is uneven, it is drawn to the block. With the real modifier among the four unrelated characteristics and four nested ones beside it, the largest row names it in 83.4% of trials and the independent-pattern fit in 74.3% — and the trials it loses go to the nested block, which it names 17.5% of the time against the largest row’s 7.4%. A block of correlated noise looks to it like four agreeing witnesses.
The joint fit pays the price a point ordinary on every axis measured as variance inflation. Holding three nested neighbours fixed leaves each nested interaction estimated from the small share of patients on whom the neighbours disagree, and the real modifier’s adjusted statistic shrinks with it. Among eight nested characteristics the joint fit names the real one in 45.6% of trials against 54.4%; with the real one inside a block of four, in 50.1% against 63.5%. Where nothing overlaps, holding the others fixed changes nothing and the two readings agree at 82.4%.
So both are worse, and for opposite reasons. The independent-pattern fit counts correlated rows as agreeing evidence and is attracted to them; the joint fit throws away the shared part of each row and with it most of the real modifier’s signal. The likelihood keeps the shared part and discounts it by exactly the right amount, and what it keeps is the row.
The ceiling, by size and by structure
The largest row is therefore the curve to read, and it can be read as the most any analysis of this table can do.
With eight unrelated characteristics the real modifier is the largest row in 39.3% of trials at and 82.4% at . Nesting it among seven overlapping neighbours cuts those to 28.5% and 54.4%, the numbers the earlier essay counted. Nesting it among three, with four unrelated characteristics beside them, gives 30.3% and 63.5%: fewer neighbours to inherit the signal, so a higher ceiling, and the loss is still most of the way to the fully nested table.
The shape of these curves is the attribution question’s whole answer at this trial size. At a modification as large as the main effect, , no reading names the real characteristic among nested neighbours even three times in ten, because its expected statistic of 1.40 is not far above its neighbours’ 1.00 and every row’s noise is a unit wide. A trial of four hundred cannot tell which of two strongly overlapping characteristics carries a modification of that size, and no method of reading the table it produced can make it.
What a table with nothing in it says
A reading that always names a characteristic needs something in front of it that says whether to name one at all. The attribution step answers “which”, and it answers it on a table with nothing in it too.
With no real modifier, the largest row is spread nearly evenly over eight nested characteristics — 12.3% of trials name the first, one in eight being chance. In the table of four nested and four unrelated, it is not even: the unrelated characteristics are the largest row 14.2% to 16.5% of the time and the nested ones 9.4% to 10.0%. Four correlated statistics have a smaller maximum than four independent ones, so a table with nothing in it tends to point away from its overlapping block, and a table with something in the block has to overcome that before its block is named. The arithmetic is the same one that thinned the subgroup table’s family: correlated rows are fewer chances.
So the family’s question comes first and the shape’s second. Whether any characteristic modifies the effect is a question about the largest row against a threshold set for the family, the one what the correction corrects describes, and the overlap lowers that threshold by an amount the baseline data supply — the same effective count how many analyses there really were computed for an analyst’s own family. Only after it is crossed does “which one” mean anything — and the answer to that second question is not a row.
A set rather than a winner
If the largest row is the best single answer and is wrong half the time, the honest answer is the set of characteristics the table cannot tell apart. The posterior above gives it directly. With the eight equally likely beforehand and a normal prior on the modification’s expected z with a scale of 2.80 — a modification up to about the size of the main effect — the posterior weight on characteristic is proportional to with , and the smallest set holding 90% of the posterior is the report. It is the same normal-prior arithmetic the weight that decides worked through for a group’s mean, applied to a choice among eight rather than to a number.
For the trial in the first figure the set is five characteristics: eight, three, four, one and two, with posterior weights from 0.213 down to 0.155. It holds the real modifier and the noise row that led the table, and it leaves out the three unrelated rows that are visibly flat. That is the right report of that table. It says that the modification is in the nested block or on characteristic eight, and that four hundred patients cannot say which.
Over trials the set behaves as a set should. At it holds the real modifier in 95.5% of trials among eight nested characteristics, 97.4% with the real one inside a block of four, and 98.9% among eight unrelated ones — above its 90% because the prior spreads its weight over smaller modifications than the one drawn. Its average size is 4.96 characteristics when all eight are nested, 4.10 for the block of four and 3.86 when nothing overlaps; at it holds 6.63, 6.16 and 6.22. With nothing real it holds 7.22 of the eight nested characteristics, which is the table saying, correctly, that it has nothing to attribute.
That set is wider than a reader of a subgroup table expects, and it is the width the data support. The ceiling figure says the same thing as a hit rate. The set says it in the form a report can carry: the characteristics compatible with the table, rather than a winner that is wrong as often as it is right. It is the same move intervals for the findings made for effect sizes after selection — report the spread of what the data allow, not the point the selection landed on.
What a protocol can take from it
Read the largest row, and know that it is the ceiling. No re-weighting of a subgroup table names the real modifier more often than its largest interaction, provided one characteristic is the modifier and none is favoured beforehand. Effort spent on a cleverer reading of the same table is effort spent on nothing; the overlap’s effect is in the data, not in the analysis.
Do not score candidates against the table as though its rows were independent. That reading is the one a spreadsheet invites, and it moves attributions into whichever block of characteristics overlaps most, at a cost of nine points when the real modifier is outside it.
Do not fit every interaction jointly to decide which one matters. The joint fit answers a different question — which characteristic modifies the effect holding the others fixed — and among nested characteristics it answers it with a fraction of the information. It names the real one less often than the largest row at every structure with overlap.
Report a set. The posterior set costs one line of arithmetic per characteristic and uses nothing but the interaction statistics. A subgroup finding reported as “one of characteristics one to four, or eight” is less quotable than “characteristic eight”, and it is true. Naming a handful in advance remains the one design change that shrinks the set rather than describing it, because a characteristic not tabulated cannot be in it.
The algebra, and what was counted
Exact: under a single modifier with split indicators correlated , the table’s expected shape is the modifier’s column of ; the generalised-least-squares score of candidate is for every ; and with equal prior weight on each characteristic the posterior is increasing in , so the largest row maximises the chance of a correct attribution at every modification size.
Counted, over two thousand trials of four hundred for each structure and size: the largest row names the real modifier in 54.4% of trials among eight nested characteristics at , 63.5% inside a block of four and 82.4% among unrelated ones; the independent-pattern fit in 53.6%, 63.8% and 81.6%, and 74.3% against 83.4% when the real one sits outside a nested block; the joint fit in 45.6%, 50.1% and 82.4%; and the 90% posterior set holds the real one in 95.5% of nested trials with 4.96 members on average.
Not claimed: that a real modifier comes alone. Everything above assumes one characteristic carries the modification, and the equivalence between the fitted shape and the largest row rests on it. Not claimed either that the prior’s scale is right — a different scale changes the set’s size, not which characteristic heads it.
Still open: two modifiers at once
The equivalence holds because the hypothesis is one cause. With two real modifiers — age and a comorbidity, say, each changing the effect — the table’s expected shape is a combination of two columns of , and the largest row is no longer the likelihood’s answer: a pair of moderately raised rows whose overlaps explain the whole table can be better supported than any single row. That is where fitting the shape stops collapsing to the obvious reading, and where the joint fit’s lost information might be bought back by the pairs it can separate.
How often a two-modifier fit names both, how often it invents a second where there is one, and how large the set of plausible pairs is among eight characteristics — twenty-eight pairs rather than eight rows — have not been measured here. They decide whether a subgroup table can ever report more than one characteristic with a straight face.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A subgroup inside its own trial — both name correlation, interaction, subgroup analysis
- A term built from the others — both name interaction, subgroup analysis, variance inflation
- False discoveries that arrive together — both name correlation, familywise error rate, multiple comparisons
- One control, many arms — both name correlation, familywise error rate, multiple comparisons
- Significant in one, not in the other — both name interaction, multiple comparisons, subgroup analysis
- A copula that halves a marginal — both name correlation, interaction
Named objects
A flat tag is an object no other essay names yet.
CorrelationFamilywise error rateGeneralised least squaresInteractionLikelihoodMultiple comparisonsPosteriorPriorSubgroup analysisVariance inflation