What a probe can see, and what a chain says about it
Each point is one probe on one fourteen-unit admissible set: horizontally what the enumeration says the two components differ by on that probe, in within-component spreads; vertically what a pair of chains 20,000 steps long reports. The two agree, which is the check — a chain that reports a large statistic where the enumeration says the components are nearly coincident would be reporting its own failure to mix. The bars through the points are the range over four chain seeds. What the picture is for is the horizontal axis: the covariate probes sit at 0.431 and 0.419, and the outcomes run from 0.001 to 4.938. A probe is chosen; an outcome is what happened.
The diagnostic after the trialwide3 views
What else it draws
The same object, drawn to answer the other questions the essays put to it.
How wrong three p-values are when they are computed over the half of the admissible set a single walk can reach, rather than over all of it, at a fourteen-unit trial where the whole set can be enumerated. The two-sided p-value on the difference in arm means — the number a trial publishes — is wrong by exactly nothing, at every row, to machine precision. That is not luck: the two components are complement pairs and the difference in arm means is exactly negated by the complement, so the distribution of its absolute value is the same on both. A one-sided p-value on the same statistic is out by as much as 0.112, and the largest response observed in the treated arm — a safety reading rather than an effect, and the one statistic here that is not odd under the complement — by as much as 0.172. The defect survived because the commonest thing anybody computes is the one quantity it cannot touch.
The two-chain statistic on a covariate probe and on the trial's own difference in arm means, at seven tolerances of a fourteen-unit rule, against the enumerated truth. Both are quiet wherever the set is one set and both fire wherever it is not, at every tolerance — which is what says the outcome probe is the same test rather than a resemblance of it. The difference between them is not accuracy and it is not power. It is *when*: the covariate probe can be run before a single outcome exists, when a practitioner can still loosen the rule or change the sampler, and it can be run again on a different function if it comes back quiet. The outcome probe runs after the trial, on the one column the trial produced, and what it can do with a positive verdict is repair the p-value rather than the design.
Where it is used
4 essays draw this figure, each at the numbers its own argument is about, so the same picture answers 4 different questions.
- The statistic the p-value is about The diagnostic after the trial
- A probe nobody chose The diagnostic after the trial
- Half a reference distribution The diagnostic after the trial
- Before the trial and after The diagnostic after the trial