Two groups, one treatment, and both readings of the same numbers
The treatment wins in group A (93.0% against 87.0%) and in group B (73.0% against 69.0%), and loses overall (78.1% against 82.6%). Nothing here is a trick; the allocation differs between the groups.
Reversals that are not errorsslider: how the groups are allocated, 0 positionswide18 views
What else it draws
The same object, drawn to answer the other questions the essays put to it.
The per-group success rates are held fixed; only the split of each group between treatment and control changes. 32% of the allocations reverse, and the worst reverses by 13.1 percentage points.
At a prevalence of one in a thousand, 98 of every hundred positives are false. At one in 10, 33 are. The test has not changed.
9 people have it and test positive. 50 do not have it and test positive anyway. So of the 59 positive results, 15% are right — and that is with a test most people would call accurate.
Pick the worst 15% on the first measurement and their average rises by 0.78 on the second. Pick the best and theirs falls by 0.48. No treatment was given to anybody.
600 units in two pre-existing groups whose true means are 1.00 apart, read once at baseline and once at follow-up, with no change for anybody. The two groups' mean changes are −0.075 and −0.032, so the change-score analysis reports a group difference of 0.043. The regression of follow-up on baseline and group reports 0.409, against a closed form of (1 − λ) × 1.00 = 0.400: at any one baseline reading the two groups' lines sit that far apart, because each group's units regress towards their own group's mean. The pooled slope in this sample is 0.614, the baseline's reliability.
2000 people read once at screening and once at follow-up, with a test–retest correlation of 0.6 and no treatment. The 215 above a cut at the top ten per cent of one reading (1.282 standard deviations) are enrolled. Their mean screening reading is 1.744 and their mean follow-up reading 0.982, a fall of 0.762 ± 0.055 with nothing done to anyone. The closed form for the fall is (1 − ρ) times the truncated-normal mean, (1 − 0.6) × 1.755 = 0.702.
The expected second reading divided by the first, for readings from 0.5 to 4 standard deviations, where the true scores come from a normal, a Laplace, a t with four degrees of freedom or a uniform parent, every one with variance 0.6, and every reading adds normal error of variance 0.4. All four put the two readings at a correlation of exactly 0.6. For the normal parent the share kept is 0.6 at every reading. At a reading of 2 the Laplace parent keeps 0.644, the t parent 0.604 and the uniform 0.507; at a reading of 4 they keep 0.817, 0.866 and 0.301. The curves are posterior means by integration, checked against Tweedie's formula to 7e-8 and against closed forms for three of the parents to 9e-10.
Each thin line is one study of 1000 readings from the t, four degrees parent: Tweedie's formula with the log-density's slope estimated by a degree-5 log-spline, drawn up to that study's largest reading. The thick line is the share kept by integration over the true score, and the dashed line the correlation, 0.6. The top ten readings of the median study begin at 2.45. Over 400 studies the corrected share kept by the top one per cent averages 0.8190, with a spread of 0.0806, against 0.7842 by integration; a Gaussian kernel averages 0.7914 with a spread of 0.0853.
Simple randomisation against randomisation stratified by group, 4,000 trials at each size. The simple design reverses on 3.40% of trials at its worst size and 0.50% at 1280 units; the stratified design reverses on none of them, at any size.
Two events on the same trials. The overall difference sitting outside the range of the two group differences happens on about 33% of trials and does not fall with the trial's size. The reversal falls from 3.23% to 0.50%. Only the second is evidence of anything.
Five strata with baseline risks from 5% to 85%, a treatment allocated by a coin in each, and a conditional odds ratio of exactly 2.5 throughout. The marginal odds ratio is 1.789. Nothing is confounded; the odds ratio is simply not a weighted average of odds ratios.
Five strata centred on a 45% baseline risk, spread by the half-width on the axis, with a conditional odds ratio of 2.5 in each. The marginal odds ratio falls from 2.500 to 1.832 — 45% of the way to no effect at all, with no confounding introduced at any point.
The direct effect stays at -0.109 throughout, because it is defined with the mediator held fixed. The total effect runs from -0.218 to 0.027 and crosses zero at 1.5. On 15% of the sweep the two have opposite signs, and both are correct answers.
A 90% sensitive, 95% specific test. One positive gives 1.77%. Two independent positives give 24.49%, which is what multiplying the likelihood ratios says. At a correlation of 0.1 between the tests' errors it is 10.16%, and at 0.5 it is 3.16%.
The published pair — 90% sensitive, 95% specific — is the open mark. With a miss costing 100 times a false alarm and a prevalence of 0.10%, the threshold that minimises expected cost sits at 74.7% sensitivity and 98.81% specificity, with a predictive value of 5.9%.
One test, one pair of costs, and a threshold that runs from 3.05 to -0.12 standard deviations of the score. At one in ten thousand the best test detects 45% of cases; at one in two it detects 99.9%. The published 90/95 is right for one population.
The positive rate reads 5.09%, which is 50.9 times the truth. The Rogan-Gladen correction averages 0.100% — unbiased — with a standard deviation of 0.820 points against the positive rate's 0.695, and it comes out negative on 48.6% of samples.
Where it is used
15 essays draw this figure, each at the numbers its own argument is about, so the same picture answers 15 different questions.
- Simpson's reversal is a region, not a table Reversals that are not errors
- The reversal a coin cannot prevent When the stratified answer and the pooled one disagree
- The second test that is not a second opinion Two tests, a threshold, and the rate they are read against
- What a positive test is worth Reversals that are not errors
- The change that is not confounding When the stratified answer and the pooled one disagree
- The test is a point somebody chose Two tests, a threshold, and the rate they are read against
- Borrowing towards a line Hierarchy past one number
- Regression to the mean Reversals that are not errors
- Conditioning on what the treatment caused When the stratified answer and the pooled one disagree
- The prevalence the test has to estimate Two tests, a threshold, and the rate they are read against
- The base rate was always Bayes The prior, doing visible work
- Two analyses of one baseline Reversals that are not errors
- The measurement that got them enrolled Reversals that are not errors
- A lead that a heavy tail keeps Reversals that are not errors
- The slope of a density nobody can see Reversals that are not errors