Every essay — page 17
The same table at seven correlations
Every cell of the copula-by-marginal table is measured at one rank correlation, and eleven of its twenty cells cancel there. Sweep the correlation from 0.1 to 0.7 and four of the twenty change the sign of their answer, all four from compounding to cancelling, all four at the most skewed covariates. The near-perfect cancellation that is the earlier field's headline is a crossing: the cell passes through zero at a Spearman of 0.38, two hundredths from where it was read, and is two orders of magnitude larger by 0.7.
An answer that changes
Eleven of twenty cells cancel and nine compound, at one rank correlation. Sweep the correlation and four of the twenty change sides — all four from compounding to cancelling, all four at the most skewed covariates.
The other dial
The table is swept along the strength of the dependence and never along the shape of the covariate. Swept along the shape at a fixed correlation, the same two copulas cross, the same way — and the near-zero cell turns out to be a minimum in both directions at once.
The interval, studentised
A percentile interval inherits the resampled distribution's skewness and its scale error together, and the standard repair is to resample a t-statistic so each resample carries its own scale. It is one extra variance per resample, and it repairs nothing: not one of the eight cells reaches its promise, the interval is twice as wide for half a point of coverage, and at the block lengths the rules choose the scale rests on two or three whole blocks. And the ordering between the two windows reverses at all four rules, on resamples that are the same resamples.
An interval that carries its scale
A percentile interval inherits the resampled distribution's skewness and its scale error together. The standard repair is one extra variance per resample. It was named and not run, so this runs it.
What studentising costs
Averaged over eight cells the studentised interval is 2.09 times as wide as the percentile one and covers 0.46 points better. At the block lengths the rules choose, the scale it divides by rests on two or three numbers.
The ordering reverses again
One field found two of four rules changing sign between two readings of one resampling. Turn the same resamples into a studentised interval instead of a percentile one and all four change sign.
The count or the length
A block length and a block count are one number read two ways at one sample size. Read at three, the studentised interval's width penalty tracks the count — with an R² of 0.9911 against a closed form that has no length in it — and its coverage tracks the length.
The interval with no resampling in it
Replace 1.96 in a normal interval on the block-means variance with Student's t on one fewer degrees of freedom than there are whole blocks, and resample nothing. Across twenty-four cells it covers at least as often as the studentised bootstrap interval at every one, by 0.42 to 10.42 points; it is narrower wherever seven blocks or fewer are left; and at fifteen blocks of 32 it covers 95.0%, which no resampled interval on the grid reaches.
The interval that holds observations, not a mean
Every interval counted before this one is about a parameter. These are about the values themselves — what a band actually holds, and how often. Two sample standard deviations around a sample mean of ten observations holds 91.1% of the population on average, and less than the advertised 95.45% on 59.9% of samples, so the average is the reading that hides it. The interval between a sample's smallest and largest reading holds a share whose distribution is Beta(n − 1, 2) for any continuous population at all, and buying with no assumption what normality buys at ten observations costs ninety-three of them. That number is the exchange rate between an assumption and data, priced in observations. And a band for all of the next m observations has no ceiling at the tolerance factor: from ten observations, holding all of the next ten needs 3.716 standard deviations, already past the 3.382 that holds 95% of the population, while a 95% prediction interval holds all ten only 67.9% of the time.
Two standard deviations of what
The 95.45% inside two standard deviations is a fact about a curve whose centre and width are given. Drawn from ten observations, the same band holds 91.1% on average and less than 95% on 59.9% of samples — and the average is the reading that hides it.
A tenth as wide, and both of them right
The interval for a mean and the interval for one future observation are both labelled 95%, and at a hundred observations one is 10.05 times the other — exactly the square root of n + 1. Read the narrow one as the wide one and it covers a new value 15.7% of the time.
Ninety-three observations, and nothing assumed
The interval between the smallest and largest of a sample holds a share of the population whose distribution does not depend on the population — Beta(n − 1, 2), for anything continuous. Buying the 95/95 that normality buys at ten observations costs 93 of them, and that number is the exchange rate between an assumption and data.
All of the next ten
A warranty, a batch release or a monitoring rule promises something about every one of the next ten observations, not about one. From a sample of ten, the band that holds all ten with 95% probability reaches 3.716 sample standard deviations either side of the mean — already wider than the 3.382 of a tolerance interval for 95% of the population — and it keeps widening: 4.942 for a hundred, 6.008 for a thousand, with no ceiling. A 95% prediction interval, read as the answer, holds all ten 67.9% of the time: more than 0.95 to the tenth power, because the ten succeed and fail together.
When the stratified answer and the pooled one disagree
Simpson's reversal is normally taught once, on one table, as a warning about confounding. None of the reversals here is confounded. A trial randomised by a coin, on a population where the treatment helps in both groups, reports it losing overall on 3.40% of trials at eighty units — and stratifying the randomisation takes that to zero at every size. An odds ratio of exactly 2.5 in all five strata pools to 1.789, because an odds ratio is not a weighted average of odds ratios, while the risk difference on the same table is exactly its stratum value. And where the grouping variable is something the treatment caused, both answers are right and answer different questions.
The reversal a coin cannot prevent
Randomisation removes Simpson's reversal in expectation, which is not the same as removing it. A correctly randomised trial of eighty units, on a population where the treatment helps in both groups, reports it losing overall on 3.40% of trials — and stratifying the randomisation takes that to zero at every size.
The change that is not confounding
Five strata, a treatment allocated by a coin in every one, and an odds ratio of exactly 2.5 in all five. The odds ratio computed on the pooled table is 1.789. Nothing is confounded — an odds ratio is not a weighted average of odds ratios, and the risk difference, on the same table, is exactly its own stratum value.
Conditioning on what the treatment caused
When the grouping variable lies on the path from treatment to outcome, the stratified answer is the direct effect and the aggregate is the total effect. Both are correct. Over 15% of a sweep of the indirect path they have opposite signs, and no arithmetic on the table says which question was being asked.
The analyses that were available and not run
A correction divides by the number of analyses, and that number is the one quantity nobody measures. Twenty analyses of one dataset are worth 11.37 independent ones at a correlation of 0.6 and 2.58 at 0.95, so the threshold controlling a stated error rate can be counted rather than assumed. What the correction costs is charged in the estimate rather than the error rate: a larger statistic is a more selected one, and what survives averages 1.69 times the truth after correction against 1.35 times before it. Naming the analysis in advance is the alternative, and it is worth exactly what correcting all twenty is worth when the chance of having named the right one is 38%.
How many analyses there really were
Bonferroni divides by twenty because twenty analyses were run. Twenty analyses of one dataset are worth 11.37 independent ones at a correlation of 0.6 and 2.58 at 0.95, and the threshold that controls exactly the same error rate is measurable rather than assumed.
What naming it in advance costs
Preregistration is argued for as free. Against an effect of two standard errors hiding in one of twenty analyses, naming the right one detects it 51.5% of the time and naming the wrong one detects it 4.7% of the time; correcting all twenty detects it 22.5% wherever it is. The two are worth the same when the chance of having named correctly is 38%.
The correction that makes the estimate worse
Correcting for twenty analyses repairs the p-value by demanding a larger statistic, and a larger statistic is a more selected one. At two standard errors the surviving estimate averages 1.35 times the truth before the correction and 1.69 times it after — so the honest error rate is bought with a more inflated effect.
What a summary of a scatter is a property of
R-squared is read as a property of a relationship and it is a property of a design. Five studies differing only in how far apart they placed their readings report it from 0.021 to 0.849, with the same line and an estimated residual spread of 0.993 in every one. For a simple regression the squared t statistic is (n − 2) times R-squared over (1 − R-squared), exactly, checked to sixteen significant figures across five hundred fits — so a report carrying both has carried one number twice. A coefficient that is zero only under independence does exist: distance correlation separates the four datasets built to share a correlation by 0.10, and the rank coefficient that guarantees nothing separates them by 0.49.
R² is a property of the design
One line, one residual spread, five studies that differ only in how far apart they placed their x values. R² runs from 0.021 to 0.849 and the estimated residual spread is 0.993 in every one of them. Nothing about the relationship changed.
The t statistic wearing different clothes
For a simple regression, t² = (n − 2)R²/(1 − R²), exactly, on every dataset — checked to sixteen significant figures over five hundred fits. So a paper reporting R² and a p-value has reported one number twice, and two studies with the same R² have points four times further from the line.
The summary that was meant to work
Distance correlation is zero if and only if two variables are independent, which is exactly the guarantee a correlation coefficient lacks. Run on the four datasets that share a correlation, it spreads them by 0.10 — and Spearman, which guarantees nothing, spreads them by 0.49.
Shape, and what it does to a two-sample test
A t interval survives a skewed source in total and not in its tails. On an exponential source at 120 observations a 95% interval covers 94.81% while missing below the mean on 4.08% of samples and above it on 1.11%. The pooled two-sample test's size runs from 0.55% to 18.91% across forty units with a true null in every cell, and Welch's stays between 4.63% and 5.51% — on normal populations. On skewed ones Welch's tails are decided by one number, the skewness of the difference of the two means: two identical exponentials balance at twenty and twenty and reject low on 7.16% of samples at eight and thirty-two. And the side a decision reads needs its own repair: widened until its total coverage is exactly 95%, a t interval's upper limit on thirty exponential observations is still exceeded on 4.69% of samples, where Hall's transformation takes it to 3.31%.
Where the two tails disagree
A 95% t interval on an exponential source at 120 observations covers 94.81%, which reads as very nearly right. It misses below the mean on 4.08% of samples and above on 1.11% — one tail 63% too heavy and the other 56% too light, and the total is the statistic that hides it.
A degrees of freedom that is not a count
The pooled two-sample test's size runs from 0.55% to 18.91% across forty units split five ways against five variance ratios, with a true null in every cell. Welch's runs from 4.63% to 5.51% — bought with a degrees of freedom that is a function of the data, not an integer, and not a count of anything.
The skewness of a difference
Welch's test holds its size to within half a point when both groups are normal. Give both groups the same skewed population and it still balances at twenty and twenty — and at eight and thirty-two it rejects low on 7.16% of samples and high on 0.66%. One number decides which: the skewness of the difference of the two means, which ranks twenty-five cells by their imbalance with a correlation of 0.997.
The side a bound is read from
On thirty exponential observations the upper limit of a 95% t interval is exceeded by the true mean 6.38% of the time, against the 2.5% a safety margin set from it assumes. Widen the interval until its total coverage is exactly 95% and the upper limit is still exceeded 4.69% of the time. A symmetric repair fixes the number that is reported and not the one that is used; Hall's transformation, which bends the interval, takes the same rate to 3.31%.