The collection

Every essay — page 17

Essays 385 to 408 of 436, in the same order.

The same table at seven correlations

Every cell of the copula-by-marginal table is measured at one rank correlation, and eleven of its twenty cells cancel there. Sweep the correlation from 0.1 to 0.7 and four of the twenty change the sign of their answer, all four from compounding to cancelling, all four at the most skewed covariates. The near-perfect cancellation that is the earlier field's headline is a crossing: the cell passes through zero at a Spearman of 0.38, two hundredths from where it was read, and is two orders of magnitude larger by 0.7.

The interval, studentised

A percentile interval inherits the resampled distribution's skewness and its scale error together, and the standard repair is to resample a t-statistic so each resample carries its own scale. It is one extra variance per resample, and it repairs nothing: not one of the eight cells reaches its promise, the interval is twice as wide for half a point of coverage, and at the block lengths the rules choose the scale rests on two or three whole blocks. And the ordering between the two windows reverses at all four rules, on resamples that are the same resamples.

Three intervals, one shortfall. What each of three intervals actually covers, at four rules and two block windows, over 300 samples of 120 rows. All three are built from the same resamples on the same draws, so a difference between them is a difference in what is done with the resampled series. Not one of the twenty-four cells reaches the ninety-five per cent it promises. The studentised interval runs from 75.7% to 92.3%, the percentile interval — the earlier field's — from 80.0% to 89.7%, and a normal interval on the same scale from 81.7% to 89.0%. The standard repair for a percentile interval's shortfall does not repair it.

An interval that carries its scale

A percentile interval inherits the resampled distribution's skewness and its scale error together. The standard repair is one extra variance per resample. It was named and not run, so this runs it.

4 figures · Bootstrap, part 20
What studentising costs. How much wider the studentised interval is than the percentile one, cell by cell, over 300 draws, with what each cell gains in coverage beside it. Averaged over the eight cells the interval is 2.09 times as wide and covers 0.46 points better. At the two rules that choose short blocks the two intervals are within a fifth of each other; at the oracle's length, where a resample holds two or three whole blocks, the studentised interval is 4.37 and 5.04 times as wide. A repair that doubles the width and buys half a point is not one a reader could not have had by widening the interval it replaced.

What studentising costs

Averaged over eight cells the studentised interval is 2.09 times as wide as the percentile one and covers 0.46 points better. At the block lengths the rules choose, the scale it divides by rests on two or three numbers.

4 figures · Bootstrap, part 21
Four rules of four change sign. The margin between the two block windows in points of coverage, under each of four rules, on each of three intervals built from the same resamples, over 300 draws. Positive is the tapered window covering better. On the percentile interval the taper wins at all four rules, by 5.33, 1.67, 5.00 and 4.00 points, which is the earlier field's own reading. On the studentised interval the rectangle wins at all four, by 4.33, 7.00, 4.33 and 2.67. And a normal interval, which uses no resampling at all, puts the two within a third of a point at every rule — so the disagreement is manufactured entirely by what is done with the resamples.

The ordering reverses again

One field found two of four rules changing sign between two readings of one resampling. Turn the same resamples into a studentised interval instead of a percentile one and all four change sign.

4 figures · Bootstrap, part 22
One penalty, read along two dials. How much wider the studentised interval is than the percentile one, at every sample size and every block length on the grid, with the number of whole blocks each cell leaves written beneath. Read across a row and the block length changes; read down a column and the sample size does. The penalty is nearly a function of the block count alone: the cells at 15 blocks read 1.16, 1.20, 1.17, 1.15, 1.13, 1.10 across three sample sizes and three block lengths, while the cells at one block length read anything from 1.10 to 3.95. The largest penalty on the grid is 3.95, at the cell with 3 whole blocks in it.

The count or the length

A block length and a block count are one number read two ways at one sample size. Read at three, the studentised interval's width penalty tracks the count — with an R² of 0.9911 against a closed form that has no length in it — and its coverage tracks the length.

7 figures · Bootstrap, part 23
Four intervals at 3 blocks of 32 rows. What four 95% intervals for the mean of a first-order autoregression at 0.7 cover, and how wide they are on average, at 120 rows cut into 3 whole blocks of 32, over 240 draws with 200 resamples each under the rectangle. The normal interval, the block-means variance with 1.96, covers 82.1% at a width of 0.708. The percentile interval covers 82.1% at 0.632 and the studentised one 90.4% at 2.495. The fourth resamples nothing: it is the normal interval with 1.96 replaced by Student's t on 2 degrees of freedom, and it covers 94.2% at 1.555, 0.62 times the studentised interval's width.

The interval with no resampling in it

Replace 1.96 in a normal interval on the block-means variance with Student's t on one fewer degrees of freedom than there are whole blocks, and resample nothing. Across twenty-four cells it covers at least as often as the studentised bootstrap interval at every one, by 0.42 to 10.42 points; it is narrower wherever seven blocks or fewer are left; and at fifteen blocks of 32 it covers 95.0%, which no resampled interval on the grid reaches.

6 figures · Bootstrap, part 24

The interval that holds observations, not a mean

Every interval counted before this one is about a parameter. These are about the values themselves — what a band actually holds, and how often. Two sample standard deviations around a sample mean of ten observations holds 91.1% of the population on average, and less than the advertised 95.45% on 59.9% of samples, so the average is the reading that hides it. The interval between a sample's smallest and largest reading holds a share whose distribution is Beta(n − 1, 2) for any continuous population at all, and buying with no assumption what normality buys at ten observations costs ninety-three of them. That number is the exchange rate between an assumption and data, priced in observations. And a band for all of the next m observations has no ceiling at the tolerance factor: from ten observations, holding all of the next ten needs 3.716 standard deviations, already past the 3.382 that holds 95% of the population, while a 95% prediction interval holds all ten only 67.9% of the time.

What x-bar plus or minus 2 sample standard deviations holds, at n = 10. The content of the band is a random variable. Across 20,000 normal samples of 10 it averages 91.1%, its fifth percentile is 74.7%, and it falls short of 95% on 59.9% of samples. The band that is drawn to show where 95% of the data lies.

Two standard deviations of what

The 95.45% inside two standard deviations is a fact about a curve whose centre and width are given. Drawn from ten observations, the same band holds 91.1% on average and less than 95% on 59.9% of samples — and the average is the reading that hides it.

7 figures · Bands, part 2
Three bands called 95%, at n = 20. Half-widths in sample standard deviations: 0.468 for the mean, 2.14 for one future observation, 2.75 to hold 95% of the population. The first two differ by exactly the square root of n + 1, which is 4.58 here.

A tenth as wide, and both of them right

The interval for a mean and the interval for one future observation are both labelled 95%, and at a hundred observations one is 10.05 times the other — exactly the square root of n + 1. Read the narrow one as the wide one and it covers a new value 15.7% of the time.

6 figures · Bands, part 3
How many observations the smallest and largest of them need. The interval between the extremes of n draws holds at least 95% of the population with probability 1 - n p^(n-1) + (n-1) p^n, whatever the population is. Reaching 95% confidence takes 93 observations.

Ninety-three observations, and nothing assumed

The interval between the smallest and largest of a sample holds a share of the population whose distribution does not depend on the population — Beta(n − 1, 2), for anything continuous. Buying the 95/95 that normality buys at ten observations costs 93 of them, and that number is the exchange rate between an assumption and data.

6 figures · Bands, part 4
The factor that holds all of the next m observations with 95% probability, from a sample of 10. From 10 observations, the band for one future value has factor 2.371, for ten 3.716, for a hundred 4.942 and for a thousand 6.008. The 95%-content tolerance factor is 3.382 and is passed by m = 10; the Bonferroni stretch of the prediction factor reaches 7.567 at a thousand.

All of the next ten

A warranty, a batch release or a monitoring rule promises something about every one of the next ten observations, not about one. From a sample of ten, the band that holds all ten with 95% probability reaches 3.716 sample standard deviations either side of the mean — already wider than the 3.382 of a tolerance interval for 95% of the population — and it keeps widening: 4.942 for a hundred, 6.008 for a thousand, with no ceiling. A 95% prediction interval, read as the answer, holds all ten 67.9% of the time: more than 0.95 to the tenth power, because the ten succeed and fail together.

5 figures · Bands, part 5

When the stratified answer and the pooled one disagree

Simpson's reversal is normally taught once, on one table, as a warning about confounding. None of the reversals here is confounded. A trial randomised by a coin, on a population where the treatment helps in both groups, reports it losing overall on 3.40% of trials at eighty units — and stratifying the randomisation takes that to zero at every size. An odds ratio of exactly 2.5 in all five strata pools to 1.789, because an odds ratio is not a weighted average of odds ratios, while the risk difference on the same table is exactly its stratum value. And where the grouping variable is something the treatment caused, both answers are right and answer different questions.

The analyses that were available and not run

A correction divides by the number of analyses, and that number is the one quantity nobody measures. Twenty analyses of one dataset are worth 11.37 independent ones at a correlation of 0.6 and 2.58 at 0.95, so the threshold controlling a stated error rate can be counted rather than assumed. What the correction costs is charged in the estimate rather than the error rate: a larger statistic is a more selected one, and what survives averages 1.69 times the truth after correction against 1.35 times before it. Naming the analysis in advance is the alternative, and it is worth exactly what correcting all twenty is worth when the chance of having named the right one is 38%.

What a summary of a scatter is a property of

R-squared is read as a property of a relationship and it is a property of a design. Five studies differing only in how far apart they placed their readings report it from 0.021 to 0.849, with the same line and an estimated residual spread of 0.993 in every one. For a simple regression the squared t statistic is (n − 2) times R-squared over (1 − R-squared), exactly, checked to sixteen significant figures across five hundred fits — so a report carrying both has carried one number twice. A coefficient that is zero only under independence does exist: distance correlation separates the four datasets built to share a correlation by 0.10, and the rank coefficient that guarantees nothing separates them by 0.49.

Shape, and what it does to a two-sample test

A t interval survives a skewed source in total and not in its tails. On an exponential source at 120 observations a 95% interval covers 94.81% while missing below the mean on 4.08% of samples and above it on 1.11%. The pooled two-sample test's size runs from 0.55% to 18.91% across forty units with a true null in every cell, and Welch's stays between 4.63% and 5.51% — on normal populations. On skewed ones Welch's tails are decided by one number, the skewness of the difference of the two means: two identical exponentials balance at twenty and twenty and reject low on 7.16% of samples at eight and thirty-two. And the side a decision reads needs its own repair: widened until its total coverage is exactly 95%, a t interval's upper limit on thirty exponential observations is still exceeded on 4.69% of samples, where Hall's transformation takes it to 3.31%.

Which side a 95% t interval misses on, exponential source. Both tails should be 2.5%. At 8 observations the interval falls short of the mean on 9.75% of samples and overshoots on 0.31%. At 500 they are 3.31% and 2.05%, and the total is 5.36% — which a coverage table reports as very nearly right.

Where the two tails disagree

A 95% t interval on an exponential source at 120 observations covers 94.81%, which reads as very nearly right. It misses below the mean on 4.08% of samples and above on 1.11% — one tail 63% too heavy and the other 56% too light, and the total is the statistic that hides it.

5 figures · Student, part 3
The pooled two-sample test's size, with a true null everywhere. Forty units split between two groups, with the second group's variance a stated multiple of the first's, and the two population means equal. A 5% test should reject 5% of the time. The pooled test runs from 0.55% to 18.91% across this region; Welch's runs from 4.63% to 5.51%.

A degrees of freedom that is not a count

The pooled two-sample test's size runs from 0.55% to 18.91% across forty units split five ways against five variance ratios, with a true null in every cell. Welch's runs from 4.63% to 5.51% — bought with a degrees of freedom that is a function of the data, not an integer, and not a count of anything.

5 figures · Student, part 4
Welch's test on skewed groups: the low and high rejection rates in every cell, with equal means throughout. Each cell should read 2.5 / 2.5. Two identical exponentials at 20 and 20 read 2.04 / 2.22; the worst cell, a wide exponential against a normal at 8 and 32, reads 9.79 / 0.47.

The skewness of a difference

Welch's test holds its size to within half a point when both groups are normal. Give both groups the same skewed population and it still balances at twenty and twenty — and at eight and thirty-two it rejects low on 7.16% of samples and high on 0.66%. One number decides which: the skewness of the difference of the two means, which ranks twenty-five cells by their imbalance with a correlation of 0.997.

5 figures · Student, part 5
Four 95% intervals for the mean of 30 exponential observations, split by the side they miss on. Both bars should read 2.5%. The t interval misses below the mean on 6.38% of samples. Widened until its total is exactly 5%, it misses below on 4.69% and above on 0.30%. Hall's transformation misses below on 3.31% and above on 1.96%.

The side a bound is read from

On thirty exponential observations the upper limit of a 95% t interval is exceeded by the true mean 6.38% of the time, against the 2.5% a safety margin set from it assumes. Widen the interval until its total coverage is exactly 95% and the upper limit is still exceeded 4.69% of the time. A symmetric repair fixes the number that is reported and not the one that is used; Hall's transformation, which bends the interval, takes the same rate to 3.31%.

6 figures · Student, part 6

FieldsThreadsSeriesConceptsFigure librarySearch