What makes it checkable

A proposal fitted to its own draws

The cross-entropy method fits an importance-sampling proposal from its own pilot draws, and asked to fit a normal's mean and spread to P(Z > 5) it fails before the question of variance arises: the spread halves at every stage, the level stalls near 3.7, and 95% of runs never reach the threshold, so the interval covers 1.1% of the time. Its ideal end point, the normal closest to Z conditioned past 5, is N(5.187, 0.181²), which has an infinite variance inside its own draws' reach and covers 93.7% at a thousand draws and 90.8% at ten thousand. Fitting the mean alone lands within a tenth of the best shift and covers 94.5%.

Worth reading first: The draws aimed at the tail · What the 95% refers to.

The draws aimed at the tail estimated P(Z > 5) = 2.8665 × 10⁻⁷ by drawing from a normal proposal placed near the tail and weighting each draw back. A proposal aimed at the threshold with unit spread covered its 95% interval about 94.7% of the time; one only slightly too narrow had an infinite variance and covered 86% however many draws it was given; and the standard diagnostic could not tell the two apart. Every proposal in that essay was chosen before any draw was made.

In practice the proposal is usually chosen by the draws. The cross-entropy method, the most common recipe, draws a batch from the current proposal, keeps the draws that landed furthest out, refits the proposal to them, and repeats until the draws reach the threshold. It is the pattern of a simulation that stops when it looks settled in a new place: a procedure that uses the data to decide how it will analyse the data. The earlier essay asked whether such a procedure could fit its way into the narrow, infinite-variance region it had shown to be dangerous. It can, and the version most often written down does something worse first.

How the cross-entropy method fits a normal

Start from the standard normal. Draw a thousand points, find the level that the top tenth of them pass, and refit the normal’s mean and spread to those top draws, each weighted by φ(x)/q(x)\varphi(x)/q(x), the ratio of the target density to the proposal it came from. Those weighted moments are the normal closest, in Kullback–Leibler distance, to the standard normal conditioned past the level. Move the proposal to them, draw again, and the next level is higher. When the top tenth pass 5, fit once more to every draw past 5 and stop. The final proposal is then used for fresh draws, independent of the pilot, and the estimate and its interval come from those.

The method’s end point can be computed without running it. The normal closest to Z conditioned on Z > 5 has the conditional mean and variance: mean φ(5)/Φˉ(5)=\varphi(5)/\bar\Phi(5) = 5.187, and spread 0.181. That is the proposal the cross-entropy method is aiming at — and its spread is far below 0.707, the boundary under which a normal proposal’s second moment is infinite.

The spread collapses before the level arrives

A cross-entropy fit of a normal proposal to P(Z > 5), stage by stage, fitting mean and spread or the mean alone. Each stage draws a thousand points from the current proposal, keeps the top tenth, and refits the normal to them, each weighted by the target density over the proposal's. Fitting the mean alone, the levels are 1.24, 2.99, 4.59, 5.00 and the fit reaches 5 at stage 4, ending at N(5.199, 1). Fitting mean and spread from the same seed, the spread goes 1.000, 0.456, 0.244, 0.096, 0.054, 0.043 over the first six stages, roughly halving each time, and after 30 stages the level has stalled at 3.728 with a spread of 1.0e-6, never reaching 5.
Fig. 1 One cross-entropy fit from one seed, stage by stage. Fitting the mean alone, the level reaches 5 at the fourth stage. Fitting mean and spread, the spread roughly halves at every stage, and with it the distance each stage can climb, until the level stalls at about 3.7 with a spread too small to measure.

The figure follows one run. Fitting mean and spread, the spread goes 1.000, 0.456, 0.244, 0.096, 0.054 and 0.043 over the first six stages, and the level climbs from 1.24 to 3.34 in five stages and by less at every stage after. After thirty stages it sits at 3.728 and the proposal’s spread is a millionth.

The mechanism is in what the refit is fitted to. Each stage keeps the top tenth of its draws, which are a narrow slice of the current proposal’s upper tail, and the spread of that slice is a fraction of the proposal’s own spread. Refitting to it shrinks the spread by about the same fraction every stage. But the next level is set by how far the new proposal’s top tenth reaches, which is its mean plus about 1.28 of its spreads, so as the spread shrinks geometrically the climb per stage does too. The sum of a geometric series is finite, and here it runs out more than a full unit short of 5.

The first stage can be computed exactly, because its weights are all one. The top tenth of a standard normal lies above 1.282, its mean is 1.755 and its spread is 0.411 of the parent’s — the run’s 0.456 is that number seen through a thousand draws. If every stage shrank the spread by that factor and climbed by 1.755 of its own spread, the whole climb would be 1.755/(1−0.411)=1.755/(1 - 0.411) = 2.98 units from the starting point, and the method would stall near 3. The later stages shrink a little less than the first, because their weights favour the lower part of each slice and so widen it, which is why the stall is nearer 3.7 than 3; but no weighting changes the shape of the arithmetic, and every fit of mean and spread from this start has a ceiling below the threshold unless its first few batches happen to be unusually spread. The method converges, confidently, to a proposal that never draws past the threshold.

Counted over two thousand runs, the full fit reaches the threshold on 5.5% of them; the median run stalls at a mean of 3.970. A proposal that never draws past 5 estimates the probability as zero with a standard error of zero, and its interval is the single point zero — the same blindness as plain simulation, now arrived at by an adaptive procedure that ran for thirty stages. The interval covers 1.1% of the time at a thousand final draws and 0.8% at ten thousand — the interval that covers nothing at a count of zero, reached the long way round.

Where the fits stop, and why smoothing does not help

One run is an anecdote, and the collapse is a property of the method only if it happens on most runs.

Where two thousand cross-entropy fits of mean and spread stop. The level the last stage of each of 2,000 independent fits reached, in bins a tenth wide. A tenth stop below 3.45, half below 3.97 and nine tenths below 4.77; 109 reach the threshold of 5.
Fig. 2 The level the last stage reached on each of two thousand independent fits of mean and spread. Most stop between 3.5 and 4.5; the bar at 5 is the 109 that arrived.

It happens on most runs, and where each one stops is set by its first few stages. A tenth of the two thousand fits stall below 3.45, half below 3.97 and nine tenths below 4.77; 109 reach 5. A fit whose early stages happened to draw a top tenth spread more widely than usual shrinks its spread more slowly and climbs further before it runs out of room, and the few that arrive are the ones whose early batches were lucky in that way. The threshold is reached by luck in the pilot draws rather than by the method’s design, which is the reverse of what an adaptive procedure is for.

The textbook guard against a collapsing update is smoothing: move each parameter only part of the way to its newly fitted value, here seven tenths, so a single stage cannot shrink the spread all at once. It slows the collapse and does not stop it. The smoothed fit reaches 5 on 6.0% of runs, stalls at a median mean of 4.185 instead of 3.970, and its interval covers 1.4% of the time at a thousand final draws. Smoothing delays a geometric shrinkage by a constant factor at each stage; the shrinkage is still geometric, and the climb it permits is still a convergent sum.

Fitting the mean alone

The same fit with the spread held at one does what the method is supposed to do.

From the same seed it climbs 1.24, 2.99, 4.59 and reaches 5 at the fourth stage, ending at a mean of 5.199. Across runs the median fitted mean is 5.187 — the conditional mean the full fit was aiming at, reached because a unit spread keeps every stage’s top tenth well spread out. The best mean for a unit-spread proposal, found from the closed form for its variance, is 5.097, and a unit-spread proposal at 5.187 has a relative error per draw of 2.382 against the best one’s 2.376. The method has no knowledge of that closed form and lands within a quarter of a per cent of its optimum.

Its interval covers 94.5% at a thousand final draws and 94.2% at ten thousand, against 94.7% for the fixed proposal aimed at 5. Adapting the mean costs nothing measurable in coverage, and it finds a proposal a person would have had to work out.

It does cost draws. Four stages of a thousand pilot draws precede a final sample of a thousand, so the adapted estimate spends four fifths of its budget finding the proposal. At ten thousand final draws the pilot is under a third of the total, and the trade improves with the size of the final run — which is the regime in which a fixed proposal chosen from the closed form would also have been cheap to find, so adaptation earns its keep on problems without a closed form, not on this one.

Where the full fit was heading

The full fit does not usually reach its end point, but some implementations would — with more pilot draws, a smoothed update, or a stage that refits only when the level is close. The question the earlier essay asked is whether that end point is safe. It can be tested directly by fixing the proposal at N(5.187, 0.181²) and counting.

How often the nominal 95% interval of P(Z > 5) covers, when the proposal was fitted from the same run's pilot draws. Counted over 2,000 runs at a thousand final draws and 600 at ten thousand. Mean and spread fitted: 1.1% and 0.8%. Mean and spread fitted, each update smoothed at 0.7: 1.4% and 1.0%. Where the full fit is heading, N(5.186, 0.181²): 93.7% and 90.8%. Spread fitted, floored at 0.71: 93.8% and 94.2%. Mean fitted, spread held at one: 94.5% and 94.2%. The full fit, a tenth of draws from N(5, 1): 92.6% and 93.8%.
Fig. 3 How often the nominal 95% interval covers, for five versions of the adaptation, at a thousand and ten thousand final draws: the full fit, its ideal end point fixed in advance, the spread floored just above the finite-variance boundary, the mean fitted alone, and the full fit with a tenth of the draws from N(5, 1).

It covers 93.7% at a thousand draws and 90.8% at ten thousand, and of its misses 89% and 95% are underestimates. That is the narrow proposal’s pattern from the earlier essay, milder: coverage that falls as the run grows, misses on one side, a typical estimate a little low — 0.994 and 0.996 of the truth at the median — because the weights that would widen the interval are the rare ones a run usually lacks. It is the same limit as a resample that cannot contain a value beyond its largest observation: a run’s own variance estimate is built only from the weights the run drew, and the weights that would correct it are, by construction, the ones it has not drawn yet. The central limit theorem that a finite variance would supply arrives late or not at all, as it does for a sum of draws with no variance.

The reason is where the divergence starts. A normal proposal with spread under 0.707 has a weight integrand that turns upward at m/(1−2s2)m/(1 - 2s^2), here 5.549. The typical largest of a thousand draws from the end point is 5.745, and of ten thousand 5.859, so every run reaches into the region where the second moment is still growing, and larger runs reach further into it.

The error a run can see, for the cross-entropy fit's ideal end point and for the mean-only fit, by the number of draws. The second moment of the weights integrated from 5 to the typical largest of R draws, as relative error per draw. N(5.187, 0.181²), the normal closest to Z conditioned past 5: 0.77, 0.78, 0.80, 0.82, 0.85, 0.89, 0.95, 1.18. N(5.187, 1): 2.38, 2.38, 2.38, 2.38, 2.38, 2.38, 2.38, 2.38 — at 100, 300, 1,000, 3,000, 10,000, 30,000, 100,000, 1,000,000 draws. The first has an infinite second moment whose divergence starts at 5.549, inside the reach of a thousand draws (5.745); the second's is finite.
Fig. 4 The relative error per draw a run of each size can see — the second moment of the weights integrated up to the typical largest draw — for the cross-entropy end point and for the mean-only fit. The end point looks three times as efficient at practical sizes, and its visible error keeps rising; the mean-only fit’s is flat.

The figure shows why a run cannot see the problem, and also why the method was drawn there. The end point’s visible relative error per draw is 0.80 at a thousand draws, against 2.38 for the mean-only fit: judged by the variance its own draws can see, it is three times as efficient, which is what minimising the Kullback–Leibler distance to the conditional distribution promises. But that number climbs — 0.85 at ten thousand, 0.95 at a hundred thousand and 1.18 at a million — while the mean-only fit’s stays at 2.38. The end point’s efficiency is borrowed from a part of the weight distribution that a run of any given size has not yet drawn, and an interval built from the part it has drawn is too narrow.

The usual diagnostic agrees with the variance the draws can see, and therefore disagrees with the coverage. The effective sample size of the weights, as a share of the draws that landed past 5, is 68.3% for the end point at a thousand draws and 26.2% for the mean-only fit. An analyst who compared the two by that number would choose the end point by a factor of two and a half — and choose the proposal whose interval covers 93.7% at a thousand draws and 90.8% at ten thousand over the one that covers 94.5% and 94.2%. That is the earlier essay’s finding about the diagnostic that reads healthy, now with an algorithm that optimises for the same thing the diagnostic rewards.

So the answer to the earlier essay’s question is yes, in both of the ways it could be. The full fit, run to completion, would land in the narrow, infinite-variance region, because the distance it minimises rewards narrowness. And the cheapest diagnostic would recommend it, because by the variance the draws can see it is the best proposal in the set.

Three ways to keep the adaptation safe

Fit the mean only. Hold the spread at one and let the draws place the centre. It is the simplest repair, it is the one measured above, and on this problem it is nearly optimal: 94.5% and 94.2%.

Floor the spread. Let the spread adapt but never fall below 0.71, just above the finite-variance boundary. The fit then lands at a median mean of 5.187 with its spread on the floor, and covers 93.8% and 94.2%. The variance is finite but very large so close to the boundary, which shows as a slightly low coverage at a thousand draws; a floor further from 0.707 would trade some of the narrow proposal’s efficiency for a variance the central limit theorem reaches sooner.

Mix in a defensive component. Draw a tenth of the final sample from N(5, 1) and weight every draw against the mixture. The fitted part can collapse entirely — here it is the stalled full fit — and the defensive tenth still carries the estimate: 92.6% at a thousand draws and 93.8% at ten thousand. It is the repair that needs no knowledge of what went wrong, which is its value, and it pays for that by spending nine tenths of its draws on a proposal that contributes nothing when the fit has failed.

What an adaptive estimate should report

Whether the adaptation reached the threshold. A cross-entropy fit that stalls reports an estimate of zero with an interval of zero width, and nothing in the final interval says the adaptation never arrived. The level the last stage reached is the one number that does.

The fitted spread, against 0.707. For a normal proposal on a normal target that boundary is a closed form, and a fitted spread below it is a proposal with an infinite variance whatever its diagnostics read.

Which parameters were adapted. “The proposal was fitted by cross-entropy” describes procedures that cover 1.1%, 93.7% and 94.5% here. The difference is whether the spread was fitted, floored or fixed.

Counted, on what

The target is P(Z > 5) for a standard normal, 2.8665 × 10⁻⁷, computed from its closed form. Each run fits its own proposal from its own pilot draws — a thousand a stage, the top tenth kept, at most thirty stages — and then draws a fresh final sample of a thousand or ten thousand from the fitted proposal, so the final estimate is unbiased given the proposal and the question is only whether its interval is honest. Two thousand runs at a thousand final draws and six hundred at ten thousand, one stated seed per run, so a coverage near 94% carries a standard error of about half a point and one point. The end point’s visible error is computed from the closed form of the weight integrand, integrated by Simpson’s rule up to each run size’s typical largest draw. Every variant starts each run from the same seed, so the smoothed and the plain fits see the same first batch and differ only in how they move. Only normal proposals are fitted, only the basic cross-entropy update and its smoothed form are used, and the pilot draws are not reused in the final estimate.

Still open: an estimate that keeps the pilot draws

Every estimate here throws the pilot away. Adaptive multiple importance sampling keeps it: every draw from every stage is reweighted against the mixture of all the proposals used, and the estimate pools them. That saves the pilot’s cost and introduces the dependence this essay avoided, since the later proposals were chosen by the earlier draws now being counted, and the pooled estimate is no longer unbiased in the usual sense.

Whether the pooled estimate’s interval covers, how much the dependence costs against the draws it saves, and whether pooling rescues the stalled full fit — whose pilot draws never pass 5 but whose earliest, wide stages might — is the measurement this leaves. It belongs beside a coverage table with its own error, since any such count of coverage is itself a simulation whose error has to be stated, and beside the seed that is part of the figure, since an adaptive proposal is a seed-dependent object in a way a fixed one is not.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Adaptive procedureClosed formCoverageCross entropy methodEffective sample sizeImportance samplingInfinite varianceKullback leibler divergenceMonte CarloNormal distributionRelative errorTail probability