The standard cross-entropy covariance is close enough to the integrability boundary to make its reported uncertainty questionable. On an eleven-factor credit book at the 99.9% level, Fang, Shen and Wang fit the covariance using exact conditional tail probabilities on the same 250,000 pilot draws. The Gaussian factor likelihood ratio then has a finite-second-moment margin of just mLR = 0.0197. Using 294 realised tail hits pushes the standard cross-entropy fit across that boundary. Its margin is -0.1158, and its minimum eigenvalue is 0.4726 against the required 0.5. The authors also give a signed eigenvector whose loadings are strictly positive for every obligor, proving that the expected-shortfall importance weight has infinite second moment in this configuration. The CEIS expected-shortfall bands consequently serve as finite-run dispersion summaries. The paper calls every band in that table nominal.

The model is familiar. Portfolio loss sums loss-given-default amounts multiplied by default indicators, with dependence supplied by a handful of common factors. Value-at-Risk contributions concern events such as {L = x} or {L >= x}. At 99.9%, ordinary Monte Carlo puts about one in a thousand paths into the tail. The exact-level event is scarcer: the paper estimates P(L = 250) at 2.75e-4. Glasserman and Li's two-stage procedure shifts the factor distribution, then exponentially twists the conditional Bernoulli defaults. Its performance depends heavily on the factor proposal. Cross-entropy calibration constructs that proposal by matching tail-conditional factor moments. Standard cross-entropy importance sampling (CEIS) weights each pilot factor draw by a binary indicator recording whether its single simulated loss exceeded the threshold.

ISCOS, the authors' method, assigns each draw a filtered Fourier-cosine estimate of q_x(u) = P(L >= x | U = u). Once the factors are fixed, defaults are independent. The conditional loss characteristic function is therefore a closed-form product over obligors, allowing the COS expansion to integrate out the default draw. The lattice requires care. Because the event is non-strict, the CDF must be evaluated at y_x = x - 0.5; evaluating at the jump makes the reconstruction return a midpoint. The remaining procedure stays unchanged: the same weighted mean and covariance fit, the same ridge (1e-8), and the same production run with the same twist. In the Student t-copula, the common state is (Z, W). Its proposal combines a Gaussian with an inverse-Gamma, and the scale component matches E[log W] and E[1/W] through a digamma equation.

The simulated benchmark is Example 4 of Glasserman (2005). It contains 100 obligors arranged in ten homogeneous blocks of ten, each with default probability 0.01. Eleven independent Gaussian factors drive the model, comprising one market-wide factor and ten industry factors. Losses-at-default are (1,1,4,4,9,9,16,16,25,25), and total loss lies on the integer lattice in [0,1100]. The market-factor loading is 0.3; the block-factor loading is 0.8. At alpha = 0.999, the Gaussian threshold is 250, while the t threshold is 504 for nu = 4. The pilot contains 250,000 draws. Production uses 250,000 per event, with common random numbers across methods.

We could not run the method on our own data. It requires obligor-level default probabilities, loss-given-default exposures, and calibrated factor loadings. It also relies specifically on conditionally independent Bernoulli defaults with an analytic conditional loss characteristic function. Equity or ETF drawdowns would substitute a different mechanism.

Calibration moments drive the ordering

The clean theoretical result is a Rao-Blackwell argument. With exact q_x, the covariance gap between the CEIS and ISCOS estimators of raw sufficient statistics equals E[R^2 q_x(1-q_x) T T'], a positive semidefinite quantity. The finite loss lattice supplies a uniform COS error bound. Calibration error then separates into O_P(M_0^{-1/2}) + O(K^{-p}), with the balancing heuristic K ~ M_0^{1/(2p)}.

That separation between statistical and numerical error is fair, though its scope is limited. The ordering applies to unnormalised moments. At the parameter level, it becomes a first-order asymptotic result through the delta method, and the constants contain 1/kappa_x. The result therefore makes no uniform claim as the event grows rarer. The paper describes ISCOS as a proposal-calibration method. It does not claim lower variance for any given CVaR or CES estimator, nor does it present COS as a direct estimator of the final risk contributions.

What the mode sweep reveals

The block structure allows exact conditional benchmarking by convolution, and this is the paper's strongest empirical feature. The mode sweep shows how badly a coarse reconstruction can distort calibration. At K = 16, relative errors in the fitted mean and covariance reach 94.59% and 58.67%. At the same resolution, factor states with exact conditional tail probability below 1e-12 receive 86.74% of total calibration weight. Raising the count to K = 32 reduces those errors to 5.67% and 8.95%, and spurious mass falls to 2.95%. At K = 1024, the errors are 2.53% and 3.07%.

The matched comparison uses K = 32. Calibration effective sample size increases from 294 binary hits to 601.8, close to the exact-weight fit's 587.60 on the same pilot. During production, event-weight ESS rises from 5337.8 to 9404.8 for the exact-level event and from 29,354.6 to 45,413.0 for the tail event. Hit rates decline slightly, from 13.87% to 13.16% and from 70.31% to 70.25%. Average nominal half-lengths contract 21.4% for CVaR, moving from 0.07623 to 0.05993, and 22.7% for CES, from 0.05068 to 0.03916. ISCOS produces narrower componentwise bands for 99 of 100 CES contributions. For CVaR, it does so for 50 of 100 contributions, effectively a coin flip on the exact-level event.

Glasserman's mean-only proposal N(mu^GL, I) follows Glasserman's own Algorithms 6.1 and 6.2 while leaving the covariance at the identity. It reaches a 75.24% CES hit rate, yet its event-weight ESS is only 1291.6. Weight balance, rather than event frequency, determines efficiency here.

The integrability check explains what saved ISCOS. Its covariance is more diffuse, with a minimum eigenvalue of 0.7296 compared with 0.4726 for CEIS. The Gaussian factor-ratio test passes for Glasserman and ISCOS and fails for CEIS; the ISCOS margin is mLR = 0.6294. The safer, more diffuse fit is also closer to the benchmark. Its d_mu is 5.67% against 7.04%, and its d_Sigma is 8.95% against 10.99%. We previously saw a similar result in a factor-graph equity model (our review).

Two warnings from the t experiment

The t run follows the same broad pattern. Calibration ESS climbs from 255.0 to 353.5, while tail-event ESS moves from 22,688.8 to 43,964.5. Average CES half-length falls by 27.0%, from 0.06891 to 0.05031. ISCOS-t gives narrower bands for all 100 CES components and 92 of 100 CVaR components.

Both proposals fail the stated second-moment condition for the common state, b_W < nu = 4. Their fitted inverse-Gamma scales are 36.198 and 33.834. The diagnostic supplies no CLT support for either set of displayed intervals. The authors qualify this finding themselves: it is a warning, rather than a complete analysis of every event-weighted estimator, because the event indicator and conditional Bernoulli twist may add tail decay. The t results also lack an exact conditional-convolution benchmark and a sweep over K, as the paper acknowledges. The implemented choice is therefore 64 modes.

Negative raw COS weights create a separate warning that proves less consequential. They account for 94.14% of weights in this run, yet clipping changes the mean weight by only 2.62e-6. Treating negative probabilities as automatic evidence of broken numerics would give the wrong reading here.

Cost barely changes

ISCOS-t calibration takes 2.656 s, compared with 0.101 s for CEIS-t. The equivalent stand-alone pipeline increases from 186.237 s to 194.226 s, a rise of 4.3%. In the Gaussian experiment, all three pipelines finish at roughly 96 to 97 s.

Those timings benefit from the block structure. The conditional characteristic function runs over ten homogeneous groups; an ungrouped calculation would run over 100 obligors. The paper states that grouping reduces the direct O(M_0 N K) cost when exposures and loadings repeat. From that expression, and the grouping factor of ten used here, we infer that a portfolio without repeated exposures incurs the undiscounted cost. The K = 16 failure also removes low mode counts as an easy economy. Timing came from one sequential run in fixed method order, and neither the processor nor the library versions were recorded.

The calibration case is convincing. Smooth conditional weights use more of the pilot, approach the exact-weight moments more closely, and produce an integrable proposal in the Gaussian experiment. Evidence for safer capital numbers remains thinner. Gaussian threshold-tail means differ by less than 0.02%, at 280.470 against 280.508, while the CVaR allocation vectors differ by 1.52% in relative L2. The practical case rests on interval and ESS diagnostics at one threshold, using one pilot and one production run per configuration.

The authors are explicit about those limits. The experiments consist of single matched runs, not repeated-run studies. Displayed intervals are nominal, and the findings do not establish a universal variance ordering. Their narrower conclusion holds: on this benchmark, smooth conditional-probability weighting improves proposal calibration.

Repeated runs across several alphas, with the inverse-Gamma scale constrained below nu so that the bands carry meaning, would change my view of the downstream claim. Until then, the portable lesson is the boundary check. Fitting a cross-entropy covariance from binary tail hits at 99.9% can leave you printing standard errors for an estimator that has none.