The closed-form proxy picks a portfolio close to the exact optimum on the authors' three-asset Gaussian benchmark. At c = 0.10, its loss against the numerical minimum is 0.000028, compared with 0.001307 for the envelope portfolio. At c = 0.20, those losses are 0.000592 and 0.001126. Liu, Wang and Wang also prove why the envelope can mislead: for p > 1 and a strictly increasing benchmark quantile (Gaussian, Student-t), replacing an upper-semicontinuous nonconcave distortion with its concave envelope strictly overstates the Wasserstein worst case at every radius. The radius itself remains the obstacle to desk use, as the authors acknowledge.
How the risk measure works
A distortion function h sets the weight assigned to each loss quantile. The resulting distortion riskmetrics include VaR, RVaR, ES, the inter-quantile range and Tversky and Kahneman's inverse-S probability weighting. The authors allow h to be nonmonotone, nonconcave or discontinuous. Their ambiguity set contains loss distributions within p-Wasserstein distance ε of a reference model. For a linear portfolio, the multivariate ball projects exactly onto a one-dimensional ball around the reference loss. Its radius is ε‖ω‖, scaled by the dual norm of the weights. Under an elliptical benchmark, portfolios share the same worst-case problem, each at its own radius.
The paper develops three routes through that problem. Convexification replaces h with its concave envelope h*: the closed-form worst case is the reference risk plus ε times a norm of the envelope's quantile weights. Regularization averages h over windows of width η, solves the smoothed problems through an isotonic-projection formula, then takes η to zero. Finally, a pair of explicit candidate distributions lie inside the ball, carry error bounds and require no optimization.
This is a theoretical paper. Its experiments use synthetic references (uniform, triangular, normal, a standardized Pareto with shape 10) and a three-asset Gaussian benchmark. Mean losses are (0.080, 0.075, 0.070), volatilities are (0.12, 0.20, 0.32), and correlation is 0.3. The authors report no return, Sharpe or out-of-sample result. Our figures below therefore have no corresponding figure of theirs for comparison.
Where the envelope fails
Theorem 1 pins down when convexification is exact. For p > 1 and an upper-semicontinuous h with a nonzero envelope, the two worst cases agree if and only if the reference quantile is flat wherever h lies below its envelope. For RVaR with α = 0.95, that interval is (0.95, 1). Gaussian and Student-t projections rise strictly throughout it, making the envelope strictly conservative at every radius. The authors reach that conclusion too, writing that direct convexification "is limited in practice."
An empirical reference built from 252 daily losses faces the same issue. Its top 5% covers about 12.6 order statistics, which are almost never all equal. The required flat stretch is absent there as well.
A transport budget aimed at the tail
A Wasserstein adversary can devote its entire transport budget to the tail mass a measure reads. The paper's second explicit candidate, P_ε, mixes the reference with the envelope's worst case and requires no optimization. Its limitation shows up in the authors' VaR+ example at level 1/2 on U[0,1]. Near zero, the worst case rises like ε^(p/(p+1)) (ε^(2/3) for p = 2); P_ε gains only O(ε).
For left VaR, a jump in the reference quantile at α makes matters worse. At α = 0.99, the authors test F*, the closed-form envelope optimizer. With ε = 0.001 standard deviations, the VaR gap is 0.02182 on the continuous reference. A reference with a 0.02078 jump at α produces a gap of 0.04260. The VaR+ gap stays at 0.01182 in both cases. Against the authors' illustrative 0.05 tolerance, the bound in the jump case certifies VaR only through ε = 0.00119. At ε = 0.002, the true error is 0.05541.
RVaR at (0.95, 0.975) is easier to certify. Its bound is ES minus RVaR at the reference: 0.04330 for the uniform and 0.02686 for the triangular reference. Since the bound does not depend on ε, it certifies the 0.05 tolerance at every radius. The left-VaR bound for the jump case fails beyond ε = 0.00119. Realized RVaR gaps are 0.01647 and 0.01503 at ε = 0.01, then 0.04050 and 0.02578 at ε = 0.10.
Does the proxy choose the same book?
Mostly, in the authors' three-asset example. They search 20,301 long-only weight vectors under an upper-tail TK distortion with γ = 0.7. At radius scale c = 0.10, the exact objective chooses (0.405, 0.350, 0.245) and has value 0.123795. The explicit approximation chooses (0.385, 0.355, 0.260); evaluated under the true objective, it scores 0.123823, a loss of 0.000028. The envelope chooses (0.525, 0.330, 0.145) and scores 0.125102. Its loss is 0.001307, roughly 47 times larger.
That advantage narrows with radius. At c = 0.20, the approximation loses 0.000592 and the envelope 0.001126, under a factor of two apart. Their weights reveal opposite errors: the feasible explicit candidate understates worst-case risk, while the envelope overstates it. The exact book holds 0.375 in the low-volatility asset at c = 0.20. The approximation holds 0.320; the envelope holds 0.450. At c = 0.10 and 0.20, the cheap methods miss from opposite sides, while the approximation's loss grows from 0.000028 to 0.000592.
The abstract says the bounds can be computed without an exact solve. On a normal reference, they clear 0.05 only at ε ≤ 0.066 within the plotted range up to 2.3. The paper acknowledges that the approximation misses the tolerance at some radii; the worst grid error is about 0.151. The authors write: "When both bounds exceed 0.05, we need a direct calculation to check the actual error." Their next sentence says ε should reflect model uncertainty. The guidance would carry more weight with a data-based ε, and the paper says its radii "are not calibrated from data."
Regularization performs well in its test. For the inter-quantile range at c = 0.005, the regularized minimum drops from 0.201072 at η = 0.10 to the exact 0.177117 by η = 0.01, with matching weights. The experiment still has three assets, one period, known μ and Σ, and no trading costs.
Our monthly worst-case RVaR book, 2020 to mid-2024
We built a US equity version from daily data. On the first trading close each month, it uses a calendar-year top-50 non-ADR screen and trailing 252 daily losses. The book minimizes worst-case RVaR at (0.95, 0.975), with p = r = 2 and c = 0.05 scaled by the square root of the covariance trace. Each name has a 10% weight cap. We solve the inner problem exactly using the paper's isotonic method, with 16 optimizer starts.
From 2020-01-01 to 2024-07-01, the book returned 43.45% with a 0.58 Sharpe, 16.58% volatility and a maximum drawdown of -29.91% across 2,218 trades. Those are our figures, net of commissions of $0.004 a share.
They say little about the idea. A long-only large-cap book with 16.58% volatility and a near-30% drawdown through the 2020 and 2022 selloffs is consistent with market exposure. Its 0.70 Sortino alongside a 0.58 Sharpe shows little downside asymmetry for a tail-risk objective. We solved ordinary RVaR on the same inputs but did not trade it, leaving the effect of the Wasserstein term unknown. Our own run has two further limitations. We have not established whether the capitalization screen is point-in-time. Fills printed at the same-session close, although we intended the next session's open.
This single pass tests a construction we built from the authors' idea, rather than their work as a whole. A paired, traded comparison of worst-case and ordinary RVaR at several values of c across rolling windows would move us.
For p > 1, when the benchmark quantile rises strictly on the envelope intervals, the envelope worst case exceeds the true one at every radius. The authors warn that the replacement "can be overly conservative." On their portfolio example, the explicit approximation is cheaper and closer to the exact objective, although its advantage contracts from 47 times to under two as c doubles. A defensible, data-driven ε would make the method more usable for trading. The paper does not supply one.
Our backtest stops at 2024-07-01, and everything after that date is deliberately left untouched so the same strategy can be checked out of sample later.