AQAI QuantAI research lab for systematic strategies

Automated analysis

This analysis was drafted by our research engine and has not been checked by a human editor. It may contain errors. It separates the paper’s own results from our tests, and any figures called ours come from our own backtest.

Our automated analysisOur backtest

Three Gaussian assets: a proxy within 0.0006 of the optimum

For p > 1, the envelope overstates worst-case risk; the cheaper proxy comes closer on the authors' benchmark.

2026-10-08 · 6 min read · Distributionally robust portfolio optimization · US equities

Reviewing: Robust distortion riskmetrics under Wasserstein ambiguity · Yang Liu, Qiuqi Wang and Yihan Wang · Read it on arxiv

Our backtest of this idea

Our automated quick test, not the paper's

Monthly Wasserstein-Robust RVaR versus Ordinary RVaR in US Equities

Backtest period 2020-01-01 to 2024-07-01 · hypothetical, net of modelled costs

Why these figures are not the paper's (1)

The paper reports no results of its own

This is a theoretical paper — derivations and proofs, with no measurement on market data. The backtest below is a strategy we built from its idea, not a test of anything the authors claimed.

Jan 2020Total 43.5%Jul 2024
Sharpe
0.58
Total Return
43.5%
Max Drawdown
-29.9%
CAGR
8.4%
Volatility
16.6%
Beta vs SPY
0.70
Trades
2,218

The closed-form proxy picks a portfolio close to the exact optimum on the authors' three-asset Gaussian benchmark. At c = 0.10, its loss against the numerical minimum is 0.000028, compared with 0.001307 for the envelope portfolio. At c = 0.20, those losses are 0.000592 and 0.001126. Liu, Wang and Wang also prove why the envelope can mislead: for p > 1 and a strictly increasing benchmark quantile (Gaussian, Student-t), replacing an upper-semicontinuous nonconcave distortion with its concave envelope strictly overstates the Wasserstein worst case at every radius. The radius itself remains the obstacle to desk use, as the authors acknowledge.

How the risk measure works

A distortion function h sets the weight assigned to each loss quantile. The resulting distortion riskmetrics include VaR, RVaR, ES, the inter-quantile range and Tversky and Kahneman's inverse-S probability weighting. The authors allow h to be nonmonotone, nonconcave or discontinuous. Their ambiguity set contains loss distributions within p-Wasserstein distance ε of a reference model. For a linear portfolio, the multivariate ball projects exactly onto a one-dimensional ball around the reference loss. Its radius is ε‖ω‖, scaled by the dual norm of the weights. Under an elliptical benchmark, portfolios share the same worst-case problem, each at its own radius.

The paper develops three routes through that problem. Convexification replaces h with its concave envelope h*: the closed-form worst case is the reference risk plus ε times a norm of the envelope's quantile weights. Regularization averages h over windows of width η, solves the smoothed problems through an isotonic-projection formula, then takes η to zero. Finally, a pair of explicit candidate distributions lie inside the ball, carry error bounds and require no optimization.

This is a theoretical paper. Its experiments use synthetic references (uniform, triangular, normal, a standardized Pareto with shape 10) and a three-asset Gaussian benchmark. Mean losses are (0.080, 0.075, 0.070), volatilities are (0.12, 0.20, 0.32), and correlation is 0.3. The authors report no return, Sharpe or out-of-sample result. Our figures below therefore have no corresponding figure of theirs for comparison.

Where the envelope fails

Theorem 1 pins down when convexification is exact. For p > 1 and an upper-semicontinuous h with a nonzero envelope, the two worst cases agree if and only if the reference quantile is flat wherever h lies below its envelope. For RVaR with α = 0.95, that interval is (0.95, 1). Gaussian and Student-t projections rise strictly throughout it, making the envelope strictly conservative at every radius. The authors reach that conclusion too, writing that direct convexification "is limited in practice."

An empirical reference built from 252 daily losses faces the same issue. Its top 5% covers about 12.6 order statistics, which are almost never all equal. The required flat stretch is absent there as well.

A transport budget aimed at the tail

A Wasserstein adversary can devote its entire transport budget to the tail mass a measure reads. The paper's second explicit candidate, P_ε, mixes the reference with the envelope's worst case and requires no optimization. Its limitation shows up in the authors' VaR+ example at level 1/2 on U[0,1]. Near zero, the worst case rises like ε^(p/(p+1)) (ε^(2/3) for p = 2); P_ε gains only O(ε).

For left VaR, a jump in the reference quantile at α makes matters worse. At α = 0.99, the authors test F*, the closed-form envelope optimizer. With ε = 0.001 standard deviations, the VaR gap is 0.02182 on the continuous reference. A reference with a 0.02078 jump at α produces a gap of 0.04260. The VaR+ gap stays at 0.01182 in both cases. Against the authors' illustrative 0.05 tolerance, the bound in the jump case certifies VaR only through ε = 0.00119. At ε = 0.002, the true error is 0.05541.

RVaR at (0.95, 0.975) is easier to certify. Its bound is ES minus RVaR at the reference: 0.04330 for the uniform and 0.02686 for the triangular reference. Since the bound does not depend on ε, it certifies the 0.05 tolerance at every radius. The left-VaR bound for the jump case fails beyond ε = 0.00119. Realized RVaR gaps are 0.01647 and 0.01503 at ε = 0.01, then 0.04050 and 0.02578 at ε = 0.10.

Does the proxy choose the same book?

Mostly, in the authors' three-asset example. They search 20,301 long-only weight vectors under an upper-tail TK distortion with γ = 0.7. At radius scale c = 0.10, the exact objective chooses (0.405, 0.350, 0.245) and has value 0.123795. The explicit approximation chooses (0.385, 0.355, 0.260); evaluated under the true objective, it scores 0.123823, a loss of 0.000028. The envelope chooses (0.525, 0.330, 0.145) and scores 0.125102. Its loss is 0.001307, roughly 47 times larger.

That advantage narrows with radius. At c = 0.20, the approximation loses 0.000592 and the envelope 0.001126, under a factor of two apart. Their weights reveal opposite errors: the feasible explicit candidate understates worst-case risk, while the envelope overstates it. The exact book holds 0.375 in the low-volatility asset at c = 0.20. The approximation holds 0.320; the envelope holds 0.450. At c = 0.10 and 0.20, the cheap methods miss from opposite sides, while the approximation's loss grows from 0.000028 to 0.000592.

The abstract says the bounds can be computed without an exact solve. On a normal reference, they clear 0.05 only at ε ≤ 0.066 within the plotted range up to 2.3. The paper acknowledges that the approximation misses the tolerance at some radii; the worst grid error is about 0.151. The authors write: "When both bounds exceed 0.05, we need a direct calculation to check the actual error." Their next sentence says ε should reflect model uncertainty. The guidance would carry more weight with a data-based ε, and the paper says its radii "are not calibrated from data."

Regularization performs well in its test. For the inter-quantile range at c = 0.005, the regularized minimum drops from 0.201072 at η = 0.10 to the exact 0.177117 by η = 0.01, with matching weights. The experiment still has three assets, one period, known μ and Σ, and no trading costs.

Our monthly worst-case RVaR book, 2020 to mid-2024

We built a US equity version from daily data. On the first trading close each month, it uses a calendar-year top-50 non-ADR screen and trailing 252 daily losses. The book minimizes worst-case RVaR at (0.95, 0.975), with p = r = 2 and c = 0.05 scaled by the square root of the covariance trace. Each name has a 10% weight cap. We solve the inner problem exactly using the paper's isotonic method, with 16 optimizer starts.

From 2020-01-01 to 2024-07-01, the book returned 43.45% with a 0.58 Sharpe, 16.58% volatility and a maximum drawdown of -29.91% across 2,218 trades. Those are our figures, net of commissions of $0.004 a share.

They say little about the idea. A long-only large-cap book with 16.58% volatility and a near-30% drawdown through the 2020 and 2022 selloffs is consistent with market exposure. Its 0.70 Sortino alongside a 0.58 Sharpe shows little downside asymmetry for a tail-risk objective. We solved ordinary RVaR on the same inputs but did not trade it, leaving the effect of the Wasserstein term unknown. Our own run has two further limitations. We have not established whether the capitalization screen is point-in-time. Fills printed at the same-session close, although we intended the next session's open.

This single pass tests a construction we built from the authors' idea, rather than their work as a whole. A paired, traded comparison of worst-case and ordinary RVaR at several values of c across rolling windows would move us.

For p > 1, when the benchmark quantile rises strictly on the envelope intervals, the envelope worst case exceeds the true one at every radius. The authors warn that the replacement "can be overly conservative." On their portfolio example, the explicit approximation is cheaper and closer to the exact objective, although its advantage contracts from 47 times to under two as c doubles. A defensible, data-driven ε would make the method more usable for trading. The paper does not supply one.

Our backtest stops at 2024-07-01, and everything after that date is deliberately left untouched so the same strategy can be checked out of sample later.

How our backtest worked

The steps the code we ran actually executed, from its strategy card. Ours, not the paper's — it is one automated implementation of the idea, not the authors' own.

At the first trading session's close each month:
  Refresh the calendar-year top-50 non-ADR US stock screen.
  Retain stocks with 252 synchronized close-to-close loss observations;
  require at least 10 eligible stocks.
  Estimate the loss covariance trace from that window.
  For each feasible trial weight vector (sum = 1; each weight in [0, 0.10]):
    Sort its 252 portfolio losses and construct the exact empirical RVaR grid.
    Set epsilon = 0.05 * sqrt(covariance trace) * ||weights||_2.
    Solve the positive-radius inner problem by weighted isotonic regression
    and radius root-finding; check monotonicity and radius residual.
  Minimize worst-case RVaR across 16 constrained starts; report solver failures.
  Separately solve zero-radius empirical RVaR on identical inputs.
Publish trades and marked returns for the robust book only; log ordinary-RVaR
weights and objectives as diagnostics. The platform reports MOC fills and marks.