AQAI QuantAI research lab for systematic strategies

Automated analysis

This analysis was drafted by our research engine and has not been checked by a human editor. It may contain errors. It separates the paper’s own results from our tests, and any figures called ours come from our own backtest.

Our automated analysisOur backtest

Worst-case CVaR still buys a Markowitz frontier portfolio

Cai, Mao and Song's wealth leader also carries the highest worst-case CVaR

2026-10-08 · 7 min read · Robust portfolio optimization · US-listed stocks and ETFs

Reviewing: Conditional value-at-risk under reward-penalty mechanism with applications to robust portfolio management · Jun Cai, Tiantian Mao and Zhiqiao Song · Read it on arxiv

Our backtest of this idea

Our automated quick test, not the paper's

Daily Worst-Case CVaR with Threshold-Exceedance Penalty

Backtest period 2020-01-01 to 2024-07-01 · hypothetical, net of modelled costs

Why these figures are not the paper's (1)

Our own audit found this run does not follow the paper faithfully (7)

  • Eqs 2.1–2.2: classical moment set M(µ,Σ) and generalized set D_n(γ₁,γ₂), including its squared Mahalanobis and Frobenius bounds.: Use M(µ̂,Σ̂) as the primary ambiguity set; retain D_n only as a reference formula, without a generalized-set test. (invalidates: Section 6 Figure 3 generalized-set sensitivity findings)
  • Eqs 4.5–4.6 and Theorem 4.1: the stated analytic weights and uniqueness condition apply to the unrestricted classical set.: Preserve the formula for reference but numerically solve the same fixed-moment objective under long-only and concentration constraints. (invalidates: Theorem 4.1 unrestricted unique-minimizer guarantee and analytic allocation; paper-sample cumulative-wealth comparisons that depend on unrestricted weights)
  • Eq 5.2: generalized-set optimization includes √γ₁, γ₂Iₙ, auxiliary bounds and the robust-return constraint.: Do not run the generalized-set optimization; implement the classical fixed-moment variant instead. (invalidates: Section 6 Figure 3 generalized-set sensitivity findings)
  • Eq 4.4: mean-variance comparator minimizes ω⊤Σω subject to ω⊤e=1 and −ω⊤µ=r₀.: Use constrained minimum variance without an equality target-return constraint. (invalidates: Paper's Figure 1 comparison with its targeted-return mean-variance portfolio)

3 further finding(s) are described in the note.

These are our findings about our own implementation, not criticisms of the paper. Read the figures below as a description of what we ran.

Jan 2020Total 77.6%Jul 2024
Sharpe
0.97
Total Return
77.6%
Max Drawdown
-27.8%
CAGR
13.6%
Volatility
17.8%
Beta vs SPY
0.74
Trades
7,456

Cai, Mao and Song's wealth leader takes more tail risk. With mean and covariance treated as known, their reward-penalty CVaR rule selects a point on the Markowitz frontier. Its wealth lead over the comparators largely reflects which point it selects. The theory earns attention; the empirical claim has to be judged against that frontier.

The bonus clause changes the loss

The fund manager earns a reward at rate β when the portfolio return exceeds −d. A loss beyond d triggers a penalty at rate λ, which the authors liken to a clawback. For the institution, the resulting loss is (1−β)X + (β−λ)(X−d)+ plus a constant. With λ < β, dividing by (1−β) gives X + κ(X−d)+, where κ = (β−λ)/(1−β). It is portfolio loss with a weighted stop-loss on the same portfolio. The authors also let κ stand as a free weight on downside risk, with d set at a risk level such as a VaR or CVaR.

Theorem 3.1 computes the worst-case CVaR of this payoff across distributions with mean µ and standard deviation σ. A three-point distribution attains the worst case. For λ < β, the result has three regimes, determined by d−µ relative to σd1 and σd2. In the middle regime, it is (1−β)d + [Δ − (2−λ−β)(d−µ)]/(2(1−α)), with Δ built from σ² + (d−µ)². The formula contains earlier results as special cases:

The paper also finds a distinction absent from the usual mean-variance set: worst-case VaR and worst-case CVaR, known to coincide there, separate when λ < β.

Which frontier point gets picked?

A univariate projection lemma reduces each portfolio's worst case to the scalar problem. Theorem 4.1 puts the optimal weights on the mean-variance frontier at a mean loss s. Choosing the portfolio comes down to choosing s: it minimizes a strictly convex g(s) and is unique under the technical condition (4.6). The budget constraint is ω⊤e = 1, with weights otherwise free in Rⁿ. Shorting and leverage are allowed. We found no position limits.

The authors call the α → 0 version M-WC-EDR: mean loss plus κ times the worst-case expected stop-loss. Its s has the explicit expression (4.8), provided κ > κ = (2+2√(b0+1))/b0. At κ ≤ κ, including κ = 0, there is no solution. The denominator in (4.8) contains 4(κ+1) − b0κ², a factor that reaches zero exactly at κ. As κ approaches κ from above, s falls without bound. Expected return and variance for the chosen portfolio both rise without limit. For M-WC-EDR, the paper's advice to wealth-maximizing, risk-seeking investors to choose κ just above κ* therefore means taking ever more leverage on the frontier.

Kang et al. (2019)'s generalized set allows the mean to move within an ellipsoid of size γ1 and the covariance to move within Frobenius distance γ2. The worst case raises portfolio mean loss by √γ1 times the portfolio standard deviation under Σ̂; portfolio variance becomes ω⊤(Σ̂ + γ2·I)ω. This yields the convex program (5.2). Explicit solutions are generally unavailable, and the Figure 3 portfolios do not lie on the frontier of (µ̂, Σ̂).

Wealth and tail risk move together

The study follows 15 US stocks, the three largest in each of five sectors, using daily Yahoo data from March 2017 to February 2025. Training takes the first three years; the test runs from March 2020 to February 2025. A 756-day window supplies the moments and rolls forward one day at a time. The settings are α = 0.9, a mean-variance target of r0 = 0.0003 daily and κ = 0.5. At the three targets −d = 0.006, 0.0003 and 0.00006, M-WC-EDR's cumulative wealth is "significantly higher" than every comparator's for most of the test period. That is the paper's description of a plot, without a test behind it. M-WC-EDR also has the highest worst-case CVaR. WC-CVaR-BPL, the paper's κ-weighted CVaR model, earns a slight but consistent cumulative-wealth lead over plain worst-case CVaR.

The authors acknowledge a "trade-off between improving expected portfolio return and controlling the worst-case portfolio CVaR" in the abstract. They also say that adding downside risk to portfolio loss "can achieve higher investment returns than considering either downside risk or portfolio loss alone." The same sentence claims better risk management. Yet Figure 1 gives M-WC-EDR, the wealth leader, the highest worst-case CVaR of the six. Theorems 4.1 and 4.2, together with the same mean-variance reduction for the comparators, make that risk-management claim difficult to treat as an edge. Every Figure 1 model chooses a point on the same frontier. Higher return accompanied by higher worst-case CVaR places a portfolio further along it. A desk would want the models compared at matched worst-case CVaR, or by Sharpe ratio. We found neither. The paper prints no return, Sharpe or drawdown figures; every empirical result appears as a plot.

The plotted CVaR is model-implied sup CVaR, computed using each window's sample moments. The paper calls its out-of-sample series "expected portfolio returns" constructed from "the expected daily loss vector". We could not determine from the text whether the wealth series compounds realized daily returns.

The mean-variance benchmark makes the choice of frontier point especially visible. It produces the lowest wealth. Its worst-case CVaR falls below M-WC-EDR's, though it sits slightly above the other four models'. An equality return constraint fixes its position at the authors' chosen r0 = 0.0003 target.

The κ sweep tells a similar story. At 1, 5, 10, 15, with d = 0.03, smaller κ brings M-WC-EDR and M-WC-CVaR more wealth and more CVaR. WC-CVaR-BPL moves in the opposite direction.

Test-period information in γ1 and γ2

To set the ambiguity sizes, the authors calculate deviations window by window using "a rolling window approach over all the data". All the data includes 2020 to 2025, so test-period information enters γ1 = 0.009 and γ2 = 0.001. Covariance ambiguity is particularly consequential here: D_n(0, 0.001) produces the highest wealth, and γ2 matters more than γ1. Figure 3 lists (0.009, 0.001) twice among its settings, while the text makes a comparison with D_n(0.009, 0). Its daily return floor, r = −0.098, looks unlikely to bind.

We found no date for the market-cap ranking. The inclusion of NVDA and LLY makes the list read like a recent one; if it is recent, the universe incorporates test-period knowledge. Nor did we find a mention of transaction costs, despite daily rebalancing. The authors are candid about κ: "it is unlikely that a single value of κ can simultaneously satisfy both objectives". They leave a selection criterion to future work. We also found no reported b0, leaving readers unable to check whether κ = 0.5 exceeds κ* for models that require it.

Our long-only run, 2020 to mid-2024

We built one version from the paper's description, in one automated pass. It traded daily from 2020-01-01 to 2024-07-01 across 14 of the paper's stocks, excluding GOOG. We used the classical fixed-moment objective, α = 0.90, β = 0.5, λ = 0 (so κ = 1) and d = 0.03, with a 756-day window. We required long-only weights, capped each name at 10%, and charged $0.004 a share in commissions.

The run returned 77.56% in total, with a 0.97 Sharpe, 17.82% volatility and a −27.78% maximum drawdown over 7,456 trades. The paper supplies no return or Sharpe figure to place beside our 77.56% total return or 0.97 Sharpe. Its setup also differs substantially from ours: the authors allow unconstrained weights, while ours were long-only. With 14 names and a 10% cap, at least 10 names have to carry weight. Our optimizer consequently had little latitude to depart from a near-equal-weight mega-cap book. Its 0.74 beta to SPY suggests it still favored the defensive names within that cap.

There is a fill discrepancy in our run. We intended to trade at the next day's open; the recorded fills occurred at the close. We also did not produce worst-case CVaR (WC-CVaR) or mean-variance (M-V) comparators. This one automated pass is no verdict on the authors' work, and its result cannot establish whether the threshold penalty beats either comparator.

The theory is a real generalization. The empirical case would persuade me if M-WC-EDR beat the other Figure 1 models at matched worst-case CVaR, and if Figure 3 held up with γ1 and γ2 estimated solely on the 2017 to 2020 training window.

Our backtest stops at 2024-07-01, and everything after that date is deliberately left untouched so the same strategy can be checked out of sample later.

How our backtest worked

The steps the code we ran actually executed, from its strategy card. Ours, not the paper's — it is one automated implementation of the idea, not the authors' own.

At each decision close T:
  Keep stocks with complete, aligned 756-loss histories and a 20-session
  mean daily dollar volume of at least $10 million, using data through T.
  If the sample covariance is not positive definite, or fewer than 10
  eligible stocks can support the 10% weight cap, skip the rebalance.
  Numerically minimize the three-branch threshold-loss worst-case CVaR
  objective over weights summing to 1, with 0 &lt;= each weight &lt;= 0.10.
  If the solve fails, retain holdings and record a missed rebalance.
  The specification calls for trades at the observed T+1 adjusted open;
  skip orders with missing opens and apply the pre-order leverage check.

The specification also calls for standard worst-case CVaR and constrained minimum-variance comparators under the same eligibility and weight rules; comparator results are not supplied here.